⚠ For in-vitro research purposes only. Strictly not for human or veterinary use.
Australian Peptide Lab kangaroo logo
Lab Technique

How to Read Peptide Sequences: Codes, Termini, Modifications and Cyclic Peptides

Peptide notation decoded: one- and three-letter codes, N→C direction, Ac- and -NH2, D-amino acids, Aib, Nal, lactams and disulfides, with worked examples.

By the APL Research Team · Updated · First published · 6 min read

A peptide sequence is a list of amino-acid symbols written from the N-terminus to the C-terminus, with prefixes, suffixes and brackets for anything that is not a plain linear chain of L-amino acids. Read correctly, a line such as Ac-Nle-cyclo[Asp-His-D-Phe-Arg-Trp-Lys]-NH2 gives the structure, the expected mass and several clues to the chemistry. The symbols and conventions used here follow the IUPAC-IUB recommendations for amino acids and peptides [1], with the informal variants that appear on supplier listings noted where they differ.

The two alphabets

Each of the 20 standard amino acids has a three-letter code and a one-letter code. Three-letter codes are joined by hyphens (Gly-His-Lys); one-letter codes are written as a continuous string (GHK). Both describe the same peptide sequence.

One-letterThree-letterAmino acidResidue mass, mono (Da)Note
GGlyGlycine57.021Achiral
AAlaAlanine71.037
SSerSerine87.032
PProProline97.053Ring constrains the backbone
VValValine99.068
TThrThreonine101.048
CCysCysteine103.009Forms disulfides
LLeuLeucine113.084Same mass as Ile
IIleIsoleucine113.084Same mass as Leu
NAsnAsparagine114.043Deamidation site
DAspAspartic acid115.027Acidic
QGlnGlutamine128.0590.036 Da lighter than Lys
KLysLysine128.095Basic; common attachment point
EGluGlutamic acid129.043Acidic
MMetMethionine131.040Oxidation site
HHisHistidine137.059Metal binding
FPhePhenylalanine147.068Aromatic
RArgArginine156.101Basic
YTyrTyrosine163.063Aromatic
WTrpTryptophan186.079Aromatic; oxidation site

The letters that do not match the name are the usual trip-ups: F is Phe, Y is Tyr, W is Trp, K is Lys, R is Arg, D is Asp, E is Glu, N is Asn and Q is Gln. Residue masses are the free amino acid minus one water, because each peptide bond releases water; a peptide's mass is the sum of its residues plus one water for the free termini.

Direction, numbering and the termini

Sequences are always written N-terminus first. Position 1 is the residue with the free amino group, and numbers count towards the C-terminus.

  • H- marks a free N-terminal amine and -OH a free C-terminal acid. Both are often left implicit.
  • Ac- is an N-terminal acetyl group, adding C2H2O (+42.011 Da).
  • -NH2 is a C-terminal amide, replacing OH with NH2 (−0.984 Da). Many signalling peptides are amidated, and the difference is visible by mass spectrometry: a confiscated GHRH analogue was found by high-resolution MS to be only partly amidated, which the analysts read as a sign of poor quality [2].

Fragment notation gives the numbered range of a parent sequence: GRH(1–29)-NH2 is the first 29 residues of growth hormone-releasing hormone with an amidated C-terminus, and GRH(1–40)-OH a 40-residue form with a free acid [3]. Numbering always follows the parent molecule. GLP-1 analogues are numbered from 7, because the circulating active hormone is GLP-1(7–36)amide [4]; semaglutide's "Aib8" is therefore the second residue of the chain, which is exactly where DPP-IV cuts [4, 5].

Substitutions are written in square brackets before the parent name. [Pro1, Val14]-hGHRH, for example, is human GHRH with proline at position 1 and valine at position 14; in that confiscated sample the "Pro1" was in fact an added N-terminal proline, an example of why the bracket shorthand needs checking against the full sequence [2].

Stereochemistry: D-amino acids

Every standard residue except glycine is chiral, and L is assumed unless stated. A D-residue is written with a D- prefix: the growth hormone-releasing hexapeptide GHRP is His-D-Trp-Ala-Trp-D-Phe-Lys-NH2 [6]. Informal one-letter strings sometimes use lower-case letters for D-residues, and some database synonyms drop stereochemistry altogether. PubChem, for instance, lists ipamorelin both with and without its D- prefixes.

D- and L-forms have identical masses, so mass spectrometry cannot tell them apart. Racemisation during synthesis produces exactly this kind of diastereomeric impurity [7], which is why a stereochemical impurity shows up only as a separate HPLC peak. D-residues are used deliberately to resist proteolysis; peptide half-life research covers that chemistry.

Non-standard residues

Many research peptides contain residues outside the standard 20. They have no one-letter code and are written with their abbreviations inside a three-letter sequence.

AbbreviationNameWhat it isCatalogue example
Aibα-Aminoisobutyric acid (2-methylalanine)Achiral; two methyl groups on the α-carbon; residue 85.053 DaPosition 1 of ipamorelin [8]; position 8 of semaglutide [5]
2-Nal3-(2-Naphthyl)alanineBulky aromatic; residue 197.084 DaD-2-Nal at position 3 of ipamorelin [8]
NleNorleucineUnbranched isomer of Leu and Ile; same formulaN-terminal Ac-Nle of Melanotan II and PT-141
D-XaaAny D-amino acidMirror-image residue; same mass as the L-formD-Phe in Melanotan II; D-Ala2 in CJC-1295 (no DAC)

Side-chain and terminal modifications

Modifications attached to a residue are written in brackets after it, or described in words when the group is large.

  • Side-chain acylation. Semaglutide is derivatised at Lys26 [5] with an 18-carbon fatty diacid linked through γ-glutamate and two short ethylene-glycol spacers, often abbreviated Lys26(…) in sequence diagrams.
  • Reactive groups. CJC-1295 with DAC carries an added C-terminal lysine bearing a 3-maleimidopropionamide group on its ε-amine [9].
  • N-terminal acyl groups other than acetyl. Tesamorelin is human GHRH(1–44)NH2 with a trans-3-hexenoyl group on Tyr1 [10].
  • Pyroglutamate. An N-terminal Gln or Glu can cyclise to pyroglutamate (pGlu), one of the degradation products seen in peptide medicines [7].
  • Metal complexes. GHK-Cu is the tripeptide GHK complexed with copper(II) [11]; the metal is noted after the sequence rather than within it.

Cyclic peptides and disulfides

Cyclisation is written by enclosing the ring in brackets, with cyclo or c, and the type of link stated or implied.

Side-chain lactam. Melanotan II is Ac-Nle-cyclo[Asp-His-D-Phe-Arg-Trp-Lys]-NH2: an amide bond joins the side-chain carboxyl of Asp2 to the side-chain amine of Lys7, closing a six-residue ring and releasing one water (−18.011 Da). PT-141 (bremelanotide), a synthetic α-MSH analogue [12], has the same ring with a free C-terminal acid instead of the amide. That one change is the whole structural difference between them: C50H69N15O9 (1024.2 g/mol) for Melanotan II against C50H68N14O10 (1025.2 g/mol) for PT-141. A forensic laboratory characterising seized samples distinguished the two by accurate mass and fragmentation [13].

Disulfide bridge. Two cysteine thiols oxidise to a disulfide, removing two hydrogens (−2.016 Da). None of the catalogue peptides contain cysteine, so the standard textbook example is oxytocin, CYIQNCPLG-NH2 with a Cys1–Cys6 disulfide, written as cyclo(1–6) or with a line joining the two cysteines. Its formula, C43H66N12O12S2 (1007.2 g/mol), is two hydrogens lighter than the reduced form.

Worked examples from the catalogue

PeptideAs writtenHow to read itFormulaAvg MW (g/mol)
BPC-157GEPPPGKPADDAGLV15 residues, linear, free terminiC62H98N16O221419.5
GHK-CuGHK·CuTripeptide Gly-His-Lys plus copper(II) [11]C14H24N6O4 (peptide)340.4 (peptide)
EpithalonAEDGAla-Glu-Asp-Gly, linearC14H22N4O9390.4
SemaxMEHFPGPACTH(4–7), Met-Glu-His-Phe, extended by Pro-Gly-Pro; an ACTH(4–10) analogue [14]C37H51N9O10S813.9
SelankTKPRPGPTuftsin (Thr-Lys-Pro-Arg) extended by Pro-Gly-Pro; a tuftsin analogue [15]C33H57N11O9751.9
IpamorelinAib-His-D-2-Nal-D-Phe-Lys-NH2Five residues, two non-standard, two D, amidated [8]C38H49N9O5711.9
Melanotan IIAc-Nle-cyclo[Asp-His-D-Phe-Arg-Trp-Lys]-NH2Acetylated, lactam ring 2–7, amidatedC50H69N15O91024.2
PT-141Ac-Nle-cyclo[Asp-His-D-Phe-Arg-Trp-Lys]-OHAs Melanotan II but free acidC50H68N14O101025.2
CJC-1295 (no DAC)Tyr-D-Ala-Asp-Ala-Ile-…-Ile-Leu-Ser-Arg-NH229 residues: GHRH(1–29)NH2 with D-Ala2, Gln8, Ala15, Leu27C152H252N44O423367.9
Tesamorelintrans-3-hexenoyl-hGHRH(1–44)NH2Full-length GHRH, N-acylated [10]C221H366N72O67S~5136

Formulas and masses were calculated from the sequences and cross-checked against PubChem records where they exist. Ipamorelin is a useful exercise. It came from a series of compounds lacking the central Ala-Trp dipeptide of GHRP-1 [8], and adding the residue formulas of Aib, His, D-2-Nal, D-Phe and Lys, plus water, minus the O-to-NH change of the amide, gives C38H49N9O5 exactly.

Mass shifts worth memorising

ChangeFormula changeMonoisotopic shift (Da)
C-terminal amide (vs acid)−O +NH−0.984
Deamidation (Asn→Asp, Gln→Glu)−NH +O+0.984
N-terminal acetylation+C2H2O+42.011
Disulfide formation−2H−2.016
Side-chain lactam−H2O−18.011
Met oxidation+O+15.995
D- for L-residuenone0

These shifts, together with the residue table above, are what a mass spectrometrist uses to assign unexpected peaks; what HPLC testing measures covers the complementary separation side.

For the bench

  • Paste standard sequences into the peptide molecular weight calculator, which handles N-acetyl, C-amide and disulfides; enter L for Nle, and add Aib, Nal or lactam corrections by hand.
  • Use the average molecular weight for weighing and molarity, and the monoisotopic mass for interpreting spectra; the molarity calculator needs the average figure.
  • Compare the sequence on the vial label, the certificate and the literature you are following, including termini and stereochemistry.
  • Remember that sequence notation describes the peptide, not the powder: counter-ions and water are additional, as net peptide content explains, and how peptides are made shows where those come from.
  • For what these reagents are and how they differ from medicines, see what are research peptides.

Frequently asked questions

Are GHK and Gly-His-Lys the same thing?

Yes. GHK is the one-letter form of glycyl-L-histidyl-L-lysine, a tripeptide found in human plasma, saliva and urine that is proposed to act as a complex with copper(II) [11]. In GHK-Cu the copper is a bound metal ion, not a residue, so it does not appear in the sequence itself; it does change the mass of the material in the vial. See the GHK-Cu page.

What do H- and -OH mean at the ends of a sequence?

They mark free termini: H- is the free amino group on the first residue and -OH the free carboxylic acid on the last. They are often omitted, so GEPPPGKPADDAGLV and H-Gly-Glu-…-Val-OH describe the same BPC-157 molecule. What matters is when they are replaced: Ac- for an acetylated N-terminus, -NH2 for a C-terminal amide.

What does a lower-case letter in a one-letter sequence mean?

In informal notation it usually marks a D-amino acid, so 'a' would be D-alanine. The formal convention writes D- before the three-letter code, as in His-D-Trp-Ala-Trp-D-Phe-Lys-NH2 for the original growth hormone-releasing hexapeptide [6]. Some listings drop stereochemistry altogether, so if a sequence matters, check it against a source that states it, because D and L forms have the same mass.

Why does a mass spectrum show a peptide about 1 Da away from the expected mass?

Two common causes sit almost exactly 0.984 Da apart from the target. A free C-terminal acid instead of an amide adds 0.984 Da, as seen in a confiscated GHRH analogue whose C-terminal amidation was incomplete [2]; deamidation of Asn or Gln to Asp or Glu also adds 0.984 Da. High-resolution MS and HPLC separate the possibilities; see understanding certificates of analysis.

Is TB-500 the same molecule as thymosin beta-4?

The name is not used consistently. A doping-control laboratory identified the N-acetylated 17–23 fragment of thymosin β4, Ac-LKKTETQ (about 889 Da), in a product sold as TB-500 [16], whereas full-length thymosin β4 has 43 residues and a mass near 4,963 Da. The two differ more than fivefold in mass, so the identity stated on the batch documentation matters. The TB-500 research guide covers the literature.

Can the molecular weight calculator handle Aib, Nal or norleucine?

Not directly; it covers the 20 standard residues with N-acetyl, C-amide and disulfide options. Norleucine has the same formula as leucine, so entering L gives the right mass, and D-residues weigh the same as L-residues. Aib (residue 85.053 Da monoisotopic) and 2-naphthylalanine (197.084 Da) have to be added by hand, and a side-chain lactam needs 18.011 Da subtracted.

References

  1. 1.IUPAC-IUB Joint Commission on Biochemical Nomenclature (JCBN) IUPAC-IUB Joint Commission on Biochemical Nomenclature (JCBN). Nomenclature and symbolism for amino acids and peptides. Recommendations 1983. Biochem J. 1984. PubMed 6743224
  2. 2.Esposito S, Deventer K, Van Eenoo P. Identification of the growth hormone-releasing hormone analogue [Pro1, Val14]-hGHRH with an incomplete C-term amidation in a confiscated product. Drug Test Anal. 2014. PubMed 25283153
  3. 3.Frohman LA, Downs TR, Heimer EP, et al. Dipeptidylpeptidase IV and trypsin-like enzymatic degradation of human growth hormone-releasing hormone in plasma. J Clin Invest. 1989. PubMed 2565342
  4. 4.Deacon CF, Johnsen AH, Holst JJ. Degradation of glucagon-like peptide-1 by human plasma in vitro yields an N-terminally truncated peptide that is a major endogenous metabolite in vivo. J Clin Endocrinol Metab. 1995. PubMed 7883856
  5. 5.Lau J, Bloch P, Schäffer L, et al. Discovery of the Once-Weekly Glucagon-Like Peptide-1 (GLP-1) Analogue Semaglutide. J Med Chem. 2015. PubMed 26308095
  6. 6.Bowers CY, Sartor AO, Reynolds GA, et al. On the actions of the growth hormone-releasing hexapeptide, GHRP. Endocrinology. 1991. PubMed 2004615
  7. 7.D'Hondt M, Bracke N, Taevernier L, et al. Related impurities in peptide medicines. J Pharm Biomed Anal. 2014. PubMed 25044089
  8. 8.Raun K, Hansen BS, Johansen NL, et al. Ipamorelin, the first selective growth hormone secretagogue. Eur J Endocrinol. 1998. PubMed 9849822
  9. 9.Jetté L, Léger R, Thibaudeau K, et al. Human growth hormone-releasing factor (hGRF)1-29-albumin bioconjugates activate the GRF receptor on the anterior pituitary in rats: identification of CJC-1295 as a long-lasting GRF analog. Endocrinology. 2005. PubMed 15817669
  10. 10.Ferdinandi ES, Brazeau P, High K, et al. Non-clinical pharmacology and safety evaluation of TH9507, a human growth hormone-releasing factor analogue. Basic Clin Pharmacol Toxicol. 2007. PubMed 17214611
  11. 11.Pickart L, Vasquez-Soltero JM, Margolina A. GHK Peptide as a Natural Modulator of Multiple Cellular Pathways in Skin Regeneration. Biomed Res Int. 2015. PubMed 26236730
  12. 12.Dhillon S, Keam SJ. Bremelanotide: First Approval. Drugs. 2019. PubMed 31429064
  13. 13.Mestria S, Odoardi S, Frison G, et al. LC-HRMS characterization of the skin pigmentation and sexual enhancers melanotan II and bremelanotide sold on the black market of performance and image enhancing drugs. Drug Test Anal. 2021. PubMed 33245851
  14. 14.Dolotov OV, Karpenko EA, Seredenina TS, et al. Semax, an analogue of adrenocorticotropin (4-10), binds specifically and increases levels of brain-derived neurotrophic factor protein in rat basal forebrain. J Neurochem. 2006. PubMed 16635254
  15. 15.Kolomin T, Morozova M, Volkova A, et al. The temporary dynamics of inflammation-related genes expression under tuftsin analog Selank action. Mol Immunol. 2014. PubMed 24291245
  16. 16.Esposito S, Deventer K, Goeman J, et al. Synthesis and characterization of the N-terminal acetylated 17-23 fragment of thymosin beta 4 identified in TB-500, a product suspected to possess doping potential. Drug Test Anal. 2012. PubMed 22962027

This article summarises published research for educational purposes. It is not medical advice. Compounds sold by Australian Peptide Lab are research reagents for in-vitro laboratory use only, not for human or veterinary use.

Research compounds discussed

Related research guides

Australian owned & operatedHPLC + mass-spec tested batchesDispatched express from Australian stockSecure Australian card payments