Key points
- Edman degradation removes and identifies one amino acid at a time from the N-terminus of a peptide, reading its sequence directly from the chemistry.
- Each cycle is slightly less than 100% efficient, so the signal decays and background rises. In practice 20โ40 residues can usually be read, rarely more than about 50.
- The method needs a free N-terminal amine. Acetylated, formylated or pyroglutamate-blocked peptides, and head-to-tail cyclic peptides, cannot be sequenced without deblocking.
- Mass spectrometry has replaced Edman sequencing for most purposes, but Edman remains a reference method for confirming the N-terminus of therapeutic peptides and proteins.
From Sanger to Edman
The first complete sequence of a protein was that of insulin, determined by Frederick Sanger between 1945 and 1955. Sanger labelled the N-terminal amino acid with fluorodinitrobenzene, hydrolysed the whole peptide and identified the labelled residue. Because hydrolysis destroyed the rest of the chain, every step required new material, and the sequence had to be assembled from many overlapping fragments. The work earned Sanger his first Nobel Prize in 1958 and proved that proteins have defined sequences.
Pehr Edman, working in Lund, found a way to remove only the N-terminal residue and leave the rest of the chain intact, ready for the next cycle. He published the method in 1950, and in 1967, with Geoffrey Begg, described an automated "sequenator" that ran the cycles unattended. For the following three decades Edman degradation was the principal way to sequence peptides and proteins.
The chemistry of one cycle
Each cycle has three chemical steps:
- Coupling. Under mildly basic conditions, around pH 9, phenyl isothiocyanate (PITC) reacts with the free N-terminal amine to form a phenylthiocarbamoyl (PTC) peptide. The basic pH is needed because only the unprotonated amine is reactive.
- Cleavage. In anhydrous trifluoroacetic acid, the sulfur atom of the PTC group attacks the carbonyl carbon of the first peptide bond. The first residue is cut off as an anilinothiazolinone (ATZ) derivative, and the remaining peptide, one residue shorter, is left with a new free N-terminus. Water is excluded at this stage so that the other peptide bonds are not hydrolysed.
- Conversion. The unstable ATZ derivative is extracted and converted in aqueous acid into the more stable phenylthiohydantoin (PTH) amino acid.
The PTH amino acid is identified by reversed-phase HPLC, where each of the twenty derivatives elutes at a characteristic time, and the shortened peptide re-enters the next cycle. A modern instrument completes a cycle in roughly 30โ45 minutes and needs only low picomole amounts of peptide.
The key to the method is the cleavage step. It is selective for the first peptide bond because only that bond is next to the PTC group that attacks it. This is a neat example of the kinetic stability of the peptide bond: the other bonds survive anhydrous acid unharmed.
Why the read length is limited
No chemical reaction is perfectly complete. If coupling or cleavage fails for a small fraction of molecules in a cycle, those molecules fall one residue behind. In the next cycle they release the residue that the majority released in the previous one, and the signal becomes a mixture. At the same time, a small amount of random acid hydrolysis elsewhere in the chain creates new N-termini, raising a background of every amino acid.
The fraction of molecules still in step is the repetitive yield raised to the number of cycles:
| Repetitive yield | After 10 cycles | After 20 cycles | After 30 cycles | After 50 cycles |
|---|---|---|---|---|
| 90% | 35% | 12% | 4% | 0.5% |
| 95% | 60% | 36% | 21% | 8% |
| 98% | 82% | 67% | 54% | 36% |
| 99% | 90% | 82% | 74% | 60% |
Good instruments achieve repetitive yields of 90โ95% or somewhat better, so by cycle 30 the correct residue is a small signal among lagging and background ones. This is the same exponential arithmetic that limits solid-phase synthesis, applied in reverse. Long proteins were therefore sequenced as overlapping fragments generated by specific cleavage with trypsin or cyanogen bromide.
Residues and modifications that cause trouble
Most amino acids give clean PTH derivatives, but some do not:
| Residue or feature | Problem | Usual solution |
|---|---|---|
| Cysteine | Free Cys gives an unstable, poorly recovered derivative; cystine appears as a mixed product | Reduce and alkylate before sequencing, then look for the alkylated PTH-Cys |
| Serine and threonine | Partial dehydration during cleavage lowers the yield | Identify from the dehydrated by-product |
| Phosphoserine, glycosylated residues | Modification is lost or the derivative is not extracted; appears as a blank cycle | Combine with mass spectrometry |
| Proline | Cleavage after proline is slower, increasing lag | Extended cleavage time |
| Blocked N-terminus | No reaction at all | Enzymatic or chemical deblocking, or internal fragments |
The last row is the most serious limitation. In eukaryotic cells, most cytosolic proteins carry an acetylated N-terminus, and many peptide hormones begin with pyroglutamate. Neither has a free amine for PITC to react with. Pyroglutamate can be removed enzymatically with pyroglutamate aminopeptidase; acetyl groups are harder to remove. The guide to post-translational modifications describes both. Head-to-tail cyclic peptides have no N-terminus at all and must be opened first.
Edman degradation and mass spectrometry compared
Since the 1990s, tandem mass spectrometry has become the dominant method for sequencing peptides, especially in combination with genome databases. The two methods have different strengths:
| Edman degradation | Tandem MS | |
|---|---|---|
| Sample | A single purified peptide or protein | Complex mixtures of thousands of peptides |
| Throughput | About one residue per 30โ45 minutes | Thousands of peptides per hour |
| Read length | Typically 20โ40 residues from the N-terminus | Short peptides; long proteins via digestion |
| Leu vs Ile | Distinguished: their PTH derivatives separate by HPLC | Not distinguished by standard CID |
| Blocked N-terminus | Fails | No problem |
| Modifications | Often lost or unidentified | Detected by mass shift and usually located |
| Database needed | No; reads directly | Usually; de novo possible but harder |
| Interpretation | Direct and unambiguous | Computational, with statistical confidence |
The two methods are complementary rather than competing. Edman gives a direct, database-independent reading of the N-terminal sequence, which is exactly what is required to confirm the identity and N-terminal integrity of a recombinant or synthetic therapeutic, and it is recognised in regulatory guidelines for biopharmaceutical characterisation. Mass spectrometry provides coverage, sensitivity to modifications and speed.
Where the method is used today
- Quality control of biopharmaceuticals: confirming the N-terminal sequence and detecting truncated or extended forms.
- Characterising unknown proteins from organisms without sequenced genomes, where database searching is impossible.
- Checking signal-peptide cleavage sites and processing of precursor proteins.
- Distinguishing isomers such as Leu and Ile, which mass spectrometry struggles with.
Edman chemistry has also returned in a new form. Single-molecule "fluorosequencing", described by Marcotte and colleagues in 2018, labels specific residues with fluorescent dyes, immobilises millions of peptides on a surface, and removes N-terminal residues one at a time with Edman chemistry while imaging the loss of fluorescence. Nanopore-based methods for protein sequencing are also under development.
Frequently asked questions
How much peptide does Edman sequencing need?
Modern instruments work with low picomole amounts, typically a few to tens of picomoles, which for a 2 kDa peptide is in the nanogram range. The sample must be highly pure, because every peptide present contributes its own residues to each cycle.
Can Edman degradation sequence from the C-terminus?
No. C-terminal chemical sequencing methods exist but have never matched Edman chemistry in reliability. The C-terminus is usually confirmed by mass spectrometry or by digestion with carboxypeptidases.
What if a cycle shows two amino acids?
Either the sample contains two peptides, or the peptide is heterogeneous at that position, for example because of partial processing. Lag from the previous cycle also contributes, so the amount of each residue has to be compared with the cycles before and after.
Why is anhydrous acid used for cleavage?
Because water would hydrolyse other peptide bonds, creating new N-termini and raising the background in every later cycle. Without water, only the first bond, activated by the neighbouring thiocarbamoyl group, is cleaved.
References
- Edman P (1950) Method for determination of the amino acid sequence in peptides. Acta Chemica Scandinavica 4:283โ293.
- Edman P, Begg G (1967) A protein sequenator. European Journal of Biochemistry 1:80โ91.
- Sanger F, Thompson EOP (1953) The amino-acid sequence in the glycyl chain of insulin. Biochemical Journal 53:353โ366.
- Swaminathan J, Boulgakov AA, Hernandez ET, et al. (2018) Highly parallel single-molecule identification of proteins in zeptomole-scale mixtures. Nature Biotechnology 36:1076โ1082.