Codons
What Are Codons?
Codons are sequences of three consecutive nucleotides in DNA or RNA that specify a single amino acid or signal the end of protein synthesis. There are 64 in total: 61 that specify amino acids and 3 that act as stop signals. Codons are the reading unit of translation, the step in which a ribosome moves along a messenger RNA transcript and assembles a polypeptide chain according to the sequence it encounters.
The triplet organization was established experimentally in the early 1960s, and it follows from a counting argument. Four nucleotide letters taken two at a time yield 16 combinations, too few for 20 amino acids, while three at a time yield 64, which is more than sufficient. The surplus makes the code degenerate: most amino acids are specified by two, four, or six alternative codons. Reading is non-overlapping and begins at a defined start codon, usually AUG, which fixes the reading frame for everything downstream. An insertion or deletion that is not a multiple of three shifts that frame and typically destroys the encoded protein.
The Genetic Code and Its Structure
The National Human Genome Research Institute describes a codon as a trinucleotide forming a unit of genomic information. The mapping from codon to amino acid is nearly universal across the tree of life, with a small number of documented variants in mitochondria and in certain ciliates and bacteria. Degenerate codons for the same amino acid usually differ at the third position, an arrangement that buffers the effect of point mutations, since many third-position substitutions change nothing about the protein produced. Transfer RNA molecules perform the physical decoding: each carries an anticodon that pairs with a codon in the ribosomal A site while delivering its charged amino acid. Wobble pairing at the third position lets a single tRNA species read more than one codon, which is why cells need fewer tRNA genes than there are sense codons.
Codon Usage Bias
Synonymous codons are not used with equal frequency. Every genome shows characteristic preferences, and a review of codon usage bias attributes them to genome GC content, gene expression level and length, codon position and context, recombination rate, mRNA secondary structure, and the relative abundance of the matching tRNAs. The consequences reach past translation speed. Work on codon usage and co-translational protein folding shows that clusters of rarely used codons slow the ribosome at particular points in a transcript, giving nascent domains time to fold before downstream sequence emerges. Because of effects like these, synonymous substitutions that leave the amino acid sequence untouched can still change protein yield, structure, and function, a point developed in the analysis of the codon usage code for gene expression and protein folding.
Codon Optimization and Genome Engineering
Applied genomics manipulates codon choice directly. Codon optimization rewrites a gene using the preferred synonymous codons of a chosen expression host, which commonly raises recombinant protein yield in bacterial, yeast, insect, or mammalian systems. The same technique shapes messenger RNA therapeutics, where codon choice interacts with secondary structure and nucleoside modification to affect translation and stability. At larger scale, genome recoding projects reassign whole codons across an organism, freeing a codon for incorporation of a non-canonical amino acid or rendering a strain resistant to viruses whose genomes still rely on the original assignment. Naive optimization can misfire by removing the translational pauses that folding depends on, so current design tools weigh speed against fidelity.
Applications
Codon analysis and engineering have applications in a range of fields, including:
- Recombinant protein and biologics manufacturing
- Messenger RNA vaccine and therapeutic design
- Synthetic biology and expanded genetic codes
- Comparative genomics and phylogenetic inference
- Clinical variant interpretation, including synonymous variants
- Gene therapy vector design
- Metabolic engineering of microbial production strains