Lab-in-a-Tab

Protein Synthesis

Your DNA is just a library of recipes. How does a cell read those recipes and build the actual machines — the proteins — that make you work?

DNARibosomeGenetic Code
Try thisPress Build and watch the ribosome move. Read each three-letter codon and the bead it adds. What special codon makes it stop?
What you're seeingThe row of letters is the mRNA — the recipe. The rounded shape sliding along it is the ribosome, reading three letters (a codon) at a time. Each codon adds one coloured bead — an amino acid — to the growing chain below. That chain is the protein.
What to notice
Three letters name one bead, and the chain grows one bead per codon until a STOP codon ends it. Just four letters, read in threes, spell out every protein your body builds — the same simple code in every living thing on Earth.

How your body reads DNA to build itself

Junior level — plain language, no maths

Every cell in your body carries a full copy of your DNA — a colossal instruction manual written in just four chemical "letters". But DNA doesn't do anything on its own. It's a library of recipes, and the things those recipes build are proteins: the tiny molecular machines that digest your food, carry oxygen in your blood, fight off germs and hold your whole body together.

To cook up a protein, the cell works in two steps. First it makes a throwaway copy of just the one recipe it needs, written in a related molecule called mRNA — a bit like photocopying a single page so the precious original never leaves the library. That copy then travels to a molecular workshop called a ribosome.

The ribosome reads the mRNA three letters at a time. Each little three-letter word, called a codon, names one amino acid — and amino acids are the beads that string together to make a protein. The ribosome shuffles along, calling out codon after codon, and a growing chain of amino acids clicks into place. When the chain is complete it folds up into a precise 3-D shape, and that shape is what lets the finished protein do its job. Watch it happen in the simulation below.

Things worth knowing

  • If you read your DNA aloud at one letter per second, it would take you nearly a hundred years to finish — 3 billion letters in every cell.
  • A ribosome adds about 10–20 amino acids every second, and a single cell can run millions of ribosomes at once, churning out proteins non-stop.
  • There are only 20 different amino acids, yet strung together in different orders they build every one of the ~100,000 kinds of protein in your body.

The central dogma: transcription and translation

Student level — the core equations

The flow of genetic information follows the central dogma: DNA → RNA → protein. In transcription, an enzyme called RNA polymerase unzips a stretch of DNA and builds a matching strand of messenger RNA, swapping the base thymine (T) for uracil (U). This mRNA carries the message out of the nucleus to the ribosomes in the cytoplasm.

The message is written in the genetic code: each codon of three bases specifies one amino acid. With four bases there are \(4^3 = 64\) possible codons but only 20 amino acids, so the code is redundant — most amino acids have several codons. One codon, AUG, doubles as the "start" signal (and codes methionine); three others (UAA, UAG, UGA) are "stop" signals that end the chain.

In translation, the ribosome clamps onto the mRNA and reads it codon by codon. For each codon a matching tRNA — carrying the right amino acid and a complementary three-base anticodon — docks in, and the ribosome links its amino acid onto the growing chain, then ratchets forward. Reach a stop codon and the finished polypeptide is released to fold into a working protein. It is a molecular assembly line of astonishing speed and fidelity.

Key Formulas

Central dogma\(\text{DNA} \to \text{RNA} \to \text{protein}\)
Codon size\(4^3 = 64\ \text{codons}\)for 20 amino acids
Start codon\(\text{AUG}\)also codes Met
Stop codons\(\text{UAA, UAG, UGA}\)
Base pairing\(\text{A–U},\ \text{G–C}\)RNA uses U not T

Things worth knowing

  • The genetic code is nearly universal — the same codons mean the same amino acids in a bacterium, a banana and a human, evidence that all life shares one ancestor.
  • tRNA molecules are the "adaptors" Francis Crick predicted before they were found: one end reads the codon, the other carries the matching amino acid.
  • mRNA vaccines work by delivering a lab-made mRNA that your own ribosomes translate into a viral protein, training your immune system without any live virus.

The code, fidelity, folding and regulation

Scholar level — full mathematical depth

01Why the code is degenerate — and robust

The 64-to-20 mapping is not random. Codons for the same amino acid usually differ only in their third base, and Crick's wobble hypothesis explains why: pairing at the third position is loose, so a single tRNA can read several synonymous codons. The upshot is a code buffered against error — many point mutations are silent, and even mistranslations tend to swap in a chemically similar amino acid. The genetic code appears optimised to minimise the damage of mistakes.

02Getting it right: translational fidelity

The ribosome makes an error only about once in \(10^{4}\) codons, far better than raw codon-anticodon binding energies allow. It buys that accuracy through kinetic proofreading: a correct tRNA is checked twice, with an irreversible, energy-burning step in between that gives wrong tRNAs extra chances to fall off before the bond is sealed. Fidelity, here, is paid for in GTP — accuracy costs energy.

03The ribosome is a ribozyme

For decades the ribosome was assumed to be a protein enzyme. The 2000 crystal structures revealed the opposite: the catalytic core that forms the peptide bond is built entirely of RNA, with no protein within reach of the reaction. The ribosome is a ribozyme. This is a molecular fossil of the RNA world — a time before DNA and proteins when RNA both stored information and did chemistry — and it won the 2009 Nobel Prize in Chemistry.

04Folding, and folding gone wrong

A protein's function lives in its three-dimensional fold, and much of that folding begins co-translationally, as the chain is still emerging from the ribosome. Chaperone proteins shepherd the process and shield sticky intermediates. When folding fails, the debris can aggregate — the molecular root of Alzheimer's, Parkinson's and prion diseases. Predicting the fold from sequence alone stumped biology for fifty years until AlphaFold cracked it in 2021.

05One gene, many proteins

In eukaryotes the DNA-to-protein path is heavily edited. Freshly made RNA is spliced: introns are cut out and exons stitched together, and alternative splicing lets a single gene yield dozens of different proteins from different exon combinations — which is how ~20,000 human genes specify a far larger proteome. Add capping, tailing and chemical base modifications and the "one gene, one protein" slogan collapses entirely.

06Turning the volume up and down

Cells control not just what they translate but how much and when. MicroRNAs silence specific mRNAs, riboswitches sense metabolites and fold to gate their own translation, and ribosomes can stall, frameshift or reinitiate under stress. This regulatory layer is where synthetic biology now intervenes — engineering mRNAs, rewriting codons, even adding wholly new amino acids to the code — turning the reading of genes into something we can deliberately program.

Key Formulas

Translation error rate\(\sim 10^{-4}\ \text{per codon}\)via kinetic proofreading
Wobble pairing\(\text{loose 3rd-base match}\)
Peptide bond\(\text{-COOH} + \text{H}_2\text{N-} \to \text{amide} + \text{H}_2\text{O}\)
Catalysis\(\text{rRNA (a ribozyme)}\)
Splicing\(\text{introns out, exons joined}\)
Proteome expansion\(\sim\!20{,}000\ \text{genes} \to 10^5\ \text{proteins}\)

Things worth knowing

  • Determining the ribosome's atomic structure — and proving its catalytic heart is RNA, not protein — won the 2009 Nobel Prize in Chemistry, confirming it as a relic of the RNA world.
  • AlphaFold (2021) solved the 50-year protein-folding problem, predicting 3-D structures from amino-acid sequence with near-experimental accuracy for over 200 million proteins.
  • Alternative splicing lets one gene make many proteins: the human Dscam gene can, in principle, be spliced into over 38,000 distinct protein variants.

Sources

Full article on Wikipedia ↗