S3-2
The double helix: backbone, base pairing, antiparallel strands
DNA is the most famous molecule in biology, and almost everyone can picture the twist. The twist is the least important thing about it. Underneath the pretty spiral sits one idea so useful that all of life's copying, reading, and repair are built on top of it. That idea is complementary base pairing, and this lesson is about earning it slowly, because if you understand it, a huge amount of molecular biology stops being memorization and starts being obvious.
Two strands, wound together
DNA is not one thread. It is two, running side by side and twisted around each other into a shape called a double helix. Think of a ladder that someone grabbed by the ends and gave a gentle spin.
Each strand has a repeating outer rail called the backbone, made of a sugar and a phosphate group over and over. Between the two backbones sit the rungs. Each rung is a pair of small molecules called bases, one reaching in from each strand, meeting in the middle.
Here is a detail worth deriving rather than accepting. The backbones are on the outside and the bases point inward. Why that way around and not the reverse? The backbone carries phosphate groups that are negatively charged and happy in water. The bases are flat, oily rings that would rather not touch water at all. So the molecule folds the way oil and water always sort themselves: the water-avoiding bases tuck into the protected core and stack on top of each other, while the water-loving backbone faces out into the cell's watery interior. The structure is not arbitrary. It is what you get when you let these parts settle.
The diagram below shows this layout. Notice the two outer rails and the colored rungs between them. Ignore the twist for a moment and just read it as a ladder.
The pairing rule, and why it is forced
There are exactly four bases in DNA. Their names are adenine, thymine, cytosine, and guanine, which everyone shortens to A, T, C, and G. On any single strand they can appear in any order at all, and that order is the information. A strand is a string over a four-letter alphabet.
Now the load-bearing fact. When two strands come together to make rungs, the bases do not pair up at random. They obey a strict rule:
A pairs with T. C pairs with G.
Always. An A on one strand sits across from a T on the other, and a C always sits across from a G. This is complementary base pairing, and it is the deepest single idea in molecular biology. Let us derive why the rule is this rule and not some other.
First, shape. The four bases come in two sizes. A and G are large, built from two fused rings (chemists call these purines). C and T are small, built from one ring (pyrimidines). A rung has to span the fixed gap between the two backbones. If you paired two large bases you would get a rung that is too wide and would bulge the backbones apart. Pair two small ones and the rung is too short to reach across. So every rung must be one large base plus one small base to keep the ladder a constant width. That single geometric constraint already rules out A with G, and C with T.
Shape alone still leaves options: A is large and C is small, so by width they could fit. They do not pair. The second filter is hydrogen bonding. A hydrogen bond is a weak attraction between a hydrogen atom on one molecule and a receptive atom on another. Any one of them is feeble, but line several up in the right places and together they hold firmly. Each base has a specific pattern of little hooks along its edge: some positions offer a hydrogen (a donor), some positions want one (an acceptor). A's pattern of donors and acceptors is a mirror image of T's, so the two edges lock together with two hydrogen bonds. G's pattern is the mirror of C's, and those two lock with three hydrogen bonds. A and C are the right combined size, but their edges do not line up: donor meets donor, and nothing grips. That is why A does not pair with C. The rule is written by geometry and chemistry, not by a lookup table someone chose.
Complementary means the other strand is not a mystery
Because the pairing rule is strict and deterministic, something powerful follows. If I show you one strand, you already know the other one. You do not need to see it.
Suppose one strand reads, in order, A T G C. Walk along it and apply the rule to each base. Across from A goes T. Across from T goes A. Across from G goes C. Across from C goes G. The partner strand is forced to be T A C G. No information about the second strand had to be stored anywhere. It was implied the whole time by the first.
This is redundancy in the exact sense a programmer means it. The molecule carries every piece of information twice, once as itself and once as its complement. Nothing is lost if you keep only one strand, because the other is fully recoverable by a fixed transform.
Now make it concrete. Type a strand into the tool below and watch the partner appear. The row labeled reverse complement is the other strand, computed by the pairing rule (we will unpack the reverse part in the next section). Before you reveal the answer in the little exercise, predict the complement of the first base yourself. If you can do that reliably, you own the rule.
Predict it yourself: what base pairs with the first base of the coding strand (A)?
The tool also shows an mRNA row, which is a preview of the next lesson. You can ignore it for now and focus on how the reverse complement tracks your edits.
Antiparallel: the strands have direction, and it is opposite
So far the two strands sound symmetric. They are not, and the asymmetry matters more than the twist ever will.
A strand has a direction. The backbone is built from a repeating unit that is not the same at both ends, the way an arrow is not the same at both ends. One end of a strand exposes a chemical group attached to what chemists number the fifth carbon of the sugar. The other end exposes a group on the third carbon. By long convention these ends are named the 5 prime end and the 3 prime end (said "five prime" and "three prime"). You do not need the organic chemistry. You need this: a strand is directional, it has a distinct head and tail, and biology always names it running from 5 prime to 3 prime.
Here is the twist that is actually a twist worth caring about. In the double helix the two strands run in opposite directions. Where one strand points its 5 prime end, the other strand has its 3 prime end, and vice versa. They are antiparallel, like a two-lane road where the lanes carry traffic in opposite directions. This is why the partner of A T G C is not simply T A C G laid out the same way. Read in its own natural 5 prime to 3 prime direction, the partner strand comes out as the reverse of that complement. That is exactly why the tool above calls the other strand the reverse complement: take the complement, then read it backwards.
Why does direction matter so much? Because every machine that reads DNA or builds a new strand can only work in one direction: it reads and extends a strand from 5 prime toward 3 prime, never the other way. A strand is best pictured as a directed, typed stream that its readers can only traverse forward. The 5 prime and 3 prime labels are not decoration. They tell the machinery which way is downstream. Hold that idea, because it is the reason DNA copying has the odd, lopsided mechanics you will meet in the next module.
Why complementarity is the whole trick
Now the payoff, and the reason all this careful setup was worth it.
Take a double helix and pull the two strands apart, unzipping the rungs down the middle. You now hold two single strands. But neither one is missing information, because each strand fully specifies its partner by the pairing rule. Give each lonely strand to the cell's machinery, let it lay down the forced complementary base across from every position, and you get two complete double helices where you had one. Each new helix is one old strand plus one freshly built partner. That is how a cell copies its genome, and it works only because the two strands were complementary backups of each other the whole time. Unzip, and each half is a template for rebuilding the other.
The same redundancy powers repair. If a base on one strand gets damaged or chemically altered, the intact partner strand still carries the correct pairing information. Repair machinery can cut out the bad base and rebuild the right one by looking across the rung. The molecule stores your genetic information twice on purpose, so that a hit to one copy is survivable.
The mirrored-backup analogy, and where it breaks
Calling complementarity a redundant backup is a good analogy, and like every analogy it has an edge where it lies. Three ways it misleads. First, the backup is not an identical copy, it is a deterministic transform (complement, reversed), so you cannot just diff the two strands and expect them to match. Second, the two copies are not stored on independent disks, they are physically zipped together in one molecule, so a force that damages the region can hit both strands at the same spot, and then the backup is gone too (a double-strand break, which is genuinely dangerous). Third, when the machinery repairs a mismatch it must guess which strand holds the correct base, and it can guess wrong, writing the error permanently into both strands. A perfect backup never guesses. This one does. The analogy earns its keep for building intuition, and you should retire it the moment you start reasoning about repair failures.
Key terms
- double helix
- The shape of DNA, two strands wound around each other with their backbones outside and their paired bases meeting in the middle.
- complementary base pairing
- The strict rule that A pairs only with T and C pairs only with G across the two strands, which lets one strand fully specify the other.
- hydrogen bond
- A weak attraction between a hydrogen on one molecule and a receptive atom on another. A-T pairs use two, C-G pairs use three.
- purine and pyrimidine
- The two base sizes. Purines (A and G) are large two-ring bases, pyrimidines (C and T) are small one-ring bases, and every rung pairs one of each.
- antiparallel
- The property that the two DNA strands run in opposite directions, so one strand's 5 prime end lines up with the other strand's 3 prime end.
- 5 prime and 3 prime ends
- The two chemically distinct ends of a strand. Sequences are written and machinery reads and builds them only from 5 prime toward 3 prime.
- reverse complement
- The partner strand written in its own 5 prime to 3 prime direction, obtained by complementing each base and then reversing the order.
Where this leaves you
Two strands, backbones out, bases in. Bases pair by a rule that shape and hydrogen bonding force on them, A with T and C with G, so one strand is always a full transform of the other. The strands run in opposite directions and are read only 5 prime to 3 prime. And that complementary redundancy is not a curiosity, it is the exact reason a cell can copy and repair its genome at all. Next we watch the cell put this to work by reading a strand and making an RNA copy of it.
Check yourself
1. In DNA, which base pairs are allowed?
2. A and C are the right combined size to fit as a rung, yet they do not pair. Why not?
3. One strand reads, in its 5 prime to 3 prime direction, ATGC. Written in its own standard 5 prime to 3 prime direction, the partner strand is:
4. Why does complementary base pairing make copying DNA possible?