S5-3
Transcription: DNA to messenger RNA
Your genome is the master copy, and the cell guards it like a read-only archive. It rarely edits that master, and it never lets the master leave the safe. So how does a single gene ever get used? The cell makes a working copy of just that one gene, uses the copy, and throws the copy away when it is done. This lesson is about how that copy gets made, one base at a time. The process is called transcription, and the copy it produces is a strand of RNA.
Here is the analogy to hold onto, because it is close: transcription is like reading or checking out a copy of one function from a big read-only repository, without touching the original. You do not modify the source. You get a copy you can work with. But hold the analogy loosely, because it breaks in a specific way we will name right now. The copy is not a faithful, same-format duplicate that sticks around. It is written in a different alphabet, it copies only one of the two strands, and it is disposable.
The machine that reads a gene
The copy is made by a molecular machine called RNA polymerase. It is an enzyme, which for our purposes means a protein that does one job well. Its job is to sit down on the DNA, unzip a short stretch of the double helix, read the bare bases underneath, and build a matching new strand next to them, one base at a time, moving along the gene like a print head.
The new strand it builds is RNA (ribonucleic acid), a single-stranded cousin of DNA. RNA is almost the same 4-letter chemistry you already know, with one swap you must remember. Where DNA uses the base T (thymine), RNA uses the base U (uracil) instead. U behaves just like T did: it pairs with A. Everywhere else the pairing is unchanged. G pairs with C, C pairs with G, and A pairs with U.
RNA polymerase builds the copy by base-pairing against the DNA it is reading. It looks at each DNA base and lays down the RNA base that pairs with it. So when the DNA it is reading shows an A, the polymerase adds a U to the RNA. When the DNA shows a T, it adds an A. When the DNA shows a G, it adds a C, and a C gets a G. That single rule, "add the complementary base," is the whole engine of transcription. The one surprise for a newcomer is that a U turns up wherever the DNA being read had an A.
Only one strand carries the message
DNA is double-stranded. The two strands are complements of each other, running in opposite directions. That raises an obvious question: when the polymerase copies a gene, which of the two strands does it actually read?
The answer is that for any given gene, the polymerase reads only one of the two strands. That strand is called the template strand, because the RNA is built as its complement, the way a mold is the complement of the thing cast in it. The other strand, the one that just sits there unread for this gene, is called the coding strand.
Walk it, base by base
The single most common confusion in this whole topic is the relationship between the mRNA, the template strand, and the coding strand. The fastest cure is to walk one tiny example by hand. Take a coding strand that reads ATG, and write it out with its two partners.
coding strand 5' A T G 3' (reads like the mRNA)
template strand 3' T A C 5' (the polymerase reads this strand)
mRNA copy 5' A U G 3' (a U sits where the template had an A)
Read down the three rows. The polymerase reads the template strand (T, A, C) and pairs each base, giving the mRNA A, U, G. Now look at the top row and the bottom row. The mRNA (AUG) is a letter-for-letter match of the coding strand (ATG), with U standing in for T. That is not luck. Both the mRNA and the coding strand are complements of the same template strand, so they must come out identical apart from the T-to-U swap.
This gives you a shortcut you will use constantly. To get the mRNA of a gene, take its coding strand and change every T to a U. You do not have to flip to the template and re-complement in your head. The coding strand already reads like the message. (This particular message, AUG, is the start codon, the first word lesson S5.5 will teach you to read.)
The promoter says start here
One more piece is missing. The genome is enormous. How does the polymerase know where a gene begins, so it copies the gene and not random neighboring junk? The answer is a stretch of DNA sequence sitting just ahead of the gene, called the promoter. The promoter is a landing pad. It is a recognizable pattern that RNA polymerase (with helpers) binds to, and it marks the spot where transcription should start and which direction to head.
The promoter is what makes transcription addressable. Because a start signal sits in front of every gene, the cell can copy exactly one gene at a time instead of dumping the whole genome. It is also where control lives: how easy a promoter is to bind sets how often that gene gets copied, which is how the cell turns genes up and down. We will spend a whole module (S8) on that control, so for now just hold the promoter as the "start here" marker.
The output is messenger RNA
The finished copy of a protein-coding gene has a name: messenger RNA, or mRNA. The "messenger" is not decoration. This strand carries the gene's message out from the DNA to the machine that will build the protein (the ribosome, coming in S5.5). DNA holds the plan, mRNA carries a working copy of one page of the plan to the shop floor, and the protein gets built from that.
That disposability is a feature, not a bug. Because each mRNA is temporary, the cell controls how much of a protein it makes largely by controlling how many mRNA copies it transcribes and how long they survive. Steady demand means a steady stream of fresh copies. Stop transcribing, and the existing messages decay away and the protein stops being made. The archive never changed. Only the flow of copies did.
Time to make some copies yourself. In the tool below, type or edit a coding DNA strand and watch it compile to its mRNA. Try to predict each output before it settles: an A stays A, a G stays G, and every T should turn into a U. Then use the built-in prediction step to call the complementary base before it is revealed.
Predict it yourself: what base pairs with the first base of the coding strand (A)?
Key terms
- transcription
- The process of copying one gene's DNA into a strand of RNA, without altering the DNA.
- RNA polymerase
- The enzyme that reads a gene's template strand and builds a complementary RNA strand base by base.
- RNA
- Ribonucleic acid, a single-stranded relative of DNA that uses the base U (uracil) in place of T.
- uracil (U)
- The RNA base that stands in for thymine and pairs with A.
- template strand
- The single DNA strand the polymerase actually reads for a given gene, so the RNA is its complement.
- coding strand
- The unread partner strand, whose sequence matches the mRNA except that T is written as U.
- promoter
- A DNA sequence just ahead of a gene that marks where transcription starts and where polymerase binds.
- messenger RNA (mRNA)
- The RNA copy of a protein-coding gene that carries the message from DNA to the ribosome.
Why U instead of T at all?
It is fair to ask why life bothers with two nearly identical letters, T for DNA and U for RNA, when U would pair with A perfectly well in both. The leading explanation, and it is an explanation rather than settled dogma, ties to error control in the permanent archive. Uracil is the cheaper base to make, and it is also what the common damage of another base (cytosine breaking down) turns into. If DNA used U normally, the cell could not tell a legitimate U from a damage-created one. By using T, which is essentially U with a small chemical tag, DNA makes damage-created U stand out as an error to be repaired. The disposable message can afford the cheaper letter, U, because it is not the thing being protected for life. Treat this as the best current story, not a law you can derive from first principles.
Check yourself
1. RNA is built with almost the same four bases as DNA, with one swap. Which base does RNA use in place of thymine (T)?
2. A gene's coding strand reads ATGC (5 prime to 3 prime). What is the mRNA sequence?
3. The mRNA of a gene looks just like its coding strand. So which strand does RNA polymerase physically read to build that mRNA?
4. What is the main job of a promoter?