In a nutshell
Your DNA is not one long free thread. It is packaged differently in different places, and only part of it actually codes for anything.
This subtopic is about where DNA sits (in prokaryotes, in the eukaryotic nucleus, and in mitochondria and chloroplasts), what a gene is and how it codes for proteins through the triplet code, and why much of the DNA in a eukaryotic nucleus does not code for polypeptides at all.
Assumed knowledge: Nucleic acids (DNA structure).
Core content
Where DNA is found: prokaryotes, eukaryotes and organelles
DNA is packaged differently depending on where it is, and AQA tests these differences precisely.
- In prokaryotic cells, DNA molecules are short, circular and not associated with proteins.
- In the nucleus of eukaryotic cells, DNA molecules are long, linear and associated with proteins, called histones. Together, a DNA molecule and its associated proteins form a chromosome.
- The mitochondria and chloroplasts of eukaryotic cells also contain DNA. Like the DNA of prokaryotes, this is short, circular and not associated with protein.
Notice that the DNA in mitochondria and chloroplasts is like prokaryotic DNA, not like the DNA in the eukaryotic nucleus that surrounds it.
| Feature | Prokaryotic DNA | Eukaryotic nuclear DNA | Mitochondrial / chloroplast DNA |
|---|---|---|---|
| Length | short | very long | short |
| Shape | circular | linear | circular |
| Associated with histone proteins? | no | yes (forms a chromosome) | no |
| Contains introns? | no | yes | no |
Both prokaryotic and eukaryotic DNA are double-stranded (a double helix). The differences above are about length, shape and packaging, not about the number of strands.
Genes and loci
A gene is a base sequence of DNA that codes for:
- the amino acid sequence of a polypeptide, or
- a functional RNA (including ribosomal RNA (rRNA) and transfer RNA (tRNA)).
A functional RNA is an RNA molecule that is used as it is and is not translated into a polypeptide. So a gene does not always code for a protein.
A gene occupies a fixed position, called a locus, on a particular DNA molecule.
The triplet code
A sequence of three DNA bases, called a triplet, codes for a specific amino acid.
- There are 4 bases (A, T, C, G). Read in threes, they give 43 = 64 possible triplets.
- Those 64 triplets code for the 20 amino acids found in proteins.
- Because a triplet is three bases long, the length of a coding sequence in bases is always three times the number of amino acids it codes for.
The three-base unit is called a triplet in DNA. When the equivalent unit is on mRNA it is called a codon, which you meet in the next subtopic on protein synthesis. In this note the term is triplet.
Features of the genetic code
The genetic code is the set of DNA base triplets that determines the amino acid sequence of a polypeptide. It has three features you must be able to explain.
| Feature | What it means | Why it matters |
|---|---|---|
| Universal | The same triplet codes for the same amino acid in all organisms | Evidence that all organisms share a common ancestor |
| Non-overlapping | Each base is part of only one triplet; triplets are read one after another, not sharing bases | The sequence is read in a fixed order, giving one correct polypeptide |
| Degenerate | More than one triplet codes for the same amino acid | A change in one base may still code for the same amino acid, reducing the effect of some mutations |
For example, the amino acid glycine can be coded for by more than one triplet, so the code is degenerate.
Still don't get it? · degenerate vs non-overlapping
Think of a keypad with letters on the buttons, like an old phone. On that keypad, the buttons 2, 7 and 9 might all give you the letter you want: several different buttons, one same result. That is degenerate: more than one input, same output.
Non-overlapping is a totally separate idea about how you read a line of buttons. Imagine the number 2-2-8-5-1-3 written out. Non-overlapping means you chop it into fixed blocks of three and never reuse a digit: (228)(513). You do NOT read (228), then slide along and read (285), then (851). Each digit belongs to one block only.
Now the exam version, and the two are easy to swap by mistake:
- Degenerate = more than one triplet codes for the same amino acid.
- Non-overlapping = each base is part of only one triplet.
A classic exam trap is stating degenerate the wrong way round ("one triplet codes for more than one amino acid"). That is false, and it describes nothing real. Learn it as "more than one triplet, same amino acid".
Coding and non-coding DNA: exons, introns and repeats
In eukaryotes, much of the nuclear DNA does not code for polypeptides.
- Between genes there are non-coding multiple repeats of base sequences.
- Even within a gene, only some sequences, called exons, code for amino acid sequences.
- Within the gene, these exons are separated by one or more non-coding sequences, called introns.
This is why the DNA of a eukaryotic gene contains more bases than are needed just to code for the amino acids: the introns and the non-coding repeats do not code for amino acids.
Still don't get it? · why so much DNA does not code for a polypeptide
Picture the recipe for one cake printed in the middle of a huge cookbook. Most of the pages are blank filler or repeated adverts, and even the recipe itself has chatty side-notes printed between the actual instructions. If you counted every letter in the book, only a small fraction would be the real cooking steps.
Rebuild it for DNA:
- The whole "cookbook" is the DNA molecule. Long stretches between the genes are the filler and repeats: non-coding multiple repeats.
- One "recipe" is a gene. Even inside that gene, the real instructions come in chunks called exons (they code for amino acids).
- The chatty side-notes printed between the exons are introns: non-coding sequences inside the gene.
Exam version: in eukaryotes much of the nuclear DNA does not code for polypeptides. Between genes there are non-coding multiple repeats; within a gene the coding exons are separated by non-coding introns. So a gene has more bases than the bare number needed to code for its amino acids.
Worked examples
A. Working between bases and amino acids
A section of the coding sequence of a gene is 30 DNA bases long. What is the maximum number of amino acids it could code for?
Each amino acid is coded for by a triplet of three bases, so divide the number of bases by three:
Now the reverse. A polypeptide is 141 amino acids long. What is the minimum number of DNA bases in the coding sequence?
Two things to notice, both of which examiners reward:
- Going from bases to amino acids you divide by 3; going from amino acids to bases you multiply by 3. Doing the wrong operation is a common lost mark.
- The whole gene is longer than 423 bases, because it also contains introns and non-coding sequences. The calculation gives the coding length, not the length of the whole gene.
B. Model comparison answer: "Compare and contrast the DNA in the nucleus of a eukaryotic cell with the DNA in a prokaryotic cell."
A compare-and-contrast answer must make statements that mention both types together, not two separate lists. Aim for one similarity and then clear differences:
- Both are made of nucleotides joined by phosphodiester bonds (a similarity: the nucleotide structure is identical).
- Eukaryotic nuclear DNA is longer, whereas prokaryotic DNA is shorter.
- Eukaryotic nuclear DNA is linear, whereas prokaryotic DNA is circular.
- Eukaryotic nuclear DNA is associated with histone proteins, whereas prokaryotic DNA is not associated with proteins.
- Eukaryotic DNA contains introns, whereas prokaryotic DNA does not.
The lesson: each numbered point compares the same feature in both types. Writing "eukaryotic DNA is long, it is linear, it has histones" and then a separate list for prokaryotes does not score, because no comparison has been made.
Common exam mistakes
- Writing that DNA is an "alpha helix". DNA is a double helix; an alpha helix is a level of protein (secondary) structure. The two are unrelated.
- Stating that prokaryotic DNA is single-stranded, or that only eukaryotic DNA is a double helix. Both prokaryotic and eukaryotic DNA are double-stranded.
- In a compare-and-contrast question, comparing the cells instead of the DNA (for example "in prokaryotes the DNA floats in the cytoplasm"), or writing two separate lists rather than comparative statements. This is the single biggest cause of lost marks on this topic.
- Defining a gene as "a strand of DNA", "part of a chromosome", "the whole DNA molecule", or "a protein". The mark needs a base sequence (length) of DNA that codes for a polypeptide (or a functional RNA).
- Naming the position of a gene as an "allele" or an "exon". The fixed position of a gene is its locus.
- Defining universal as "the code is found everywhere". The mark needs the same triplet codes for the same amino acid in all organisms.
- Stating degenerate the wrong way round ("one triplet codes for more than one amino acid"), or confusing it with non-overlapping. Degenerate = more than one triplet codes for the same amino acid.
- Writing that triplets "produce" or "make" amino acids. Triplets code for amino acids; the code does not build them.
- Explaining why a gene has more bases than three times the number of amino acids by saying "the code is degenerate". The extra bases are due to introns and non-coding sequences, not degeneracy.
- Failing to name the non-coding sequences within a gene. They are introns; the coding sequences are exons.
Key definitions
- Gene: a base sequence (length) of DNA that codes for the amino acid sequence of a polypeptide, or for a functional RNA (such as rRNA or tRNA).
- Locus: the fixed position of a gene on a particular DNA molecule (chromosome).
- Chromosome: a DNA molecule together with its associated proteins (histones), found in the eukaryotic nucleus.
- Histones: the proteins that eukaryotic nuclear DNA is associated with (wound around).
- Triplet: a sequence of three DNA bases that codes for a specific amino acid.
- Genetic code: the sequence of DNA base triplets that codes for the amino acid sequence of a polypeptide.
- Universal (of the genetic code): the same triplet codes for the same amino acid in all organisms.
- Non-overlapping (of the genetic code): each base is part of only one triplet.
- Degenerate (of the genetic code): more than one triplet codes for the same amino acid.
- Exon: a base sequence of a gene that codes for (part of) the amino acid sequence of a polypeptide.
- Intron: a non-coding base sequence found within a gene.
Specification
- I can state that prokaryotic DNA molecules are short, circular and not associated with proteins.
- I can state that eukaryotic nuclear DNA is long, linear and associated with proteins called histones, and that a DNA molecule with its associated proteins forms a chromosome.
- I can state that mitochondria and chloroplasts contain DNA that, like prokaryotic DNA, is short, circular and not associated with protein.
- I can define a gene as a base sequence of DNA that codes for the amino acid sequence of a polypeptide or a functional RNA (including rRNA and tRNA).
- I can state that a gene occupies a fixed position, called a locus, on a particular DNA molecule.
- I can state that a sequence of three DNA bases, called a triplet, codes for a specific amino acid.
- I can explain what is meant by the genetic code being universal, non-overlapping and degenerate.
- I can explain that much eukaryotic nuclear DNA does not code for polypeptides, including non-coding multiple repeats between genes and non-coding introns that separate the coding exons within a gene.
Related notes
Ready to test yourself?
Put DNA, genes and chromosomes into practice with exam-style questions and full mark schemes.
Practise DNA, genes and chromosomes