Sign in

In a nutshell

Genetic diversity can be compared within a species or between species. The most accurate way is to compare the molecules that carry the genetic information itself: DNA, mRNA and the proteins they code for.

This subtopic is about those comparisons (and how they replaced the older method of judging by observable characteristics), and about how we quantify variation within a species using random sampling, the mean and the standard deviation.

Assumed knowledge: Genetic diversity and adaptation, DNA and protein synthesis, Species and taxonomy.

Core content

Four ways to compare genetic diversity

Genetic diversity within or between species can be compared using:

  • the frequency of measurable or observable characteristics (the older method: e.g. height, flower colour)
  • the base sequence of DNA
  • the base sequence of mRNA
  • the amino acid sequence of the proteins encoded by DNA and mRNA

The last three are molecular methods. They compare the genetic material (or its direct products) rather than the visible features it produces, so they give a far more accurate picture of how closely two organisms are related.

The rule that runs through all of them:

  • More similar sequence = more closely related = more recent common ancestor.
  • More differences in sequence = less closely related = a less recent (earlier) common ancestor.

The reason is mutation: base sequences change by mutation, and mutations build up over time. Two organisms that shared a common ancestor recently have had less time for differences to accumulate, so their sequences are more alike.

Why observable characteristics were replaced

The old approach inferred genetic differences from measurable or observable characteristics. It has two weaknesses:

  • Observable characteristics are affected by the environment, not just the genes. Two genetically identical plants can look different if grown in different light or soil.
  • Different species can share a similar feature without being closely related, so appearances can mislead.

Gene technology now lets us investigate the DNA base sequence directly, which is more accurate, so inferring differences from observable characteristics has been replaced by direct DNA sequencing. (You are not expected to know how the sequencing technology works.)

Comparing DNA, mRNA and amino acid sequences

All three molecular methods use the same logic (fewer differences means a closer relationship), but they are not equivalent.

ComparisonWhat is comparedNote
DNA base sequencethe sequence of bases in DNAshows the most variation, because it includes introns (non-coding DNA) as well as coding regions
mRNA base sequencethe sequence of bases in mRNAmRNA is made only from coding regions (introns are removed), so it shows less variation than DNA; useful when DNA is not available
Amino acid sequencethe sequence of amino acids in a proteinthe amino acid sequence is determined by the DNA base sequence, so more similar proteins mean more similar DNA and a closer relationship

To compare amino acid sequences fairly you must compare the same protein in each organism. A common choice is cytochrome c, because it is used in respiration and so is found in all organisms that respire (all eukaryotes), which lets very different organisms be compared.

Comparing amino acid sequences underestimates the true genetic variation, because the genetic code is degenerate: more than one DNA base triplet can code for the same amino acid. Two species can therefore have an identical amino acid sequence but different DNA base sequences.

Still don't get it? · why comparing proteins misses genetic variation

Think of a finished cake. Two bakers can follow different recipes (one uses caster sugar, the other icing sugar ground down) and still hand you cakes that taste and look identical. If you only inspect the cakes, you would swear the recipes were the same. You cannot see the differences that got cancelled out on the way.

Now the biology. DNA is the recipe; the protein is the cake. The genetic code is degenerate, which means several different base triplets can code for the same amino acid. So two organisms can end up with exactly the same amino acid sequence (the same cake) even though their DNA base sequences differ (different recipes). Comparing the protein hides those DNA differences. DNA also contains introns (non-coding stretches) that never reach the protein at all, and these vary too.

The exam version: comparing amino acid sequences underestimates genetic diversity because the code is degenerate, so different DNA base sequences can produce the same amino acid sequence. Comparing DNA base sequences directly reveals more variation.

Required skill: quantitative investigations of variation

To measure variation within a species you collect numerical data, then summarise it with a mean and a standard deviation. There are three parts.

1. Collect data from random samples.

Sampling must be random to remove bias, so the sample is representative of the whole population and valid conclusions can be drawn. A worked method for sampling an area:

  • Lay out a grid over the area (for example, two long tape measures at right angles as the axes) and number the positions along each axis.
  • Use a random number generator (or a calculator, or a random number table) to produce pairs of numbers, and use each pair as the coordinates for a sample point.
  • Place a quadrat (or take a measurement) at each coordinate, and take a large sample (many readings), so the results are representative and anomalies can be identified.

2. Calculate the mean.

xˉ=∑xn\bar{x} = \frac{\sum x}{n}

Quote the mean to an appropriate number of significant figures; do not spread it over five or six decimal places.

3. Calculate and interpret the standard deviation.

The standard deviation measures the spread of the data around the mean:

  • a small standard deviation means the values are closely grouped around the mean, so there is little variation (the data are consistent)
  • a large standard deviation means the values are spread out, so there is more variation

The standard deviation is usually drawn as error bars on a bar chart. For normally distributed data, about 95% of values lie within ±2 standard deviations of the mean.

You compare two means by looking at whether their error bars overlap:

  • if the error bars (standard deviations) do not overlap, the difference between the means is likely to be significant (a real difference, not due to chance)
  • if the error bars do overlap, the difference is likely to be due to chance and is not significant

A statistical test is needed to confirm significance, but the overlap gives a strong indication.

Mean leaf length in two light conditions (error bars = 1 SD)02468Mean leaf length / cmShadedExposedMean

Here the error bars do not overlap (4.2 + 0.3 = 4.5 is below 6.8 − 0.4 = 6.4), so the difference between the two means is likely to be significant.

Note: you will not be asked to calculate a standard deviation in a written paper. You are expected to interpret means and standard deviations, as above.

Still don't get it? · what overlapping error bars actually tell you

Imagine two people step on two different bathroom scales. One reads 70 kg, the other 72 kg. But each scale is only trustworthy to about ±3 kg. So the first person is really "somewhere between 67 and 73", and the second is "somewhere between 69 and 75". Those ranges overlap, so you genuinely cannot say the two people are different weights: the 2 kg gap could just be the scales being imprecise.

Now swap in the biology. The mean is your best single estimate. The error bar (the standard deviation) is the "give or take" around it, the spread of the real readings. If the error bars of two groups overlap, the true means could actually be the same, so any apparent difference could be down to chance. If the error bars do not overlap, the two groups really do sit apart, so the difference is unlikely to be chance.

The exam version: if the error bars (standard deviations) do not overlap, the difference between the means is likely to be significant and not due to chance; if they overlap, it is likely due to chance. A statistical test confirms it.

Worked examples

Model 3-mark answer: interpreting amino acid sequence data.

The table shows the number of amino acid differences in the same respiratory protein between species P and four other species.

Species compared with PAmino acid differences
A2
B5
C14
D27

Question: "Use the data to suggest the relationships between P and species A to D."

  1. Species A has the fewest amino acid differences from P (only 2), so A and P have the most similar amino acid sequence, and therefore the most similar DNA base sequence.
  2. This means A and P are the most closely related and share the most recent common ancestor.
  3. Species D has the most differences (27), so D is the least closely related to P and shares a less recent (earlier) common ancestor.

The lesson: each point must connect the number of differences to a relationship (and, for full credit, back to the DNA sequence). Stating "A has 2 differences" alone earns nothing.

Model 2-mark answer: interpreting a mean and standard deviation.

Using the leaf-length chart above. Question: "The mean leaf length is greater in the exposed site. Do the data support the conclusion that light affects leaf length?"

  1. The mean leaf length is greater in the exposed site (6.8 cm) than in the shaded site (4.2 cm).
  2. The error bars (standard deviations) do not overlap, so the difference between the means is likely to be significant and not due to chance, which supports the conclusion.

Common exam mistakes

  • Saying random sampling makes the results "a fair test" or "controls variables". The credited reason is that it removes bias so the sample is representative of the population.
  • Describing the sampling method incompletely: forgetting the first step (generating a grid / coordinates), or numbering each individual organism instead of using a coordinate system. Some students wrongly describe mark, release and recapture instead.
  • Defining standard deviation using the word "range". The mark requires the spread of the data around the mean ("range" is rejected).
  • Writing that non-overlapping error bars show "the results are significant". The mark needs the difference between the means to be significant (not due to chance).
  • Saying a smaller standard deviation means the data are "more accurate". It means less variation / the readings are more consistent.
  • Confusing the molecules: writing "amino acid base sequence", or treating DNA as if it were made of amino acids.
  • Comparing amino acid sequences but not linking the similarity back to the DNA base sequence or to a relationship. "The amino acids are similar" on its own does not score.
  • Concluding organisms are the same species or closely related because sequences are "similar" when the mark requires "identical".
  • Forgetting that comparing amino acid sequences underestimates genetic variation because the code is degenerate (different DNA triplets can code for the same amino acid).
  • Saying cytochrome c is "present in all species". It is present in all organisms that respire (all eukaryotes), which is why it is useful.
  • Being asked for the ways to compare genetic diversity and naming only one. There are three molecular methods: DNA base sequence, mRNA base sequence and amino acid sequence.
  • Quoting a mean to too many decimal places instead of an appropriate number of significant figures.

Key definitions

  • Random sampling: selecting sample points (or individuals) using randomly generated coordinates or numbers, so that every part of the area (or every member of the population) has an equal chance of being chosen and bias is removed.
  • Representative sample: a sample that reflects the whole population, so that valid conclusions can be drawn from it.
  • Mean: the average of the data, calculated as the sum of the values divided by the number of values.
  • Standard deviation: a measure of the spread (dispersion) of the data around the mean.
  • Degenerate (genetic code): more than one DNA base triplet codes for the same amino acid.

Specification

  • I can state that genetic diversity within or between species can be compared using the frequency of measurable or observable characteristics, the base sequence of DNA, the base sequence of mRNA, and the amino acid sequence of proteins.
  • I can interpret data on similarities and differences in DNA base sequences and amino acid sequences to suggest relationships within a species and between species.
  • I can explain that gene technology has replaced inferring differences from observable characteristics with direct investigation of DNA sequences.
  • I can describe how to collect data from random samples (design and carry out random sampling).
  • I can calculate a mean value from collected data.
  • I can interpret mean values and their standard deviations, including what the overlap of standard deviations shows. (You will not be asked to calculate a standard deviation in a written paper.)

Ready to test yourself?

Put Investigating diversity into practice with exam-style questions and full mark schemes.

Practise Investigating diversity