Sign in to save

Bookmark this page so you can find it later.

Sign in to save

Bookmark this page so you can find it later.

Medical Technology: Next-Generation Sequencing infographic - Reading Millions of DNA Fragments

Click image to open full size

Medical Technology

Medical Technology: Next-Generation Sequencing

Reading Millions of DNA Fragments

Next-generation sequencing, or NGS, is a medical technology that reads DNA much faster than older one-fragment-at-a-time methods. Instead of sequencing a single DNA molecule, an NGS instrument reads millions of small DNA fragments in parallel on a flow cell. This makes it possible to study whole genomes, cancer mutations, inherited disease genes, and infectious organisms in one test.

The result is a powerful bridge between biology, medicine, optics, chemistry, and data science.

In a typical NGS workflow, DNA is cut into fragments, adapters are attached, and the fragments bind to a patterned surface inside the flow cell. Each fragment is copied into a cluster, then fluorescent signals reveal which base is added during each sequencing cycle. Cameras and optical sensors capture the signals from many clusters at once, and software converts the images into base calls and sequence reads.

The final data are aligned to a reference genome or assembled to identify variants that may affect health.

Understanding Medical Technology: Next-Generation Sequencing

Before a sample reaches a sequencer, its quality matters greatly. DNA from blood is often fairly intact, while DNA from an old tissue sample may be broken or chemically damaged. Laboratory workers measure how much DNA is present and check its fragment sizes.

They then make a library, which is a prepared collection of DNA pieces carrying known tag sequences. These tags let the instrument hold, copy, and identify the pieces. Some tests target a small set of medically important genes.

Others examine all protein coding regions, called the exome, or nearly the entire genome. The choice depends on the medical problem, cost, and amount of usable DNA.

The instrument does not simply produce a perfect DNA message. It produces many short reads, each with a quality score that estimates confidence in every base call. Errors can come from faint light signals, repeated DNA regions, sample contamination, or mistakes made while copying fragments.

Repeated regions are especially difficult because several places in the genome can look nearly identical. Software may not know where a short read belongs.

Longer reads can help resolve these areas, though they may have different error patterns or cost more. Scientists filter weak reads and compare results across many overlapping reads before trusting a change in DNA.

Coverage is important because a single read is weak evidence. When many independent reads show the same base at one position, confidence rises. A coverage value of thirty times means that, on average, each genome position was measured by about thirty reads.

Average coverage does not guarantee equal coverage everywhere. Some regions receive far fewer reads because their DNA sequence is hard to copy or contains many G and C bases. A clinical report may therefore list areas that were not measured reliably.

In cancer testing, the fraction of reads carrying a variant can be small because a tumour sample contains a mixture of cancer cells and normal cells. Low fractions need careful checking.

After sequencing, bioinformatics turns raw reads into information that doctors can use. Programs remove adapter sequences, align reads, find possible variants, and label how likely each result is to be real. Finding a variant does not prove that it causes disease.

Many DNA differences are harmless and common in the population. Specialists compare a finding with databases, family history, symptoms, laboratory evidence, and published research. Important results are often confirmed with another method.

Students should separate three ideas clearly. A sequencing read is a measurement. A variant is a difference from a reference sequence.

A diagnosis requires evidence that connects that difference to a person’s health. This distinction explains why genomic medicine needs laboratory science, computing, and careful human judgement.

Key Facts

  • DNA bases are read as A, T, C, and G, where A pairs with T and C pairs with G.
  • NGS uses massively parallel sequencing to read millions to billions of DNA fragments in one run.
  • Coverage = total bases sequenced / genome size.
  • If 90 billion bases are sequenced for a 3 billion base genome, coverage = 90,000,000,000 / 3,000,000,000 = 30x.
  • Read length is the number of bases measured in one DNA fragment, such as 150 bases per read.
  • Variant allele fraction = variant reads / total reads at that DNA position.

Vocabulary

Next-generation sequencing
A set of high-throughput methods that determine DNA sequences by reading many DNA fragments at the same time.
Flow cell
A glass or silicon cartridge with tiny lanes or patterned sites where DNA fragments attach and are sequenced.
Library
A prepared collection of DNA fragments with adapters added so the fragments can bind, amplify, and be read by the sequencer.
Cluster
A group of many copied DNA molecules in one location on the flow cell that produces a strong enough signal to detect.
Base call
The computer-assigned identity of a DNA base, A, T, C, or G, based on the detected signal during sequencing.

Common Mistakes to Avoid

  • Confusing NGS with reading one long DNA molecule, because most NGS methods read many short fragments and then use computation to place them in order.
  • Ignoring coverage, because a single read at a position is not enough to confidently identify many variants or sequencing errors.
  • Assuming every detected variant causes disease, because many DNA variants are harmless and clinical interpretation depends on evidence, frequency, and gene function.
  • Forgetting the library preparation step, because the sequencer usually cannot read raw genomic DNA until it is fragmented and fitted with adapters.

Practice Questions

  1. 1 A sequencing run produces 120 billion bases of data for a human genome of 3 billion bases. What is the average coverage?
  2. 2 A DNA position has 80 total reads, and 20 reads show a mutation. What is the variant allele fraction?
  3. 3 A sequencing machine detects fluorescent signals from millions of clusters at once. Explain why this massively parallel design makes NGS faster than sequencing one DNA fragment at a time.