Metagenomics

What Is Metagenomics?

Metagenomics is the branch of genomics that sequences and analyzes the collective genetic material recovered directly from an environmental or clinical sample, without first isolating and culturing the organisms it contains. The term was coined in 1998 by Jo Handelsman and colleagues to describe the study of soil microbial DNA as a single collective genome. Its central premise is that the great majority of microorganisms in soil, seawater, sediment, and the human body cannot be grown in laboratory culture, so any survey built on cultivation misses most of what is present. By extracting total DNA from a sample and sequencing it directly, metagenomics captures the uncultured majority along with the cultivable minority.

The discipline sits at the intersection of molecular biology, high-throughput sequencing, and computational biology. Its outputs are of two kinds: taxonomic profiles that describe which organisms are present and in what proportion, and functional profiles that describe which genes and metabolic pathways the community encodes.

Amplicon Surveys and Whole-Genome Shotgun Sequencing

Two experimental strategies dominate. Amplicon sequencing amplifies a conserved marker gene, usually the bacterial 16S ribosomal RNA gene or the fungal internal transcribed spacer, and reads that single locus across every organism in the sample. It is inexpensive and tolerant of host DNA contamination, but it resolves taxonomy only to genus level in most cases and says nothing directly about gene content. Whole-genome shotgun sequencing instead fragments all DNA in the sample and sequences it randomly, which yields species-level and strain-level resolution and captures viruses, fungi, and archaea alongside bacteria. Direct comparisons of 16S amplicon and shotgun approaches for gut microbial communities report substantially different sensitivity and different biases, so the choice of method shapes the conclusions.

Assembly, Binning, and Annotation

Shotgun reads arrive as a mixture from hundreds or thousands of genomes at widely different abundances, which makes assembly harder than for a single isolate. Metagenomic assemblers reconstruct contigs from overlapping reads, then binning algorithms group contigs into draft genomes using tetranucleotide composition, coverage depth across samples, and marker-gene completeness. The resulting draft genomes are called metagenome-assembled genomes, or MAGs, and quality is reported as estimated completeness and contamination. Production pipelines such as the DOE Joint Genome Institute metagenome workflow chain assembly, structural and functional annotation, and binning into a single reproducible process that runs across thousands of samples per year. Annotated results are deposited in comparative systems such as IMG/M, the integrated microbial genomes and microbiomes platform, which places new datasets alongside reference genomes from bacteria, archaea, eukaryotes, plasmids, and viruses.

Functional and Sequence-Based Discovery

Beyond cataloguing, metagenomics is used to find genes with useful biochemistry. Sequence-based discovery searches assembled data for homologs of known enzyme families, while function-based discovery clones environmental DNA fragments into a culturable host and screens the resulting library for activity, an approach that can recover enzymes with no recognizable similarity to anything in the databases. Large-scale mining of environmental data has expanded the known diversity of bacterial and archaeal lineages considerably, with JGI reporting that thousands of novel genomes were reconstructed from Earth's microbiomes without cultivation. Adding transcriptomic, proteomic, or metabolomic measurements to the same samples extends the analysis from what the community could do to what it is actually doing.

Applications

Metagenomics has applications in a wide range of disciplines, including:

  • Human health research, including gut, skin, and oral microbiome studies
  • Clinical diagnostics, where unbiased sequencing identifies pathogens that targeted assays miss
  • Environmental monitoring of soil, freshwater, marine, and wastewater communities
  • Bioprospecting for industrial enzymes, antibiotics, and other natural products
  • Agriculture and plant science, particularly rhizosphere and rumen community studies
  • Biodefense and public health surveillance from environmental samples
Loading…