The Human Genome

A comprehensive overview of the human genome — from its fundamental structure to the era of personal genomics and precision medicine.

⚕️ Medical Disclaimer: The information on this page is for educational and informational purposes only. It is not intended to be a substitute for professional medical advice, diagnosis, or treatment. Always seek the advice of your physician or other qualified healthcare provider with any questions you may have regarding a medical condition.

Human Genome at a Glance

3.2B

Base pairs (haploid)

~20,000

Protein-coding genes

46

Chromosomes (diploid)

~1.5%

Protein-coding DNA

Overview of the Human Genome

The human genome is the complete set of genetic information encoded in the DNA of Homo sapiens. It contains approximately 3.2 billion (3.2 × 10⁹) base pairs of DNA in the haploid state (one copy of each chromosome). In the diploid state of most somatic cells, the genome contains approximately 6.4 billion base pairs organized into 46 chromosomes (22 pairs of autosomes + 1 pair of sex chromosomes).

The Human Genome Project (HGP), completed in April 2003, produced the first reference human genome sequence — a monumental collaborative achievement involving laboratories in the United States, United Kingdom, France, Germany, Japan, and China. The HGP cost approximately $2.7 billion and took 13 years. Today, a human genome can be sequenced for approximately $200–500 and the raw data produced in about 24 hours using next-generation sequencing technology — a price reduction of over 10 million-fold in 20 years.

The T2T Reference Genome: The original HGP left approximately 8% of the genome unsequenced — mainly highly repetitive regions (centromeres, telomeres, pericentromeric heterochromatin) too difficult to assemble with short-read technology. In 2022, the Telomere-to-Telomere (T2T) consortium published the first truly complete human genome sequence (T2T-CHM13), filling in the remaining gaps. This complete reference adds approximately 200 million base pairs to the previously published sequence and reveals new gene candidates and repeat structures.

Genome Size and Complexity

Paradoxically, the human genome is neither the largest nor the most gene-dense genome known. The Paris japonica flower has a genome 50× larger than humans (~150 billion base pairs). The relationship between organism complexity and genome size is called the "C-value paradox" — there is no simple correlation. What distinguishes the human genome is not its size but the extraordinary complexity of its gene regulation, splicing diversity, and developmental gene expression programs.

If the DNA in a single human cell were stretched end-to-end, it would be approximately 2 meters long. Yet it is coiled and compacted into a nucleus approximately 6 micrometers in diameter — a compaction ratio of roughly 300,000:1. This extraordinary packaging is achieved through multiple levels of DNA compaction: nucleosome formation (DNA wraps around histone octamers), 30-nm chromatin fiber formation, looping, and chromosome condensation.

Chromosomes and Karyotype

The human genome is organized into 23 pairs of chromosomes in somatic cells — 22 pairs of autosomes (numbered 1–22, roughly in order of decreasing size) and 1 pair of sex chromosomes (XX in females, XY in males). The total diploid chromosome number is 46. Each chromosome consists of a single, continuous, linear DNA molecule associated with proteins (mainly histones).

Human Chromosome Reference Data

Chr 1

248.9 Mb

~2,058 genes

Chr 2

242.2 Mb

~1,309 genes

Chr 3

198.3 Mb

~1,078 genes

Chr 4

190.2 Mb

~796 genes

Chr 5

181.5 Mb

~923 genes

Chr 6

170.8 Mb

~1,049 genes

Chr 7

159.3 Mb

~989 genes

Chr 8

145.1 Mb

~741 genes

Chr 9

138.4 Mb

~847 genes

Chr 10

133.8 Mb

~816 genes

Chr 11

135.1 Mb

~1,317 genes

Chr 12

133.3 Mb

~1,074 genes

Chr 13

114.4 Mb

~327 genes

Chr 14

107.0 Mb

~860 genes

Chr 15

101.9 Mb

~613 genes

Chr 16

90.3 Mb

~954 genes

Chr 17

83.3 Mb

~1,228 genes

Chr 18

80.4 Mb

~291 genes

Chr 19

58.6 Mb

~1,476 genes

Chr 20

64.4 Mb

~564 genes

Chr 21

46.7 Mb

~265 genes

Chr 22

50.8 Mb

~515 genes

Chr X

156.0 Mb

~855 genes

Chr Y

57.2 Mb

~71 genes

Karyotype Analysis

A karyotype is an ordered image of an individual's full complement of chromosomes, typically produced from cultured white blood cells stimulated to divide, arrested in metaphase (when chromosomes are most condensed), and stained with Giemsa or similar dyes to produce characteristic banding patterns (G-bands). Karyotyping allows identification of chromosomal abnormalities including:

  • Aneuploidy: Abnormal chromosome number — trisomy 21 (Down syndrome, +21), trisomy 18 (Edwards syndrome), trisomy 13 (Patau syndrome), monosomy X (Turner syndrome, 45,X), 47,XXY (Klinefelter syndrome)
  • Structural rearrangements: Deletions, duplications, inversions, translocations — e.g., the Philadelphia chromosome t(9;22) in CML
  • Mosaicism: Two or more cell populations with different karyotypes in one individual

Protein-Coding Genes

Despite constituting approximately 1.5% of the total genome, protein-coding genes are the primary functional units that have traditionally defined genomics. Current estimates from GENCODE (the reference gene annotation project) identify approximately 19,000–20,000 protein-coding genes in the human genome — remarkably similar to simpler organisms like the nematode worm C. elegans (~20,000 genes) and far fewer than many plant species. The C. elegans comparison challenged the assumption that complexity correlates with gene number.

Human protein-coding genes vary enormously in size:

  • Smallest protein-coding genes: Some encode proteins of only ~30–100 amino acids (e.g., histone genes)
  • Largest gene: DMD (dystrophin) — 2.4 Mb genomic span, 79 exons, ~14 kb mRNA
  • Most transcripts per gene: KCNQ5 has over 30 alternative transcripts
  • Gene density: Chromosome 19 has the highest gene density (~23 genes/Mb); chromosomes 13 and 18 are gene-poor

The complexity of the human proteome far exceeds the gene number through alternative splicing — a single gene can produce multiple distinct proteins by including or excluding different exons. The total human proteome is estimated to include over 100,000 distinct protein isoforms. Post-translational modifications (phosphorylation, ubiquitination, glycosylation, etc.) add another layer of functional diversity.

Introns, Exons, and Pre-mRNA Processing

One of the most remarkable features of eukaryotic genes is that the protein-coding information is split into discontinuous segments called exons, separated by non-coding intervening sequences called introns.

Gene Structure

A typical human gene contains:

  • Promoter and regulatory elements (5' of the transcription start site)
  • 5' UTR (untranslated region in the mRNA, between TSS and start codon AUG)
  • Exons — coding sequences that appear in the mature mRNA; average human exon is ~170 bp
  • Introns — intervening sequences removed during pre-mRNA splicing; average human intron is ~3,000 bp (can exceed 100 kb)
  • 3' UTR (after stop codon — contains signals for polyadenylation and mRNA stability)
  • Poly-A tail — added post-transcriptionally to the 3' end
Pre-mRNA Splicing

After transcription, introns are removed from the pre-mRNA by the spliceosome — a large ribonucleoprotein complex containing 5 snRNAs (U1, U2, U4, U5, U6) and over 150 proteins. Splicing recognizes conserved splice site sequences (5' splice site: GU; 3' splice site: AG — the "GT-AG rule") and branch point sequences within the intron. The intron is excised as a lariat structure and degraded.

Alternative Splicing

Alternative splicing allows different exon combinations to be included in the mature mRNA, generating multiple protein isoforms from a single gene. Approximately 95% of multi-exon human genes undergo alternative splicing. Types include: exon skipping (most common, ~39%), alternative 5' or 3' splice site selection, intron retention, and mutually exclusive exon usage. Alternative splicing is regulated by splicing factors (SR proteins, hnRNPs) that bind to exonic/intronic splicing enhancers and silencers.

Non-Coding RNA (ncRNA)

While protein-coding genes were historically the focus of genomics, it is now recognized that the vast majority of the genome is transcribed but most transcripts do not encode proteins. Non-coding RNAs (ncRNAs) include several major classes with distinct biogenesis and functions:

ClassSizeCount (human)Function
microRNA (miRNA)~22 nt~2,600Post-transcriptional gene silencing via RISC complex; bind 3' UTR of target mRNAs
Small interfering RNA (siRNA)~21 ntVariableRNAi pathway — typically derived from exogenous dsRNA or transposons
Long non-coding RNA (lncRNA)>200 nt~17,000Diverse: chromatin remodeling, transcription regulation, splicing, X-inactivation (XIST)
PIWI-interacting RNA (piRNA)24–30 ntMillionsTransposon silencing in germline cells via PIWI proteins
Small nucleolar RNA (snoRNA)60–300 nt~700Ribosomal RNA processing and modification in the nucleolus
Circular RNA (circRNA)Variable>15,000miRNA sponging, protein decoys; highly stable (no free ends to degrade)
tRNA, rRNA, snRNAVariableHundredsTranslation machinery, ribosome structure, splicing

The ENCODE project demonstrated that approximately 80% of the human genome shows biochemical activity (transcription, chromatin modification, protein binding) in at least one cell type — challenging the notion of "junk DNA." However, whether all of this activity represents functional elements remains actively debated.

Repetitive Elements

Approximately 45–50% of the human genome consists of repetitive sequences — transposable elements (mobile genetic elements) and tandem repeats. Far from being uniformly "junk," repetitive elements have profoundly shaped genome evolution and some have been co-opted for regulatory functions.

  • LINEs (Long Interspersed Nuclear Elements): ~20% of genome; L1 elements (6 kb) can transpose via RNA intermediate; ~100 L1 elements remain potentially active
  • SINEs (Short Interspersed Nuclear Elements): ~13%; Alu elements (~300 bp) are the most abundant — ~1.1 million copies; B2 and MIRs in other mammals
  • LTR retrotransposons (HERVs): ~8%; Human Endogenous Retroviruses — ancient retroviral integrations; some HERV-derived proteins have been co-opted (syncytins for placental development)
  • DNA transposons: ~3%; currently mostly inactive in humans
  • Tandem repeats: Satellite DNA (constitutive heterochromatin, centromeres), microsatellites (STRs, used in forensic DNA profiling), and minisatellites

The Personal Genomics Era

The dramatic drop in sequencing costs has ushered in an era of personal genomics — where an individual's entire genome can be sequenced, analyzed, and interpreted at clinical scale. This revolution is transforming medicine, ancestry research, and our understanding of disease.

Sequencing Technologies
  • Whole Genome Sequencing (WGS): Sequences all ~3.2 billion base pairs; detects SNPs, indels, structural variants, CNVs; ~$200–500 for 30× coverage (2026)
  • Whole Exome Sequencing (WES): Sequences only the coding exons (~1.5% of genome); ~$200–300; misses non-coding regulatory variants
  • SNP Array / Genotyping Chips: Tests 500,000–4 million known variant positions; cost ~$50–200; used in consumer genomics (23andMe, AncestryDNA)
  • Long-read sequencing (PacBio, Oxford Nanopore): Reads of 10–100+ kb; resolves complex structural variants, haplotype phasing, repetitive regions
Clinical Applications
  • Rare disease diagnosis: WGS/WES solves ~35–40% of previously undiagnosed rare disease cases
  • Cancer genomics: Tumor sequencing identifies actionable mutations (EGFR, KRAS, BRCA1/2, PIK3CA) guiding targeted therapy selection
  • Pharmacogenomics: CYP2D6, CYP2C19, TPMT, DPYD variants predict drug metabolism and toxicity risk
  • Newborn screening: Genomic sequencing replacing biochemical newborn screening panels — identifies hundreds of conditions
  • Polygenic risk scores: Aggregate thousands of common variants to predict individual risk for diabetes, coronary artery disease, breast cancer, and other common diseases
Privacy and Ethical Considerations

Genomic data is uniquely identifying — it cannot be de-identified and it reveals information about biological relatives who have not consented. Key ethical issues include: informed consent for secondary uses of genomic data, risk of genetic discrimination (GINA in the US provides some workplace and insurance protection), data security and commercial exploitation of genomic databases, return of incidental findings (variants not related to the indication for testing), and equity in access to genomic medicine.

Stay Updated on Human Genetics Research

Subscribe for updates on genome research discoveries, new technologies, and clinical advances.