A comprehensive overview of the human genome — from its fundamental structure to the era of personal genomics and precision medicine.
Base pairs (haploid)
Protein-coding genes
Chromosomes (diploid)
Protein-coding DNA
The human genome is the complete set of genetic information encoded in the DNA of Homo sapiens. It contains approximately 3.2 billion (3.2 × 10⁹) base pairs of DNA in the haploid state (one copy of each chromosome). In the diploid state of most somatic cells, the genome contains approximately 6.4 billion base pairs organized into 46 chromosomes (22 pairs of autosomes + 1 pair of sex chromosomes).
The Human Genome Project (HGP), completed in April 2003, produced the first reference human genome sequence — a monumental collaborative achievement involving laboratories in the United States, United Kingdom, France, Germany, Japan, and China. The HGP cost approximately $2.7 billion and took 13 years. Today, a human genome can be sequenced for approximately $200–500 and the raw data produced in about 24 hours using next-generation sequencing technology — a price reduction of over 10 million-fold in 20 years.
Paradoxically, the human genome is neither the largest nor the most gene-dense genome known. The Paris japonica flower has a genome 50× larger than humans (~150 billion base pairs). The relationship between organism complexity and genome size is called the "C-value paradox" — there is no simple correlation. What distinguishes the human genome is not its size but the extraordinary complexity of its gene regulation, splicing diversity, and developmental gene expression programs.
If the DNA in a single human cell were stretched end-to-end, it would be approximately 2 meters long. Yet it is coiled and compacted into a nucleus approximately 6 micrometers in diameter — a compaction ratio of roughly 300,000:1. This extraordinary packaging is achieved through multiple levels of DNA compaction: nucleosome formation (DNA wraps around histone octamers), 30-nm chromatin fiber formation, looping, and chromosome condensation.
The human genome is organized into 23 pairs of chromosomes in somatic cells — 22 pairs of autosomes (numbered 1–22, roughly in order of decreasing size) and 1 pair of sex chromosomes (XX in females, XY in males). The total diploid chromosome number is 46. Each chromosome consists of a single, continuous, linear DNA molecule associated with proteins (mainly histones).
248.9 Mb
~2,058 genes
242.2 Mb
~1,309 genes
198.3 Mb
~1,078 genes
190.2 Mb
~796 genes
181.5 Mb
~923 genes
170.8 Mb
~1,049 genes
159.3 Mb
~989 genes
145.1 Mb
~741 genes
138.4 Mb
~847 genes
133.8 Mb
~816 genes
135.1 Mb
~1,317 genes
133.3 Mb
~1,074 genes
114.4 Mb
~327 genes
107.0 Mb
~860 genes
101.9 Mb
~613 genes
90.3 Mb
~954 genes
83.3 Mb
~1,228 genes
80.4 Mb
~291 genes
58.6 Mb
~1,476 genes
64.4 Mb
~564 genes
46.7 Mb
~265 genes
50.8 Mb
~515 genes
156.0 Mb
~855 genes
57.2 Mb
~71 genes
A karyotype is an ordered image of an individual's full complement of chromosomes, typically produced from cultured white blood cells stimulated to divide, arrested in metaphase (when chromosomes are most condensed), and stained with Giemsa or similar dyes to produce characteristic banding patterns (G-bands). Karyotyping allows identification of chromosomal abnormalities including:
Despite constituting approximately 1.5% of the total genome, protein-coding genes are the primary functional units that have traditionally defined genomics. Current estimates from GENCODE (the reference gene annotation project) identify approximately 19,000–20,000 protein-coding genes in the human genome — remarkably similar to simpler organisms like the nematode worm C. elegans (~20,000 genes) and far fewer than many plant species. The C. elegans comparison challenged the assumption that complexity correlates with gene number.
Human protein-coding genes vary enormously in size:
The complexity of the human proteome far exceeds the gene number through alternative splicing — a single gene can produce multiple distinct proteins by including or excluding different exons. The total human proteome is estimated to include over 100,000 distinct protein isoforms. Post-translational modifications (phosphorylation, ubiquitination, glycosylation, etc.) add another layer of functional diversity.
One of the most remarkable features of eukaryotic genes is that the protein-coding information is split into discontinuous segments called exons, separated by non-coding intervening sequences called introns.
A typical human gene contains:
After transcription, introns are removed from the pre-mRNA by the spliceosome — a large ribonucleoprotein complex containing 5 snRNAs (U1, U2, U4, U5, U6) and over 150 proteins. Splicing recognizes conserved splice site sequences (5' splice site: GU; 3' splice site: AG — the "GT-AG rule") and branch point sequences within the intron. The intron is excised as a lariat structure and degraded.
Alternative splicing allows different exon combinations to be included in the mature mRNA, generating multiple protein isoforms from a single gene. Approximately 95% of multi-exon human genes undergo alternative splicing. Types include: exon skipping (most common, ~39%), alternative 5' or 3' splice site selection, intron retention, and mutually exclusive exon usage. Alternative splicing is regulated by splicing factors (SR proteins, hnRNPs) that bind to exonic/intronic splicing enhancers and silencers.
While protein-coding genes were historically the focus of genomics, it is now recognized that the vast majority of the genome is transcribed but most transcripts do not encode proteins. Non-coding RNAs (ncRNAs) include several major classes with distinct biogenesis and functions:
| Class | Size | Count (human) | Function |
|---|---|---|---|
| microRNA (miRNA) | ~22 nt | ~2,600 | Post-transcriptional gene silencing via RISC complex; bind 3' UTR of target mRNAs |
| Small interfering RNA (siRNA) | ~21 nt | Variable | RNAi pathway — typically derived from exogenous dsRNA or transposons |
| Long non-coding RNA (lncRNA) | >200 nt | ~17,000 | Diverse: chromatin remodeling, transcription regulation, splicing, X-inactivation (XIST) |
| PIWI-interacting RNA (piRNA) | 24–30 nt | Millions | Transposon silencing in germline cells via PIWI proteins |
| Small nucleolar RNA (snoRNA) | 60–300 nt | ~700 | Ribosomal RNA processing and modification in the nucleolus |
| Circular RNA (circRNA) | Variable | >15,000 | miRNA sponging, protein decoys; highly stable (no free ends to degrade) |
| tRNA, rRNA, snRNA | Variable | Hundreds | Translation machinery, ribosome structure, splicing |
The ENCODE project demonstrated that approximately 80% of the human genome shows biochemical activity (transcription, chromatin modification, protein binding) in at least one cell type — challenging the notion of "junk DNA." However, whether all of this activity represents functional elements remains actively debated.
Approximately 45–50% of the human genome consists of repetitive sequences — transposable elements (mobile genetic elements) and tandem repeats. Far from being uniformly "junk," repetitive elements have profoundly shaped genome evolution and some have been co-opted for regulatory functions.
The dramatic drop in sequencing costs has ushered in an era of personal genomics — where an individual's entire genome can be sequenced, analyzed, and interpreted at clinical scale. This revolution is transforming medicine, ancestry research, and our understanding of disease.
Genomic data is uniquely identifying — it cannot be de-identified and it reveals information about biological relatives who have not consented. Key ethical issues include: informed consent for secondary uses of genomic data, risk of genetic discrimination (GINA in the US provides some workplace and insurance protection), data security and commercial exploitation of genomic databases, return of incidental findings (variants not related to the indication for testing), and equity in access to genomic medicine.
Subscribe for updates on genome research discoveries, new technologies, and clinical advances.