A comprehensive guide to the types of genetic variation in the human genome — from single base changes to chromosomal rearrangements — and their discovery through GWAS.
Any two randomly chosen humans differ at approximately 0.1% of their genomic sequence — roughly 3 million positions in the 3.2 billion base pair genome. While this seems small, it generates enormous phenotypic diversity and disease susceptibility differences. Understanding the types and scale of genetic variation is fundamental to human genetics, disease biology, and precision medicine.
SNVs per person vs reference
CNV events per genome
Known SNPs in dbSNP
De novo mutations per birth
Genetic variants are classified by their size and mechanism. Most common is the single nucleotide variant (SNV); when found at a frequency above 1% in a population, it is called a single nucleotide polymorphism (SNP). Larger variants include copy number variants, structural variants, and chromosomal abnormalities.
The human genome contains approximately 4–7 million SNPs per individual relative to the reference genome. The NCBI dbSNP database catalogs over 700 billion known human variants. The gnomAD database (v4.0) contains over 800 million variants observed in ~730,000 sequenced individuals.
| Location | Type | Functional Consequence |
|---|---|---|
| Coding exon — synonymous | Silent mutation | Same amino acid (wobble position); may still affect splicing or mRNA stability |
| Coding exon — missense | Amino acid change | Altered protein — may be benign, damaging, or pathogenic (e.g., BRCA1 p.Ser1040Asn) |
| Coding exon — nonsense | Premature stop codon | Truncated protein; often pathogenic (NMD may degrade mRNA) |
| Splice site | Affects GT-AG dinucleotides | Exon skipping, intron retention, cryptic splice site activation |
| Promoter / regulatory | eQTL (expression QTL) | Altered transcription factor binding; changes gene expression level |
| Intronic | Deep intronic | Often neutral; some create cryptic splice sites or affect regulatory elements |
| 3' UTR | Affects mRNA stability / miRNA binding | Altered mRNA half-life, translation efficiency, or miRNA-mediated silencing |
Nearby SNPs are often inherited together as haplotype blocks due to linkage disequilibrium (LD) — the non-random association of alleles at different loci. LD arises because recombination between close markers is rare. A single tag SNP within an LD block can represent all variants in that region, enabling efficient genotyping for association studies. The HapMap project (2005) and 1000 Genomes project (2015) systematically characterized LD structure and haplotypes across human populations.
SNP allele frequencies vary among human populations as a consequence of demographic history (population bottlenecks, founder effects, migration) and natural selection. African populations have the highest genetic diversity (more SNPs, lower LD blocks) reflecting the origin of Homo sapiens in Africa. Non-African populations show lower diversity reflecting the out-of-Africa bottleneck ~60,000 years ago. These population differences affect polygenic risk score transferability between ancestral groups — a major equity challenge in genomic medicine.
Each person carries approximately 1,000 CNV regions that differ from the reference genome. CNVs collectively account for more base pair differences between individuals than all SNPs combined, yet affect far fewer genomic positions. Landmark studies by Iafrate et al. and Sebat et al. (2004) discovered that large structural variation is pervasive in the human genome — contrary to the prevailing view at the time that the genome was nearly identical in structure between individuals.
Some CNV regions are highly associated with disease:
When indels in coding regions are not a multiple of 3, they shift the reading frame of translation — changing all subsequent amino acids and usually generating a premature stop codon. Frameshift mutations are typically loss-of-function. Examples:
Indels that are multiples of 3 bp add or remove whole amino acids without shifting the reading frame. These can be pathogenic or neutral depending on the affected protein domain. Example: a 3 bp deletion removing a critical residue in an enzyme's active site.
| Gene / Disease | Mutation | Consequence |
|---|---|---|
| BRCA1 / Hereditary Breast Cancer | 185delAG (c.68_69del) | Frameshift → premature stop; ~1% of Ashkenazi Jewish women |
| CFTR / Cystic Fibrosis | ΔF508 (c.1521_1523del) | In-frame 3 bp deletion; removes phenylalanine at position 508; protein misfolding |
| HTT / Huntington Disease | CAG trinucleotide repeat expansion | Polyglutamine expansion >35 repeats → toxic protein aggregation |
| BRCA2 / Hereditary Cancer | 6174delT | Frameshift; ~1.4% of Ashkenazi Jewish women |
| FMR1 / Fragile X Syndrome | CGG repeat expansion | >200 repeats → gene silencing → intellectual disability |
Microsatellites (short tandem repeats, STRs) are particularly prone to indel mutations due to slippage during DNA replication. Deficiency in DNA mismatch repair (MMR) proteins (MLH1, MSH2, MSH6, PMS2) leads to microsatellite instability (MSI) — widespread accumulation of indels at STR loci throughout the genome. MSI-high cancers (colorectal, endometrial, gastric) have high mutation burdens and are particularly responsive to immune checkpoint inhibitor immunotherapy.
Large-scale structural changes affect the organization of chromosomal segments and are a major class of human genetic variation, particularly in cancer and developmental disorders.
Transfer of chromosomal segments between non-homologous chromosomes. Reciprocal translocations exchange segments between two chromosomes; Robertsonian translocations fuse two acrocentric chromosomes (13, 14, 15, 21, 22) at their centromeres.
Example: t(9;22)(q34;q11) — Philadelphia chromosome in CML; fuses BCR and ABL1 genes creating BCR-ABL1 fusion oncogene (target of imatinib)
A chromosomal segment is reversed end-to-end. Pericentric inversions include the centromere; paracentric inversions do not. Inversions can disrupt genes at breakpoints or bring together gene regulatory elements with inappropriate targets.
Example: inv(16)(p13q22) in acute myeloid leukemia (AML-M4) — creates CBFbeta-MYH11 fusion
Loss of a chromosomal segment. Interstitial deletions remove internal segments; terminal deletions remove chromosome ends. Large deletions visible by karyotype; small deletions (microdeletions) require FISH or microarray.
Example: del(5q) in MDS; del(17p) (TP53 deletion) in CLL — poor prognosis
Extra copy of a chromosomal region. Tandem duplications are adjacent; interchromosomal duplications occur on different chromosomes. Can create gene dosage effects or novel fusion genes.
Example: dup(17p12) — Charcot-Marie-Tooth neuropathy type 1A; MYC amplification in various cancers
Chromothripsis ("chromosome shattering") is a catastrophic event in which a chromosome undergoes dozens to hundreds of simultaneous double-strand breaks and is reassembled in a disorganized fashion through aberrant DNA repair. First described in 2011, chromothripsis is observed in ~2–3% of cancers (more common in bone cancers) and some congenital developmental disorders. Similarly, chromoplexy involves coordinated rearrangements of multiple chromosomes in a single event — particularly common in prostate cancer.
Genome-wide association studies (GWAS) systematically scan the entire genome to identify common genetic variants (typically SNPs) associated with complex traits and diseases. GWAS transformed human genetics by enabling unbiased discovery of disease-associated loci without prior biological hypotheses.
The NHGRI-EBI GWAS Catalog contains over 500,000 significant associations across ~4,000 traits as of 2026. Key examples:
| Disease/Trait | Loci Discovered | Notable Genes / Insights |
|---|---|---|
| Type 2 Diabetes | >700 | TCF7L2 (strongest signal), PPARG, KCNJ11; revealed beta-cell biology importance |
| Coronary Artery Disease | >300 | 9p21.3 (non-coding, ANRIL lncRNA); many novel loci beyond traditional lipid genes |
| Height (adult stature) | >5,000 | Thousands of loci, each tiny effect; illuminates skeletal and growth biology |
| Schizophrenia | >300 | MHC region (strongest); complement system (C4A); synaptic genes |
| Inflammatory Bowel Disease | >200 | IL23R, NOD2, ATG16L1 (autophagy); ileal vs. colonic Crohn's differ genetically |
| Breast Cancer | >200 | FGFR2, LSP1, MAP3K1; many explain only small fraction of heritability |
Polygenic risk scores aggregate the effects of hundreds to millions of common variants into a single score predicting individual disease risk. PRS for coronary artery disease, breast cancer, and diabetes can identify individuals at 3–5× higher lifetime risk — comparable in risk magnitude to some monogenic risk factors. PRS are beginning to enter clinical practice for coronary artery disease screening in young adults.
| Database | URL | Content |
|---|---|---|
| dbSNP | ncbi.nlm.nih.gov/snp | Comprehensive repository of all known human variants; >700B variants |
| gnomAD | gnomad.broadinstitute.org | Population allele frequencies from ~730K WGS/WES; essential for variant pathogenicity assessment |
| ClinVar | ncbi.nlm.nih.gov/clinvar | Clinical significance classifications (P, LP, VUS, LB, B) for disease-associated variants |
| OMIM | omim.org | Gene-disease associations; mutation catalogs for Mendelian disorders |
| GWAS Catalog | ebi.ac.uk/gwas | Curated repository of all published GWAS associations (NHGRI-EBI) |
| ClinGen | clinicalgenome.org | Expert curated gene-disease validity and variant pathogenicity classifications |
| COSMIC | cancer.sanger.ac.uk/cosmic | Somatic mutations in human cancers; driver gene catalog |
| UniProt Variants | uniprot.org | Protein-level variant annotations with functional consequences |
Subscribe for updates on GWAS discoveries, variant databases, and clinical genetics advances.