Human Genetic Variation

A comprehensive guide to the types of genetic variation in the human genome — from single base changes to chromosomal rearrangements — and their discovery through GWAS.

⚕️ Medical Disclaimer: The information on this page is for educational and informational purposes only. It is not intended to be a substitute for professional medical advice, diagnosis, or treatment. Always seek the advice of your physician or other qualified healthcare provider with any questions you may have regarding a medical condition.

Overview: The Scale of Human Genetic Variation

Any two randomly chosen humans differ at approximately 0.1% of their genomic sequence — roughly 3 million positions in the 3.2 billion base pair genome. While this seems small, it generates enormous phenotypic diversity and disease susceptibility differences. Understanding the types and scale of genetic variation is fundamental to human genetics, disease biology, and precision medicine.

~3M

SNVs per person vs reference

1,000

CNV events per genome

~700B

Known SNPs in dbSNP

~200

De novo mutations per birth

Genetic variants are classified by their size and mechanism. Most common is the single nucleotide variant (SNV); when found at a frequency above 1% in a population, it is called a single nucleotide polymorphism (SNP). Larger variants include copy number variants, structural variants, and chromosomal abnormalities.

Single Nucleotide Polymorphisms (SNPs)

Definition: A SNP (pronounced "snip") is a single base position in the genome where two or more nucleotide alternatives exist at a frequency ≥1% in the population. When the alternative allele frequency is <1%, it is called a rare variant or single nucleotide variant (SNV).
Scale of SNP Variation

The human genome contains approximately 4–7 million SNPs per individual relative to the reference genome. The NCBI dbSNP database catalogs over 700 billion known human variants. The gnomAD database (v4.0) contains over 800 million variants observed in ~730,000 sequenced individuals.

Types of SNPs by Genomic Location and Functional Effect
LocationTypeFunctional Consequence
Coding exon — synonymousSilent mutationSame amino acid (wobble position); may still affect splicing or mRNA stability
Coding exon — missenseAmino acid changeAltered protein — may be benign, damaging, or pathogenic (e.g., BRCA1 p.Ser1040Asn)
Coding exon — nonsensePremature stop codonTruncated protein; often pathogenic (NMD may degrade mRNA)
Splice siteAffects GT-AG dinucleotidesExon skipping, intron retention, cryptic splice site activation
Promoter / regulatoryeQTL (expression QTL)Altered transcription factor binding; changes gene expression level
IntronicDeep intronicOften neutral; some create cryptic splice sites or affect regulatory elements
3' UTRAffects mRNA stability / miRNA bindingAltered mRNA half-life, translation efficiency, or miRNA-mediated silencing
Linkage Disequilibrium and Haplotypes

Nearby SNPs are often inherited together as haplotype blocks due to linkage disequilibrium (LD) — the non-random association of alleles at different loci. LD arises because recombination between close markers is rare. A single tag SNP within an LD block can represent all variants in that region, enabling efficient genotyping for association studies. The HapMap project (2005) and 1000 Genomes project (2015) systematically characterized LD structure and haplotypes across human populations.

Population Differences

SNP allele frequencies vary among human populations as a consequence of demographic history (population bottlenecks, founder effects, migration) and natural selection. African populations have the highest genetic diversity (more SNPs, lower LD blocks) reflecting the origin of Homo sapiens in Africa. Non-African populations show lower diversity reflecting the out-of-Africa bottleneck ~60,000 years ago. These population differences affect polygenic risk score transferability between ancestral groups — a major equity challenge in genomic medicine.

Copy Number Variations (CNVs)

Definition: A CNV is a segment of DNA of at least 1 kb that is present in a variable number of copies compared to a reference genome. CNVs include duplications (extra copies), deletions (missing copies), and can affect a single exon to an entire chromosome arm. They represent the second most common type of human genetic variation after SNPs.
Scale and Discovery

Each person carries approximately 1,000 CNV regions that differ from the reference genome. CNVs collectively account for more base pair differences between individuals than all SNPs combined, yet affect far fewer genomic positions. Landmark studies by Iafrate et al. and Sebat et al. (2004) discovered that large structural variation is pervasive in the human genome — contrary to the prevailing view at the time that the genome was nearly identical in structure between individuals.

CNV Detection Methods
  • Chromosomal Microarrays (SNP arrays / CGH arrays): Standard clinical tool; detects CNVs ≥50–500 kb depending on probe density
  • Read Depth Analysis from WGS: Detects CNVs as regions of abnormal read coverage (duplications = elevated coverage; deletions = reduced coverage)
  • FISH (Fluorescence In Situ Hybridization): Detects specific known CNVs at individual loci using fluorescent probes
  • qPCR: Quantifies copy number at specific loci
CNVs and Disease

Some CNV regions are highly associated with disease:

  • 22q11.2 deletion syndrome (DiGeorge syndrome): ~3 Mb deletion; most common chromosomal microdeletion (~1/4,000 births); congenital heart defects, cleft palate, T cell immunodeficiency, intellectual disability, schizophrenia risk
  • 15q11-q13 duplications/deletions: Angelman syndrome (maternal deletion or paternal UPD) / Prader-Willi syndrome (paternal deletion or maternal UPD) — imprinting-dependent
  • SMN1 deletion: Spinal muscular atrophy (SMA) — deletion of SMN1 on chromosome 5q
  • Williams syndrome: ~1.8 Mb deletion at 7q11.23 including ELN (elastin)
  • Charcot-Marie-Tooth type 1A: 1.5 Mb duplication at 17p11.2 including PMP22
  • Autism spectrum disorder: Multiple rare CNVs (16p11.2, 15q13.3, 1q21.1) contribute to ASD risk

Insertions and Deletions (Indels)

Definition: Indels (insertions and deletions) are mutations where one or more nucleotides are inserted into or deleted from the genome, ranging from a single base to several kilobases. Short indels (1–50 bp) are the most common structural variants. Indels in coding regions often cause frameshift mutations.
Types and Functional Consequences
Frameshift Mutations

When indels in coding regions are not a multiple of 3, they shift the reading frame of translation — changing all subsequent amino acids and usually generating a premature stop codon. Frameshift mutations are typically loss-of-function. Examples:

Normal sequence: ATG CAT CGT AAA GCT TAA Normal protein: Met-His-Arg-Lys-Ala-Stop +1 insertion (A): ATG ACA TCG TAA AGC TTA A... Frameshift protein: Met-Thr-Ser-Stop (truncated after 4 amino acids) -2 deletion (CA): ATG CTG TAA AGC TTA A... Frameshift protein: Met-Leu-Stop (truncated after 3 amino acids)
In-Frame Indels

Indels that are multiples of 3 bp add or remove whole amino acids without shifting the reading frame. These can be pathogenic or neutral depending on the affected protein domain. Example: a 3 bp deletion removing a critical residue in an enzyme's active site.

Examples of Disease-Causing Indels
Gene / DiseaseMutationConsequence
BRCA1 / Hereditary Breast Cancer185delAG (c.68_69del)Frameshift → premature stop; ~1% of Ashkenazi Jewish women
CFTR / Cystic FibrosisΔF508 (c.1521_1523del)In-frame 3 bp deletion; removes phenylalanine at position 508; protein misfolding
HTT / Huntington DiseaseCAG trinucleotide repeat expansionPolyglutamine expansion >35 repeats → toxic protein aggregation
BRCA2 / Hereditary Cancer6174delTFrameshift; ~1.4% of Ashkenazi Jewish women
FMR1 / Fragile X SyndromeCGG repeat expansion>200 repeats → gene silencing → intellectual disability
Microsatellite Instability (MSI)

Microsatellites (short tandem repeats, STRs) are particularly prone to indel mutations due to slippage during DNA replication. Deficiency in DNA mismatch repair (MMR) proteins (MLH1, MSH2, MSH6, PMS2) leads to microsatellite instability (MSI) — widespread accumulation of indels at STR loci throughout the genome. MSI-high cancers (colorectal, endometrial, gastric) have high mutation burdens and are particularly responsive to immune checkpoint inhibitor immunotherapy.

Chromosomal Rearrangements

Large-scale structural changes affect the organization of chromosomal segments and are a major class of human genetic variation, particularly in cancer and developmental disorders.

Types of Chromosomal Rearrangements
Translocations

Transfer of chromosomal segments between non-homologous chromosomes. Reciprocal translocations exchange segments between two chromosomes; Robertsonian translocations fuse two acrocentric chromosomes (13, 14, 15, 21, 22) at their centromeres.

Example: t(9;22)(q34;q11) — Philadelphia chromosome in CML; fuses BCR and ABL1 genes creating BCR-ABL1 fusion oncogene (target of imatinib)

Inversions

A chromosomal segment is reversed end-to-end. Pericentric inversions include the centromere; paracentric inversions do not. Inversions can disrupt genes at breakpoints or bring together gene regulatory elements with inappropriate targets.

Example: inv(16)(p13q22) in acute myeloid leukemia (AML-M4) — creates CBFbeta-MYH11 fusion

Deletions

Loss of a chromosomal segment. Interstitial deletions remove internal segments; terminal deletions remove chromosome ends. Large deletions visible by karyotype; small deletions (microdeletions) require FISH or microarray.

Example: del(5q) in MDS; del(17p) (TP53 deletion) in CLL — poor prognosis

Duplications

Extra copy of a chromosomal region. Tandem duplications are adjacent; interchromosomal duplications occur on different chromosomes. Can create gene dosage effects or novel fusion genes.

Example: dup(17p12) — Charcot-Marie-Tooth neuropathy type 1A; MYC amplification in various cancers

Chromothripsis and Chromoplexy

Chromothripsis ("chromosome shattering") is a catastrophic event in which a chromosome undergoes dozens to hundreds of simultaneous double-strand breaks and is reassembled in a disorganized fashion through aberrant DNA repair. First described in 2011, chromothripsis is observed in ~2–3% of cancers (more common in bone cancers) and some congenital developmental disorders. Similarly, chromoplexy involves coordinated rearrangements of multiple chromosomes in a single event — particularly common in prostate cancer.

GWAS — Genome-Wide Association Studies

Genome-wide association studies (GWAS) systematically scan the entire genome to identify common genetic variants (typically SNPs) associated with complex traits and diseases. GWAS transformed human genetics by enabling unbiased discovery of disease-associated loci without prior biological hypotheses.

How GWAS Works
  1. Cohort assembly: Cases (individuals with disease) and controls (without disease) are recruited — typically thousands to hundreds of thousands of individuals
  2. Genotyping: DNA from each participant is genotyped for 500,000–5 million SNPs using SNP arrays (imputation expands coverage to ~10–30 million common variants)
  3. Association testing: For each SNP, a statistical test (logistic regression for case-control, linear regression for quantitative traits) tests whether the SNP allele frequency differs between cases and controls
  4. Multiple testing correction: The genome-wide significance threshold is P < 5 × 10⁻⁸ (Bonferroni correction for ~1 million independent tests)
  5. Replication: Nominally significant associations are replicated in independent cohorts
  6. Fine-mapping and functional annotation: Identify the likely causal variant(s) and mechanism
GWAS Discoveries

The NHGRI-EBI GWAS Catalog contains over 500,000 significant associations across ~4,000 traits as of 2026. Key examples:

Disease/TraitLoci DiscoveredNotable Genes / Insights
Type 2 Diabetes>700TCF7L2 (strongest signal), PPARG, KCNJ11; revealed beta-cell biology importance
Coronary Artery Disease>3009p21.3 (non-coding, ANRIL lncRNA); many novel loci beyond traditional lipid genes
Height (adult stature)>5,000Thousands of loci, each tiny effect; illuminates skeletal and growth biology
Schizophrenia>300MHC region (strongest); complement system (C4A); synaptic genes
Inflammatory Bowel Disease>200IL23R, NOD2, ATG16L1 (autophagy); ileal vs. colonic Crohn's differ genetically
Breast Cancer>200FGFR2, LSP1, MAP3K1; many explain only small fraction of heritability
Limitations of GWAS
  • Missing heritability: Even thousands of GWAS variants explain only a fraction of the genetic contribution to common diseases — rare variants, gene-gene interactions, and epigenetic effects are poorly captured
  • Population bias: ~80% of GWAS participants are of European ancestry; PRS and association results often do not transfer well to non-European populations
  • Correlation, not causation: GWAS identifies associated SNPs, not necessarily causal variants — fine-mapping and functional studies are required
  • Effect sizes are small: Most common variants have odds ratios of 1.05–1.3; not individually clinically actionable
Polygenic Risk Scores (PRS)

Polygenic risk scores aggregate the effects of hundreds to millions of common variants into a single score predicting individual disease risk. PRS for coronary artery disease, breast cancer, and diabetes can identify individuals at 3–5× higher lifetime risk — comparable in risk magnitude to some monogenic risk factors. PRS are beginning to enter clinical practice for coronary artery disease screening in young adults.

Major Human Genetic Variant Databases

DatabaseURLContent
dbSNPncbi.nlm.nih.gov/snpComprehensive repository of all known human variants; >700B variants
gnomADgnomad.broadinstitute.orgPopulation allele frequencies from ~730K WGS/WES; essential for variant pathogenicity assessment
ClinVarncbi.nlm.nih.gov/clinvarClinical significance classifications (P, LP, VUS, LB, B) for disease-associated variants
OMIMomim.orgGene-disease associations; mutation catalogs for Mendelian disorders
GWAS Catalogebi.ac.uk/gwasCurated repository of all published GWAS associations (NHGRI-EBI)
ClinGenclinicalgenome.orgExpert curated gene-disease validity and variant pathogenicity classifications
COSMICcancer.sanger.ac.uk/cosmicSomatic mutations in human cancers; driver gene catalog
UniProt Variantsuniprot.orgProtein-level variant annotations with functional consequences

Stay Updated on Human Genetics Research

Subscribe for updates on GWAS discoveries, variant databases, and clinical genetics advances.