Hardy-Weinberg Equilibrium Calculator
Free Hardy-Weinberg calculator — enter one frequency and get p, q, p², 2pq and q². Shows the five assumptions and expected carrier counts.
Hardy-Weinberg Equilibrium Calculator
Background.
This Hardy-Weinberg calculator takes the one frequency you actually know and expands it into the whole picture: the two allele frequencies p and q, the three genotype frequencies p², 2pq and q², the frequency of each phenotype, and the expected number of carriers in a population of whatever size you name. Choose from the dropdown which quantity you are starting from — the frequency of the recessive phenotype, the frequency of the homozygous dominant genotype, either allele frequency on its own, or the frequency of the dominant phenotype — and the calculator reconstructs p and q and prints the full distribution.
The classic version of the problem looks like this. A recessive condition affects one person in ten thousand, so the frequency of the aa genotype is q² = 0.0001. Take the square root: q = 0.01. The other allele must make up the rest, so p = 1 − 0.01 = 0.99. Now expand: p² = 0.9801 of the population are homozygous dominant, 2pq = 2 × 0.99 × 0.01 = 0.0198 are heterozygotes, and q² = 0.0001 are affected. The three add to exactly 1.0000, which is the arithmetic check you should always run. In a town of 50,000 people the model expects 0.0198 × 50,000 = 990 carriers. Restated the way a textbook usually puts it: about 1 in 51 people carry the allele even though only 1 in 10,000 shows the trait. That gap between what you can see and what is actually in the gene pool is the single most useful thing the Hardy-Weinberg principle tells you, and it is why the principle survived a century of being derived, in 1908, twice and independently — by the Cambridge mathematician G. H. Hardy in a short letter to Science, and by the Stuttgart physician Wilhelm Weinberg, whose German-language paper was largely ignored by the English-speaking genetics community for roughly thirty-five years.
The result is only as good as five conditions, and they are strong ones. The population must be effectively infinite, so that random sampling of gametes does not drift allele frequencies from one generation to the next. Mating must be random with respect to this locus, so that gametes combine independently. There must be no natural selection, so that no genotype leaves more offspring than another. There must be no mutation converting one allele into the other. And there must be no migration adding or removing gene copies from outside. Violate any one of them and the proportions above are no longer what you should expect — which is precisely why the principle is useful as a null hypothesis. Population geneticists do not usually believe Hardy-Weinberg holds; they test whether it holds, and the size and direction of the departure tells them which of the five conditions has failed. A heterozygote deficit points to inbreeding or to two subpopulations having been pooled by mistake; in modern genotyping data it far more often points to a technical genotyping error at that marker.
Two scope limits are worth reading before you trust the number on screen. First, the heterozygote frequency this page returns is a **population expectation, not a person's risk**. Knowing that 1 in 51 people carries an allele tells you nothing about whether any named individual carries it; that question needs family history or an actual test, and the arithmetic for it lives on the carrier probability calculator. Second, this page covers a **single autosomal locus with exactly two alleles**. X-linked loci do not reach these proportions in one generation — they approach them asymptotically, oscillating between the sexes — and loci with three or more alleles need the full multinomial expansion rather than a binomial square.
Below the calculator you will find the derivation of p² + 2pq + q² = 1 from a Punnett square of gametes, each of the five modes worked through, a discussion of what a departure from equilibrium actually means and why the chi-square test for it uses one degree of freedom rather than two, the difference between the carrier frequency 2pq and the conditional carrier probability 2q/(1 + q), and the significant-figure discipline that stops a two-significant-figure disease incidence from being reported as a ten-digit allele frequency.
What is hardy-weinberg equilibrium calculator?
The Hardy-Weinberg principle states that in an idealised population the frequencies of alleles and genotypes stay constant from generation to generation, and that after a single round of random mating the three genotypes appear in the proportions p², 2pq and q². The National Research Council's forensic DNA report puts the rule in words rather than symbols: 'The proportion of persons with two copies of the same allele is the square of that allele's frequency, and the proportion of persons with two different alleles is twice the product of the two frequencies.' Here p is the frequency of one allele (conventionally the dominant or reference allele, written A) and q is the frequency of the other (the recessive or alternate allele, a). Because there are only two alleles at the locus, p + q = 1, and squaring both sides of that identity produces the genotype equation directly: (p + q)² = p² + 2pq + q² = 1. Every quantity on this page is a dimensionless proportion between 0 and 1, never a percentage and never a count — the only count is the expected number of carriers, which is just 2pq multiplied by the population size you supply. A worked second example makes the shape of the relationship clear. If a recessive phenotype is common rather than rare — say 4 percent of the population, q² = 0.04 — then q = 0.2 and p = 0.8. The genotype frequencies become p² = 0.64, 2pq = 0.32 and q² = 0.04, so 96 percent of the population shows the dominant phenotype and, of those unaffected people, 0.32 ÷ 0.96 = one third are carriers. In a population of 50,000 that is 16,000 heterozygotes. Compare this with the rare-allele case in the introduction and a pattern appears that trips people up constantly: as the recessive allele becomes rarer, the absolute number of carriers falls (from 0.32 down to 0.0198), but the share of all copies of that allele which are hidden inside carriers rather than exposed in affected individuals rises (from 0.8 up to 0.99). Both statements are true at once, and both fall straight out of the algebra: the hidden share is 2pq ÷ (2pq + 2q²), which simplifies to exactly p.
How to use this calculator.
- Choose from the dropdown which frequency you already have. If a problem gives you the number of affected individuals for a recessive trait, that is the frequency of the recessive phenotype (q²) — the most common starting point by far.
- Enter the frequency as a decimal proportion between 0 and 1, not as a percentage. One in ten thousand is 0.0001. Four percent is 0.04. Thirty-six percent is 0.36. Values of exactly 0 or exactly 1 are rejected, because at those points one allele has been lost and there are no heterozygotes left to describe.
- Set the population size if you want an expected head-count of carriers. It has no effect on any of the frequencies — Hardy-Weinberg assumes an infinite population, so the frequencies are size-independent by construction.
- Read p and q first, then check that p² + 2pq + q² comes to 1. If it does not, you have entered a percentage where a proportion was wanted.
- Read the heterozygote frequency 2pq as the headline answer to 'what fraction are carriers'. If your question is instead 'what fraction of the people who look normal are carriers', use the carrier-share-of-unaffected output, which divides 2pq by 1 − q².
- Do not report more digits than your input supports. An incidence quoted as 'about 1 in 10,000' justifies two significant figures in q at most, so q = 0.01 and 2pq ≈ 0.020 — not 0.0198019802.
- Before you use the result for anything real, confirm the five assumptions are plausible for your population. If the population is small, structured, inbred, under selection, or recently mixed, the expected proportions are the wrong benchmark and the departure is the interesting quantity.
The formula.
The derivation is a Punnett square drawn over gametes rather than over parents. In an infinite, randomly mating population, a sperm picked at random carries allele A with probability p and allele a with probability q; the same is independently true of an egg. Combining them gives four equally-weighted gamete pairings: A from father × A from mother, with probability p × p = p²; A × a, probability p × q; a × A, probability q × p; and a × a, probability q × q = q². The two heterozygous routes produce the same genotype Aa, so they are added, giving 2pq. Hence p² + 2pq + q² = 1, which is nothing more than the binomial expansion of (p + q)² = 1² = 1. The independence of the two gametes is exactly the assumption of random mating, and the infinite population size is what makes 'probability' and 'frequency' interchangeable.
The five modes on this page are five different ways of recovering p and q before that expansion is applied. If you know the frequency of the recessive phenotype, that quantity is q² and q = √(q²). If you know the frequency of the homozygous dominant genotype, that is p² and p = √(p²). If you know either allele frequency directly, the other is one minus it. If you know the frequency of the dominant phenotype, note that it equals p² + 2pq = 1 − q², so q = √(1 − dominant). In every case the calculator then computes p² , 2pq and q² from the recovered pair, which is why all five entry points produce an identical output table for the same underlying population — a round trip the test suite asserts explicitly.
ROUNDING STAGE — FINAL ONLY. Every intermediate value (the square root, the subtraction 1 − q, each product, and the division 2pq ÷ (p² + 2pq)) is carried in arbitrary-precision decimal arithmetic. Rounding happens exactly once, at the moment each number is returned, to ten decimal places. Nothing is rounded and then reused. One consequence is worth stating plainly: because the numbers land on a ten-decimal grid, a genotype frequency smaller than 5 × 10⁻¹¹ displays as exactly 0. The carrier-odds text is derived from the unrounded value, so at q = 10⁻¹² it still reads 'about 1 in 500,000,000,001' while the frequency column shows 0.
Two related quantities are easy to confuse, so this page prints both. The heterozygote frequency 2pq is the share of the WHOLE population that carries one copy. The carrier share of unaffected people is 2pq ÷ (p² + 2pq), which simplifies to 2q ÷ (1 + q) — the probability that someone who does not show the recessive phenotype is nevertheless a carrier. With q = 0.01 these are 0.0198 and 0.0198019802 respectively: almost identical for a rare allele, because removing the 0.0001 of affected people barely changes the denominator. With q = 0.2 they are 0.32 and 0.3333333333, a visible gap. Ogino and Wilson, writing for the genetic-counselling literature, use the unconditional 2(1 − q)q and note it is approximately 2q when q is small; at q = 0.01 that approximation gives 0.02 against the exact 0.0198, an overstatement of exactly q/p = 1.01 percent. The approximation always overstates, never understates.
When the model is used as a null hypothesis rather than as a predictor, the departure is tested with a chi-square goodness-of-fit statistic comparing observed genotype counts with the expected p²N, 2pqN and q²N. That test uses ONE degree of freedom, not two: there are three genotype classes, one is lost to the constraint that the counts sum to N, and a second is lost because p was estimated from the same data rather than specified in advance. Wigginton, Cutler and Abecasis showed that this chi-square approximation can have inflated type I error rates even in samples of a thousand individuals when the minor allele is uncommon, and recommended an exact test in that regime — a caveat that matters for genotyping quality control far more than for classroom problems.
A worked example.
A recessive condition affects one person in ten thousand. What fraction of the population carries the allele without showing it, and how many carriers would you expect in a town of 50,000? Start from the only thing you can observe directly: for a fully recessive trait, the affected individuals are exactly the aa homozygotes, so q² = 0.0001. Take the square root to get the recessive allele frequency, q = 0.01. Since there are only two alleles at the locus, p = 1 − 0.01 = 0.99. Now expand the square. The homozygous dominant frequency is p² = 0.99 × 0.99 = 0.9801. The heterozygote frequency is 2pq = 2 × 0.99 × 0.01 = 0.0198. The homozygous recessive frequency is q² = 0.01 × 0.01 = 0.0001. Adding the three gives 0.9801 + 0.0198 + 0.0001 = 1.0000 exactly, which is the check that the arithmetic held. The headline answer is 2pq = 0.0198 — just under 2 percent of the population carries one copy. Stated as odds that is about 1 in 51. The dominant phenotype frequency is 0.9801 + 0.0198 = 0.9999, so 99.99 percent of people show no sign of the trait; of those unaffected people, 0.0198 ÷ 0.9999 = 0.0198019802 are carriers, which is the same thing as the closed form 2q ÷ (1 + q) = 0.02 ÷ 1.01. In a town of 50,000 the model expects 0.0198 × 50,000 = 990 carriers against just 5 affected individuals. That 990-to-5 ratio is the point of the exercise. Nearly 200 times as many people carry the allele as express it, and none of them can be identified by looking. Put the other way round, the fraction of all copies of the recessive allele that are sitting hidden in heterozygotes rather than exposed in affected people is 2pq ÷ (2pq + 2q²), which simplifies to exactly p = 0.99. Ninety-nine percent of the recessive alleles in this population are invisible to selection acting on the phenotype, which is the classical explanation for why selection against a rare recessive is so slow to remove it. Two caveats belong with the number rather than after it: 990 is an expected value for an idealised infinite population, so a real town of 50,000 will scatter around it, and the incidence 'one in ten thousand' carries two significant figures at best, so the honest report is 'roughly 2 percent, about a thousand people', not 'exactly 990'.
Frequently asked questions.
What is the Hardy-Weinberg equation?
What are the five assumptions of Hardy-Weinberg equilibrium?
Does random mating have to hold for Hardy-Weinberg proportions to appear?
Why do I take the square root of the affected frequency and not of something else?
What is the difference between 2pq and the carrier probability for one person?
How many degrees of freedom does the chi-square test for Hardy-Weinberg equilibrium use?
Does the Hardy-Weinberg principle work for X-linked genes?
What does it mean when a population is NOT in Hardy-Weinberg equilibrium?
Can I use this calculator for a gene with three or more alleles?
Who actually discovered the Hardy-Weinberg principle?
References& sources.
- [1]Hardy, G. H. (1908). Mendelian proportions in a mixed population. Science 28(706):49–50. The original half-page letter showing that under random mating the genotype proportions p², 2pq and q² are stable from one generation to the next. Freely readable as an archival page-image reprint in the Yale Journal of Biology and Medicine 76(2):79–80 (2003). Independent, free, archival scan; retrieved 2026-07-29.
- [2]National Research Council (US) Committee on DNA Forensic Science (1996). The Evaluation of Forensic DNA Evidence, Chapter 4, section 'Random Mating and Hardy-Weinberg Proportions'. National Academies Press; NCBI Bookshelf NBK232608. Source of the plain-language statement of the proportions used in the 'What is' section: the proportion of persons with two copies of the same allele is the square of that allele's frequency, and the proportion with two different alleles is twice the product of the two frequencies. Independent, free, US National Academies; retrieved 2026-07-29.
- [3]Ogino, S. & Wilson, R. B. (2004). Bayesian analysis and risk assessment in genetic counseling and testing. Journal of Molecular Diagnostics 6(1):1–9. PMC1867463. Source for deriving the carrier frequency 2(1 − q)q from the disease frequency q² under Hardy-Weinberg equilibrium, and for the stated small-q approximation 2q. Independent, free full text; retrieved 2026-07-29.
- [4]Saadat, M. (2024). The importance of examining the Hardy-Weinberg Equilibrium in genetic association studies. Molecular Biology Research Communications 13(1):1–2, doi:10.22099/mbrc.2023.48386.1872. PMC10644312. Source for the conventional five-condition statement — a large population with random mating, and in the absence of mutation, migration, and natural selection. Independent, free, peer-reviewed; retrieved 2026-07-29.
- [5]Stark, A. E. (2023). Stable populations and Hardy-Weinberg equilibrium. Hereditas 160:19. PMC10161561. Consulted as an independent second authority: it confirms the p² : 2pq : q² arithmetic but dissents from the conventional teaching, arguing that random mating is a sufficient but not a necessary condition and exhibiting non-random mating schemes that also hold the proportions. Independent, free; retrieved 2026-07-29.
- [6]Wigginton, J. E., Cutler, D. J. & Abecasis, G. R. (2005). A note on exact tests of Hardy-Weinberg equilibrium. American Journal of Human Genetics 76(5):887–893. PubMed 15789306. Source for the caveat that the chi-square goodness-of-fit test for Hardy-Weinberg can have inflated type I error rates even in samples of about 1,000 individuals containing roughly 100 copies of the minor allele, and that an exact test should be preferred there. Independent; abstract free on PubMed, full text paywalled at the publisher; retrieved 2026-07-29.
- [7]NIST/SEMATECH (2012). e-Handbook of Statistical Methods, §1.3.5.15 'Chi-Square Goodness-of-Fit Test'. Source for the general degrees-of-freedom rule k − c, where c is the number of estimated parameters plus one — the rule that gives 1 degree of freedom for a two-allele Hardy-Weinberg test and k − 1 for a fixed Mendelian ratio. Independent, free, US National Institute of Standards and Technology; retrieved 2026-07-29.
- [8]Crow, J. F. (1999). Hardy, Weinberg and language impediments. Genetics 152(3):821–825. PMC1460671. Cited only for the historical point that Wilhelm Weinberg's independent 1908 German-language derivation was largely unrecognised in the English-speaking literature for roughly thirty-five years. Independent, free, PDF-only full text; retrieved 2026-07-29.
In this category
Embed
Quanta Pro
Paid features are coming later.
- All 682 calculators remain free
- No billing is enabled