Audited ·Last updated 29 Jul 2026·8 citations·Tier 1·0 uses

Hardy-Weinberg Equilibrium Calculator

Free Hardy-Weinberg calculator — enter one frequency and get p, q, p², 2pq and q². Shows the five assumptions and expected carrier counts.

Hardy-Weinberg Equilibrium Calculator

What do you already know?
A dimensionless proportion strictly between 0 and 1 — not a percentage. One in ten thousand is 0.0001; 36 percent is 0.36. Exactly 0 and exactly 1 are rejected because the locus would be fixed and there would be no heterozygotes left to describe.
Head-count used only to turn the carrier frequency 2pq into an expected number of carriers. It does not affect p, q or any frequency — Hardy-Weinberg assumes an infinitely large population.
people
Heterozygote (carrier) frequency, 2pq
0.0198
The expected proportion of the population carrying one copy of each allele (genotype Aa), computed as 2pq. This is a population-level expectation under five idealised conditions — it is not any individual person's risk, and it is not a test result.
Dominant allele frequency (p)
0.99
Recessive allele frequency (q)
0.01
Homozygous dominant frequency (p²)
0.9801
Homozygous recessive frequency (q²)
0.0001
Dominant phenotype frequency (p² + 2pq)
0.9999
Carriers as a share of unaffected people
0.0198
Expected carriers in this population
990 people
Carrier odds
about 1 in 51

Background.

This Hardy-Weinberg calculator takes the one frequency you actually know and expands it into the whole picture: the two allele frequencies p and q, the three genotype frequencies p², 2pq and q², the frequency of each phenotype, and the expected number of carriers in a population of whatever size you name. Choose from the dropdown which quantity you are starting from — the frequency of the recessive phenotype, the frequency of the homozygous dominant genotype, either allele frequency on its own, or the frequency of the dominant phenotype — and the calculator reconstructs p and q and prints the full distribution.

The classic version of the problem looks like this. A recessive condition affects one person in ten thousand, so the frequency of the aa genotype is q² = 0.0001. Take the square root: q = 0.01. The other allele must make up the rest, so p = 1 − 0.01 = 0.99. Now expand: p² = 0.9801 of the population are homozygous dominant, 2pq = 2 × 0.99 × 0.01 = 0.0198 are heterozygotes, and q² = 0.0001 are affected. The three add to exactly 1.0000, which is the arithmetic check you should always run. In a town of 50,000 people the model expects 0.0198 × 50,000 = 990 carriers. Restated the way a textbook usually puts it: about 1 in 51 people carry the allele even though only 1 in 10,000 shows the trait. That gap between what you can see and what is actually in the gene pool is the single most useful thing the Hardy-Weinberg principle tells you, and it is why the principle survived a century of being derived, in 1908, twice and independently — by the Cambridge mathematician G. H. Hardy in a short letter to Science, and by the Stuttgart physician Wilhelm Weinberg, whose German-language paper was largely ignored by the English-speaking genetics community for roughly thirty-five years.

The result is only as good as five conditions, and they are strong ones. The population must be effectively infinite, so that random sampling of gametes does not drift allele frequencies from one generation to the next. Mating must be random with respect to this locus, so that gametes combine independently. There must be no natural selection, so that no genotype leaves more offspring than another. There must be no mutation converting one allele into the other. And there must be no migration adding or removing gene copies from outside. Violate any one of them and the proportions above are no longer what you should expect — which is precisely why the principle is useful as a null hypothesis. Population geneticists do not usually believe Hardy-Weinberg holds; they test whether it holds, and the size and direction of the departure tells them which of the five conditions has failed. A heterozygote deficit points to inbreeding or to two subpopulations having been pooled by mistake; in modern genotyping data it far more often points to a technical genotyping error at that marker.

Two scope limits are worth reading before you trust the number on screen. First, the heterozygote frequency this page returns is a **population expectation, not a person's risk**. Knowing that 1 in 51 people carries an allele tells you nothing about whether any named individual carries it; that question needs family history or an actual test, and the arithmetic for it lives on the carrier probability calculator. Second, this page covers a **single autosomal locus with exactly two alleles**. X-linked loci do not reach these proportions in one generation — they approach them asymptotically, oscillating between the sexes — and loci with three or more alleles need the full multinomial expansion rather than a binomial square.

Below the calculator you will find the derivation of p² + 2pq + q² = 1 from a Punnett square of gametes, each of the five modes worked through, a discussion of what a departure from equilibrium actually means and why the chi-square test for it uses one degree of freedom rather than two, the difference between the carrier frequency 2pq and the conditional carrier probability 2q/(1 + q), and the significant-figure discipline that stops a two-significant-figure disease incidence from being reported as a ten-digit allele frequency.

What is hardy-weinberg equilibrium calculator?

The Hardy-Weinberg principle states that in an idealised population the frequencies of alleles and genotypes stay constant from generation to generation, and that after a single round of random mating the three genotypes appear in the proportions p², 2pq and q². The National Research Council's forensic DNA report puts the rule in words rather than symbols: 'The proportion of persons with two copies of the same allele is the square of that allele's frequency, and the proportion of persons with two different alleles is twice the product of the two frequencies.' Here p is the frequency of one allele (conventionally the dominant or reference allele, written A) and q is the frequency of the other (the recessive or alternate allele, a). Because there are only two alleles at the locus, p + q = 1, and squaring both sides of that identity produces the genotype equation directly: (p + q)² = p² + 2pq + q² = 1. Every quantity on this page is a dimensionless proportion between 0 and 1, never a percentage and never a count — the only count is the expected number of carriers, which is just 2pq multiplied by the population size you supply. A worked second example makes the shape of the relationship clear. If a recessive phenotype is common rather than rare — say 4 percent of the population, q² = 0.04 — then q = 0.2 and p = 0.8. The genotype frequencies become p² = 0.64, 2pq = 0.32 and q² = 0.04, so 96 percent of the population shows the dominant phenotype and, of those unaffected people, 0.32 ÷ 0.96 = one third are carriers. In a population of 50,000 that is 16,000 heterozygotes. Compare this with the rare-allele case in the introduction and a pattern appears that trips people up constantly: as the recessive allele becomes rarer, the absolute number of carriers falls (from 0.32 down to 0.0198), but the share of all copies of that allele which are hidden inside carriers rather than exposed in affected individuals rises (from 0.8 up to 0.99). Both statements are true at once, and both fall straight out of the algebra: the hidden share is 2pq ÷ (2pq + 2q²), which simplifies to exactly p.

How to use this calculator.

  1. Choose from the dropdown which frequency you already have. If a problem gives you the number of affected individuals for a recessive trait, that is the frequency of the recessive phenotype (q²) — the most common starting point by far.
  2. Enter the frequency as a decimal proportion between 0 and 1, not as a percentage. One in ten thousand is 0.0001. Four percent is 0.04. Thirty-six percent is 0.36. Values of exactly 0 or exactly 1 are rejected, because at those points one allele has been lost and there are no heterozygotes left to describe.
  3. Set the population size if you want an expected head-count of carriers. It has no effect on any of the frequencies — Hardy-Weinberg assumes an infinite population, so the frequencies are size-independent by construction.
  4. Read p and q first, then check that p² + 2pq + q² comes to 1. If it does not, you have entered a percentage where a proportion was wanted.
  5. Read the heterozygote frequency 2pq as the headline answer to 'what fraction are carriers'. If your question is instead 'what fraction of the people who look normal are carriers', use the carrier-share-of-unaffected output, which divides 2pq by 1 − q².
  6. Do not report more digits than your input supports. An incidence quoted as 'about 1 in 10,000' justifies two significant figures in q at most, so q = 0.01 and 2pq ≈ 0.020 — not 0.0198019802.
  7. Before you use the result for anything real, confirm the five assumptions are plausible for your population. If the population is small, structured, inbred, under selection, or recently mixed, the expected proportions are the wrong benchmark and the departure is the interesting quantity.

The formula.

p + q = 1 → p² + 2pq + q² = 1

The derivation is a Punnett square drawn over gametes rather than over parents. In an infinite, randomly mating population, a sperm picked at random carries allele A with probability p and allele a with probability q; the same is independently true of an egg. Combining them gives four equally-weighted gamete pairings: A from father × A from mother, with probability p × p = p²; A × a, probability p × q; a × A, probability q × p; and a × a, probability q × q = q². The two heterozygous routes produce the same genotype Aa, so they are added, giving 2pq. Hence p² + 2pq + q² = 1, which is nothing more than the binomial expansion of (p + q)² = 1² = 1. The independence of the two gametes is exactly the assumption of random mating, and the infinite population size is what makes 'probability' and 'frequency' interchangeable.

The five modes on this page are five different ways of recovering p and q before that expansion is applied. If you know the frequency of the recessive phenotype, that quantity is q² and q = √(q²). If you know the frequency of the homozygous dominant genotype, that is p² and p = √(p²). If you know either allele frequency directly, the other is one minus it. If you know the frequency of the dominant phenotype, note that it equals p² + 2pq = 1 − q², so q = √(1 − dominant). In every case the calculator then computes p² , 2pq and q² from the recovered pair, which is why all five entry points produce an identical output table for the same underlying population — a round trip the test suite asserts explicitly.

ROUNDING STAGE — FINAL ONLY. Every intermediate value (the square root, the subtraction 1 − q, each product, and the division 2pq ÷ (p² + 2pq)) is carried in arbitrary-precision decimal arithmetic. Rounding happens exactly once, at the moment each number is returned, to ten decimal places. Nothing is rounded and then reused. One consequence is worth stating plainly: because the numbers land on a ten-decimal grid, a genotype frequency smaller than 5 × 10⁻¹¹ displays as exactly 0. The carrier-odds text is derived from the unrounded value, so at q = 10⁻¹² it still reads 'about 1 in 500,000,000,001' while the frequency column shows 0.

Two related quantities are easy to confuse, so this page prints both. The heterozygote frequency 2pq is the share of the WHOLE population that carries one copy. The carrier share of unaffected people is 2pq ÷ (p² + 2pq), which simplifies to 2q ÷ (1 + q) — the probability that someone who does not show the recessive phenotype is nevertheless a carrier. With q = 0.01 these are 0.0198 and 0.0198019802 respectively: almost identical for a rare allele, because removing the 0.0001 of affected people barely changes the denominator. With q = 0.2 they are 0.32 and 0.3333333333, a visible gap. Ogino and Wilson, writing for the genetic-counselling literature, use the unconditional 2(1 − q)q and note it is approximately 2q when q is small; at q = 0.01 that approximation gives 0.02 against the exact 0.0198, an overstatement of exactly q/p = 1.01 percent. The approximation always overstates, never understates.

When the model is used as a null hypothesis rather than as a predictor, the departure is tested with a chi-square goodness-of-fit statistic comparing observed genotype counts with the expected p²N, 2pqN and q²N. That test uses ONE degree of freedom, not two: there are three genotype classes, one is lost to the constraint that the counts sum to N, and a second is lost because p was estimated from the same data rather than specified in advance. Wigginton, Cutler and Abecasis showed that this chi-square approximation can have inflated type I error rates even in samples of a thousand individuals when the minor allele is uncommon, and recommended an exact test in that regime — a caveat that matters for genotyping quality control far more than for classroom problems.

A worked example.

Example

A recessive condition affects one person in ten thousand. What fraction of the population carries the allele without showing it, and how many carriers would you expect in a town of 50,000? Start from the only thing you can observe directly: for a fully recessive trait, the affected individuals are exactly the aa homozygotes, so q² = 0.0001. Take the square root to get the recessive allele frequency, q = 0.01. Since there are only two alleles at the locus, p = 1 − 0.01 = 0.99. Now expand the square. The homozygous dominant frequency is p² = 0.99 × 0.99 = 0.9801. The heterozygote frequency is 2pq = 2 × 0.99 × 0.01 = 0.0198. The homozygous recessive frequency is q² = 0.01 × 0.01 = 0.0001. Adding the three gives 0.9801 + 0.0198 + 0.0001 = 1.0000 exactly, which is the check that the arithmetic held. The headline answer is 2pq = 0.0198 — just under 2 percent of the population carries one copy. Stated as odds that is about 1 in 51. The dominant phenotype frequency is 0.9801 + 0.0198 = 0.9999, so 99.99 percent of people show no sign of the trait; of those unaffected people, 0.0198 ÷ 0.9999 = 0.0198019802 are carriers, which is the same thing as the closed form 2q ÷ (1 + q) = 0.02 ÷ 1.01. In a town of 50,000 the model expects 0.0198 × 50,000 = 990 carriers against just 5 affected individuals. That 990-to-5 ratio is the point of the exercise. Nearly 200 times as many people carry the allele as express it, and none of them can be identified by looking. Put the other way round, the fraction of all copies of the recessive allele that are sitting hidden in heterozygotes rather than exposed in affected people is 2pq ÷ (2pq + 2q²), which simplifies to exactly p = 0.99. Ninety-nine percent of the recessive alleles in this population are invisible to selection acting on the phenotype, which is the classical explanation for why selection against a rare recessive is so slow to remove it. Two caveats belong with the number rather than after it: 990 is an expected value for an idealised infinite population, so a real town of 50,000 will scatter around it, and the incidence 'one in ten thousand' carries two significant figures at best, so the honest report is 'roughly 2 percent, about a thousand people', not 'exactly 990'.

population Size50,000
value0
known Quantityrecessive-phenotype-frequency

Frequently asked questions.

What is the Hardy-Weinberg equation?
Two equations, used together. The first is p + q = 1: at a locus with exactly two alleles, the two allele frequencies must account for every gene copy in the population. The second is p² + 2pq + q² = 1, which is just the first one squared — it gives the expected genotype frequencies, with p² for the homozygous dominant genotype AA, 2pq for the heterozygote Aa, and q² for the homozygous recessive aa. The factor of 2 on the middle term exists because there are two ways to make a heterozygote: the A allele can come from either parent. Because these are frequencies rather than counts, they are dimensionless proportions between 0 and 1 and must sum to exactly 1 — that sum is the arithmetic check to run before trusting any answer.
What are the five assumptions of Hardy-Weinberg equilibrium?
An infinitely large population, so that random sampling of gametes does not drift the allele frequencies. Random mating with respect to the locus in question, so that gametes pair independently. No natural selection, so that no genotype has a survival or fertility advantage. No mutation, so that neither allele is being converted into the other. And no migration, so that no gene copies enter or leave the population. Saadat's 2024 editorial states the standard form of the list: the observed and expected genotype frequencies should agree 'in a large population with random mating, and in the absence of mutation, migration, and natural selection'. Violating any one of them invalidates the expected proportions — which is what makes the principle useful as a null hypothesis rather than as a description of any real population.
Does random mating have to hold for Hardy-Weinberg proportions to appear?
It is the standard sufficient condition, and it is what every course teaches, but it is not strictly necessary. Stark's 2023 paper in Hereditas argues explicitly that random mating is sufficient but not necessary, and exhibits non-random mating schemes — mating matrices satisfying a particular constraint — that also hold the p² : 2pq : q² proportions stable across generations. This does not change any number this calculator produces, and it does not change what you should write in an exam. It does change the precise wording: the five conditions are the standard conditions under which the proportions are expected and stable, not a list of logical necessities. We flag the dissent rather than pretend the textbook framing is unanimous.
Why do I take the square root of the affected frequency and not of something else?
Because for a completely recessive trait, the affected individuals are exactly the aa homozygotes and nobody else. Heterozygotes look identical to homozygous dominants, so the only genotype you can count by inspecting phenotypes is aa. Its expected frequency is q², so q = √(observed affected frequency). This is the single most useful move in the whole topic, and it is also the step where the assumptions bite hardest: if the population is inbred, structured, or under selection, the aa frequency is no longer q² and the square root gives you the wrong q. It also fails if the trait is not fully recessive — with incomplete dominance or codominance you can see heterozygotes directly, so you should count them instead of inferring them.
What is the difference between 2pq and the carrier probability for one person?
2pq is a property of a population; a carrier probability is a property of a person. 2pq says what proportion of everybody is heterozygous — with q = 0.01 that is 0.0198, about 1 in 51. It does not say that any named individual has a 1-in-51 chance of being a carrier, because that individual may have information the population average does not: an affected sibling, a known-carrier parent, a negative screening test, or ancestry from a population with a different allele frequency. This page also prints the closely related conditional quantity 2pq ÷ (p² + 2pq) = 2q ÷ (1 + q), the probability that someone who does NOT show the recessive phenotype is a carrier — 0.0198019802 at q = 0.01. For anything shaped like a personal or reproductive risk question, use the carrier probability calculator, which takes family history and test results into account. A Hardy-Weinberg frequency is not a clinical result and must not be used as one.
How many degrees of freedom does the chi-square test for Hardy-Weinberg equilibrium use?
One, for a locus with two alleles. There are three genotype classes, so you start with three. One degree of freedom is spent on the constraint that the observed counts sum to the sample size. A second is spent because the allele frequency p was estimated from the very data being tested rather than specified in advance — the general rule from the NIST/SEMATECH handbook is that the statistic has k − c degrees of freedom, where c is the number of estimated parameters plus one. Three minus two leaves one. This is the standard trap: testing against a fixed Mendelian ratio such as 3:1 uses k − 1 because nothing was estimated, but testing Hardy-Weinberg uses k − 2 because p was. Note also that Wigginton, Cutler and Abecasis showed the chi-square approximation can have inflated type I error rates even in samples of a thousand individuals when the minor allele is uncommon, and recommend an exact test in that situation.
Does the Hardy-Weinberg principle work for X-linked genes?
Not in the one-generation form this calculator implements. On the X chromosome, individuals with two X chromosomes have the usual three genotypes at frequencies p², 2pq and q², but individuals with a single X are hemizygous and have only two genotypes, at frequencies p and q. Because the two sexes carry different fractions of the population's X chromosomes, allele frequencies that start out unequal between the sexes do not equalise in a single generation — they oscillate, with the difference halving and changing sign each generation, and converge on the equilibrium only asymptotically. That is why an X-linked recessive condition appears far more often in the hemizygous sex: its frequency there is q rather than q². This page is scoped to a single autosomal locus with two alleles; use the sex-linked inheritance calculator for the X-linked case.
What does it mean when a population is NOT in Hardy-Weinberg equilibrium?
It means at least one of the five conditions has failed, and the direction of the departure narrows down which. An excess of homozygotes and a deficit of heterozygotes points to inbreeding, assortative mating, or — very commonly — the Wahlund effect, where two subpopulations with different allele frequencies have been pooled and analysed as one. An excess of heterozygotes points to outbreeding, heterozygote advantage, or a recent admixture event. In modern genotyping data, however, the single most frequent cause of a significant departure at one marker is not biology at all but a technical genotyping failure at that marker — allele dropout, a poorly performing probe, or a null allele — which is exactly why quality-control pipelines test every marker for Hardy-Weinberg and discard the outliers before any biological interpretation is attempted.
Can I use this calculator for a gene with three or more alleles?
No. The p² + 2pq + q² form is the binomial expansion of (p + q)², and it only covers two alleles. With three alleles at frequencies p, q and r summing to 1, the expected genotype frequencies come from the trinomial expansion (p + q + r)² = p² + q² + r² + 2pq + 2pr + 2qr — three homozygous classes and three heterozygous classes rather than two and one. The ABO blood group is the standard example, and its analysis is genuinely more involved because the A and B alleles are codominant with each other while both are dominant to O, so several genotypes share a phenotype. The general rule still holds — every homozygote is the square of its allele frequency, every heterozygote is twice the product of the two — but this calculator implements only the two-allele case and will give a wrong answer if you force a multi-allelic locus into it.
Who actually discovered the Hardy-Weinberg principle?
Both of them, independently, in 1908. G. H. Hardy, a Cambridge pure mathematician with no particular interest in biology, wrote a short letter to Science in response to a claim that a dominant allele must inevitably spread through a population, showing algebraically that in the absence of disturbing forces the proportions simply stay put. Wilhelm Weinberg, a physician in Stuttgart, published a fuller and in some respects more general treatment in German in the same year. As Crow documented in Genetics in 1999, Weinberg's contribution went largely unrecognised in the English-speaking literature for roughly thirty-five years — a language barrier, not a priority dispute. The convention of naming the principle after both men only became standard in the mid-twentieth century, and some older English textbooks still refer to it as 'Hardy's law' alone.

References& sources.

  1. [1]Hardy, G. H. (1908). Mendelian proportions in a mixed population. Science 28(706):49–50. The original half-page letter showing that under random mating the genotype proportions p², 2pq and q² are stable from one generation to the next. Freely readable as an archival page-image reprint in the Yale Journal of Biology and Medicine 76(2):79–80 (2003). Independent, free, archival scan; retrieved 2026-07-29.
  2. [2]National Research Council (US) Committee on DNA Forensic Science (1996). The Evaluation of Forensic DNA Evidence, Chapter 4, section 'Random Mating and Hardy-Weinberg Proportions'. National Academies Press; NCBI Bookshelf NBK232608. Source of the plain-language statement of the proportions used in the 'What is' section: the proportion of persons with two copies of the same allele is the square of that allele's frequency, and the proportion with two different alleles is twice the product of the two frequencies. Independent, free, US National Academies; retrieved 2026-07-29.
  3. [3]Ogino, S. & Wilson, R. B. (2004). Bayesian analysis and risk assessment in genetic counseling and testing. Journal of Molecular Diagnostics 6(1):1–9. PMC1867463. Source for deriving the carrier frequency 2(1 − q)q from the disease frequency q² under Hardy-Weinberg equilibrium, and for the stated small-q approximation 2q. Independent, free full text; retrieved 2026-07-29.
  4. [4]Saadat, M. (2024). The importance of examining the Hardy-Weinberg Equilibrium in genetic association studies. Molecular Biology Research Communications 13(1):1–2, doi:10.22099/mbrc.2023.48386.1872. PMC10644312. Source for the conventional five-condition statement — a large population with random mating, and in the absence of mutation, migration, and natural selection. Independent, free, peer-reviewed; retrieved 2026-07-29.
  5. [5]Stark, A. E. (2023). Stable populations and Hardy-Weinberg equilibrium. Hereditas 160:19. PMC10161561. Consulted as an independent second authority: it confirms the p² : 2pq : q² arithmetic but dissents from the conventional teaching, arguing that random mating is a sufficient but not a necessary condition and exhibiting non-random mating schemes that also hold the proportions. Independent, free; retrieved 2026-07-29.
  6. [6]Wigginton, J. E., Cutler, D. J. & Abecasis, G. R. (2005). A note on exact tests of Hardy-Weinberg equilibrium. American Journal of Human Genetics 76(5):887–893. PubMed 15789306. Source for the caveat that the chi-square goodness-of-fit test for Hardy-Weinberg can have inflated type I error rates even in samples of about 1,000 individuals containing roughly 100 copies of the minor allele, and that an exact test should be preferred there. Independent; abstract free on PubMed, full text paywalled at the publisher; retrieved 2026-07-29.
  7. [7]NIST/SEMATECH (2012). e-Handbook of Statistical Methods, §1.3.5.15 'Chi-Square Goodness-of-Fit Test'. Source for the general degrees-of-freedom rule k − c, where c is the number of estimated parameters plus one — the rule that gives 1 degree of freedom for a two-allele Hardy-Weinberg test and k − 1 for a fixed Mendelian ratio. Independent, free, US National Institute of Standards and Technology; retrieved 2026-07-29.
  8. [8]Crow, J. F. (1999). Hardy, Weinberg and language impediments. Genetics 152(3):821–825. PMC1460671. Cited only for the historical point that Wilhelm Weinberg's independent 1908 German-language derivation was largely unrecognised in the English-speaking literature for roughly thirty-five years. Independent, free, PDF-only full text; retrieved 2026-07-29.

In this category

Embed

Quanta Pro

Paid features are coming later.

  • All 682 calculators remain free
  • No billing is enabled
Coming soon