Meridia Insight Science Breakthroughs Knowledge

The Scientists' Blind Spot: How Evolution's Math Forgets Alleles That Lose

A beneficial allele's 'winning' trajectory is only half the story — the losing trajectories it ignores can be the dominant fact.

Even beneficial alleles get wiped out — and the standard math quietly deletes those histories, causing large errors.

Evolution is full of surprises that look like clean rules until you look closer. Here's one: a beneficial allele — a genetic variant that helps its carriers survive and reproduce — should, by the standard account, march toward fixation, spreading through the population until everyone carries it. The math seems to say so. Yet in reality, a helpful allele frequently vanishes entirely, wiped out by the lottery of genetic drift before it can do its work.

The discrepancy between those two stories—the clean deterministic march and the messy stochastic one—is the subject of a new theoretical paper by David Waxman of Fudan University. The paper delivers a striking verdict about how population geneticists have been computing allele-frequency statistics under strong selection: the standard shortcut, which keeps only the "winning" trajectories, can quietly delete a whole class of histories—the ones where the allele is lost—and produce large errors. Waxman shows how to fix this, and the fix matters for anyone who wants to predict how adaptation actually unfolds.

Consider what "strong selection" really means. The paper defines a compound parameter , where is the effective population size and is the selection coefficient of a focal allele. Selection is "strong" when . Crucially, this does not require dramatic biology. With a modest population of and a selection coefficient of just , you already get . In larger populations— around to —even a tiny gives anywhere from 20 to 2,000. "Strong selection," in other words, is an everyday phenomenon, not an exotic one. It says nothing about how fast allele frequencies change; it only says that deterministic selection dominates the random churn of drift. An allele can change slowly and still be strongly selected.

The Science

The paper works within the diffusion approximation, the continuous-state, continuous-time framework descended from the foundational work of Fisher (1922), Wright (1931), and Kimura (1955). Under this approximation, the random allele frequency evolves according to a stochastic differential equation with two competing terms: a deterministic push from selection and a random push from drift:

The first term is the directional bias of natural selection; the second is the unpredictable randomness of a finite population. Both terms vanish at frequencies 0 and 1, the absorbing boundaries where an allele is either lost or fixed.

Waxman rescales time to , which conveniently sets the selection coefficient to unit magnitude and leaves drift with a noise strength proportional to , where . This is the "small noise" regime: when is large, is small, and the noise contribution looks negligible. That visual impression is precisely the trap.

Here is the central observation. If you naively expand the frequency in powers of the noise, the leading term is a purely deterministic trajectory . For a beneficial allele, this trajectory climbs smoothly toward frequency 1.

Figure 1: Deterministic trajectories following from the vanishing ε\varepsilon limit of
Eq. (10). In this figure we plot two deterministic trajectories, Yτ,0Y_{\tau,0}, against
the rescaled time, τ\tau. The trajectories follow from the
zero ε\varepsilon limit of dY=τσY(1−Y)ττdτ+ε⋅Yτ(1−Yτ)dWτdY{{}_{\tau}}=\sigma\,Y{{}_{\tau}}(1-Y{{}_{\tau}})\,d\tau+\sqrt{\varepsilon}\cdot\sqrt{Y{{}_{\tau}}(1-Y{{}_{\tau}})}\,dW{{}_{\tau}}. The initial
value is Y0=0.2Y_{0}=0.2 and σ\sigma takes the values ±1\pm 1.
Figure 1: Deterministic trajectories following from the vanishing ε\varepsilon limit of Eq. (10). In this figure we plot two deterministic trajectories, Yτ,0Y_{\tau,0}, against the rescaled time, τ\tau. The trajectories follow from the zero ε\varepsilon limit of dY=τσY(1−Y)ττdτ+ε⋅Yτ(1−Yτ)dWτdY{{}_{\tau}}=\sigma\,Y{{}_{\tau}}(1-Y{{}_{\tau}})\,d\tau+\sqrt{\varepsilon}\cdot\sqrt{Y{{}_{\tau}}(1-Y{{}_{\tau}})}\,dW{{}_{\tau}}. The initial value is Y0=0.2Y_{0}=0.2 and σ\sigma takes the values ±1\pm 1. Source: David Waxman

shows two such deterministic paths from an initial frequency of 0.2: one rising to fixation for positive selection, one falling to loss for negative selection. The natural intuition is that the real, noisy trajectory is just the deterministic one with small fluctuations on top.

But that picture is wrong in a way that matters. A beneficial allele starting at low frequency has a real, sometimes large probability of being lost entirely. The deterministic trajectory—which always approaches fixation—gives the wrong long-time answer, because the true long-time mean allele frequency is precisely the fixation probability, not 1. If loss is common, the deterministic trajectory and the truth can disagree badly.

Waxman's solution is to force both possible outcomes into the analysis from the start. He writes any statistic as a sum of two contributions, one conditioned on ultimate fixation and one on ultimate loss, each weighted by its probability:

Here and are the fixation and loss probabilities, given by Kimura's classic results, and and are modified processes that guarantee eventual loss or fixation respectively, constructed using Doob's h-transform. Rather than discarding the "losing" trajectories as higher-order corrections, this representation insists they contribute. Then Waxman applies a small-noise expansion to each conditional piece, keeping the deterministic contribution plus the fluctuation correction.

What They Found

The headline result is that omitting the loss trajectories is not a small correction—it is a structural error. The naive single-trajectory approximation, which keeps only the path to fixation, can badly misrepresent statistics even when selection is strong. The fix, by contrast, works.

The paper tests the theory against numerically exact Wright-Fisher simulations, and the agreement is the heart of the empirical case.

Fixation probability vs. initial frequency (R=50)

The three initial frequencies used in the numerical comparisons, and the corresponding fixation probabilities under the diffusion approximation. Effective population size 2e3, selection coefficient 0.0125, R=50.

Fixation probability vs. initial frequency (R=50)
LabelValue
Initial frequency 0.0010.1
Initial frequency 0.0070.5
Initial frequency 0.0230.9

summarizes the key parameter regime for the numerical comparisons: effective population size , selection coefficient , giving . Three initial frequencies were used, corresponding to fixation probabilities of roughly 0.1, 0.5, and 0.9.

The mean frequency results are shown in

Figure 2: Results
for the mean frequency as a function of the rescaled time. The figure
contains three well-separated sets of curves, one set for each of the initial
frequencies y=0.001y=0.001, y=0.007y=0.007, and y=0.023y=0.023. 
Each set of
three curves contains (i) the deterministic contribution, ⟨Zτ,0⟩π\langle Z_{\tau,0}\rangle_{\pi}; (ii) the deterministic plus fluctuation contribution,
⟨Zτ,0⟩π+⟨Mτ⟩π/R\langle Z_{\tau,0}\rangle_{\pi}+\langle M_{\tau}\rangle_{\pi}/R; (iii) the exact Wright-Fisher result, 𝐄y​[Yτ(W​F)]\mathbf{E}_{y}[Y_{\tau}^{(WF)}].
For the figure, the census and
effective population sizes were set equal. For all curves, the effective
population size is Ne=2×103N_{e}=2\times 10^{3}, the selection coefficient is
s=0.0125s=0.0125, and the scaled measure of selection R=50R=50. The three initial frequencies correspond,
under the diffusion approximation, to fixation probabilities of approximately
0.10.1, 0.50.5, and 0.90.9.
Figure 2: Results for the mean frequency as a function of the rescaled time. The figure contains three well-separated sets of curves, one set for each of the initial frequencies y=0.001y=0.001, y=0.007y=0.007, and y=0.023y=0.023. Each set of three curves contains (i) the deterministic contribution, ⟨Zτ,0⟩π\langle Z_{\tau,0}\rangle_{\pi}; (ii) the deterministic plus fluctuation contribution, ⟨Zτ,0⟩π+⟨Mτ⟩π/R\langle Z_{\tau,0}\rangle_{\pi}+\langle M_{\tau}\rangle_{\pi}/R; (iii) the exact Wright-Fisher result, 𝐄y​[Yτ(W​F)]\mathbf{E}_{y}[Y_{\tau}^{(WF)}]. For the figure, the census and effective population sizes were set equal. For all curves, the effective population size is Ne=2×103N_{e}=2\times 10^{3}, the selection coefficient is s=0.0125s=0.0125, and the scaled measure of selection R=50R=50. The three initial frequencies correspond, under the diffusion approximation, to fixation probabilities of approximately 0.10.1, 0.50.5, and 0.90.9. Source: David Waxman

. For each initial frequency there are three curves: the deterministic-only contribution, the deterministic-plus-fluctuation approximation, and the exact Wright-Fisher result. The pattern is illuminating. When the initial frequency is high (fixation probability near 0.9), even the deterministic-only curve tracks the exact result reasonably well, because loss is rare. But as the initial frequency drops—fixation probability near 0.5 or 0.1—the deterministic-only approximation begins to drift away from the truth. Adding the fluctuation term pulls it back into line with the exact result. The paper is explicit that the deterministic contribution alone "may not always be sufficient for good accuracy," and that including fluctuations is what recovers "reasonable accuracy."

The variance tells the same story, and it is arguably the more demanding test.

Figure 3: Results for the variance of the frequency as
a function of rescaled time, τ\tau. In all panels we plot three versions of the
variance of allele frequency as a function of τ\tau: (i) the contribution
from deterministic trajectories, V0V_{0}; (ii) the contribution from
deterministic trajectories and fluctuations, V0+V1/RV_{0}+V_{1}/R; (iii)
the exact Wright-Fisher result, Vary⁡(Yτ(W​F))\operatorname{Var}_{y}\big(Y_{\tau}^{(WF)}\big).

For all panels, the census size was set equal to the effective population size.
The parameter values that
were common to all three panels are as follows: effective population size
Ne=2×103N_{e}=2\times 10^{3}, selection coefficient s=0.0125s=0.0125, scaled measure of
selection R=50R=50. The single parameter that was different for all three panels
was the initial frequency, yy, with the value indicated in the top right-hand corner
of each panel.
Figure 3: Results for the variance of the frequency as a function of rescaled time, τ\tau. In all panels we plot three versions of the variance of allele frequency as a function of τ\tau: (i) the contribution from deterministic trajectories, V0V_{0}; (ii) the contribution from deterministic trajectories and fluctuations, V0+V1/RV_{0}+V_{1}/R; (iii) the exact Wright-Fisher result, Vary⁡(Yτ(W​F))\operatorname{Var}_{y}\big(Y_{\tau}^{(WF)}\big). For all panels, the census size was set equal to the effective population size. The parameter values that were common to all three panels are as follows: effective population size Ne=2×103N_{e}=2\times 10^{3}, selection coefficient s=0.0125s=0.0125, scaled measure of selection R=50R=50. The single parameter that was different for all three panels was the initial frequency, yy, with the value indicated in the top right-hand corner of each panel. Source: David Waxman

shows the variance of allele frequency as a function of rescaled time across three panels, one for each initial frequency, each again comparing the deterministic contribution, the deterministic-plus-fluctuation version, and the exact Wright-Fisher result. Variance is a tricky statistic: it can rise as trajectories fan out, then fall as they coalesce at the absorbing boundaries. The transient humps and valleys in these curves are precisely where approximations tend to fail, and precisely where the fluctuation term earns its keep.

The key finding, captured in

Long-time mean allele frequency equals the fixation probability

Reference values of the three fixation probabilities studied, corresponding to initial frequencies y=0.001, 0.007, and 0.023. These are the long-time limits of the mean allele frequency, and are the quantities a naive deterministic trajectory fails to capture.

Long-time mean allele frequency equals the fixation probability
LabelValue
Fixation probability 0.10.1
Fixation probability 0.50.5
Fixation probability 0.90.9

, is that the variance's deviation from the deterministic-only prediction is roughly as large as—or larger than—the deep-signal value it is trying to capture, particularly at intermediate and low fixation probabilities. The fluctuation correction is what rescues the approximation in these regimes.

Why This Changes Things

This work is a corrective to a quiet assumption that has been baked into much of theoretical population genetics. The intuition that "when selection is strong, just follow the deterministic path and add small wobbles" is widespread, and Waxman shows it is incomplete in a specific, quantifiable way. The losing class of trajectories is not a negligible minority; it can represent an appreciable fraction of realized evolutionary histories, especially when the allele starts rare. And once those histories are admitted, they do not merely add a small perturbation—they change the structure of the statistic.

The practical payoff is two-fold. First, the paper supplies a computationally efficient analytical tool. Rather than running large numbers of Wright-Fisher simulations—expensive when populations are big and time spans long—one can evaluate the two-term approximation directly. The expressions are closed-form in spirit and fast to evaluate, making them attractive for inference problems where many parameter combinations must be scanned.

Second, the correct handling of conditioning matters for how we interpret allele-frequency data. The diffusion approximation is a workhorse for inferring selection coefficients from observed allele trajectories (Tataru et al., 2017), and the field is actively extending inference across the full range of selection strengths, from weak to strong (Lacerda and Seoighe, 2014). If the strong-selection regime is exactly where standard approximations break down—and this paper shows it is—then the conditioning-based expansion is a natural foundation for more accurate inference in biologically important contexts where selection is strong relative to drift (Gillespie, 1983).

There is also a deeper conceptual reframing here. The paper makes vivid that evolution's outcomes are not settled by the "average" trajectory but by the two competing destinations—fixation and loss—that every trajectory is ultimately heading toward. The deterministic picture that "beneficial alleles fix" is a statement about the winning class only. The honest picture requires holding both classes in mind simultaneously, weighted by their probabilities, and acknowledging that the allele's fate is not written in its selection coefficient alone but in the race between selection and the finite size of the population.

The loss trajectories are not noise in the statistical sense; they are structure. They are the reason a beneficial allele can start at frequency 0.023—near enough to the 0.1 fixation-probability threshold to be "fixed" in a naive sense—and still be lost most of the time in the exact simulation. That is not a second-order effect; it is the dominant fact about the allele's short- and long-term behavior.

What's Next

The paper is deliberately focused: it treats the mean and the variance, with mutation neglected and timescales where no new variants arrive. That is a controlled and defensible starting point, and the appendices lay out the machinery in careful detail. But it also suggests a clear direction for extension.

Natural next steps would include generalizing to other statistics—heterozygosity, the distribution itself, fixation-time distributions—and bringing mutation back into the picture over longer timescales. The two-class framing (fixation versus loss) is exact only when mutation is negligible; with recurrent mutation, the absorbing boundaries become reflecting, and the clean dichotomy dissolves. Extending the conditioning approach to that regime would be a substantial and valuable contribution.

There is also the question of how far beyond the approximation holds, and where its edges fray. The paper is honest about accuracy being "reasonable" rather than exact, and mapping the boundaries of that reasonableness—across initial frequencies, selection strengths, and higher moments—is a natural program for follow-up work.

For the nonspecialist, the takeaway is elegantly simple. When a trait or allele is under strong selection, the temptation is to describe its fate with a single smooth curve marching toward victory. This paper is a reminder that evolution always runs two scripts in parallel—triumph and disappearance—and that understanding either requires keeping both in view. The mathematics to do so, Waxman shows, is within reach, and it is more faithful to reality than the shortcut we have been leaning on.

"Loss does not represent a higher-order correction but rather a substantial contribution to the statistic."

Comments (0)

No comments yet. Be the first to share your thoughts.