The Cutoff Frequency: A New Mathematical Framework for Understanding Cancer Evolution
Mathematicians derive exact formulas for how advantageous mutations spread through growing cell populations, revealing a sharp "cutoff frequency" that partition
A precise mathematical 'cutoff frequency' partitions clones into two evolutionary regimes, with implications for
The Hidden Geometry of Cancer's Evolution
In the mathematics of life, every tumor tells a story. Not of the patient it inhabits, but of the evolutionary forces that shaped it — the mutations that struck lucky, the clones that surged to dominance, the vast silent populations that never quite made it. For decades, scientists have tried to read these stories from the scattered wreckage of genomic data: the mutations found in biopsy samples, the frequencies of different genetic variants, the shape of what's called the site frequency spectrum.
Now, a team of mathematicians from Johns Hopkins, UC Berkeley, the University of Minnesota, and the University of Iceland has cracked open a fundamental problem in cancer evolutionary biology. Their paper, "The Site Frequency Spectrum in an Exponentially-Growing Population with Selection," published in July 2026, provides the first complete mathematical description of how advantageous mutations spread through a growing population — the exact situation that plays out every time a precancerous clone begins to expand (Ahmed et al., 2026).
The finding that matters most: they've discovered a precise "cutoff frequency" — a specific proportion of the population at which the behavior of mutant clones changes dramatically. Below this threshold, clones are numerous but small, their fates dominated by random chance. Above it, clones are rare but powerful, their growth driven by selection's unforgiving logic. This cutoff isn't a fuzzy concept or an approximation — it's a number, derived from first principles, that could eventually help oncologists read the evolutionary history of a tumor from a single tissue sample.
"Understanding genetic tumor heterogeneity remains an important goal as it can provide clinically-relevant information regarding the mechanisms of tumor formation, predicting therapeutic response, classifying tumors," the researchers write. Their paper is a mathematical instrument for doing exactly that.
The Science
Reading the Genome Like a Historian
When cells divide, they occasionally make mistakes — mutations that get copied into daughter cells and then into their daughters, and so on. Over a lifetime, these mutations accumulate like archaeological layers. By the time a tumor is detected, it contains thousands of distinct genetic variants, each present at some frequency in the overall cell population.
The site frequency spectrum — SFS, to those who study it — is a way of organizing this information. If you sort all the mutations in a tumor by how common they are, you get a distribution: how many mutations appear in just 1% of cells, how many in 2%, how many in 50%, how many in nearly every cell. This distribution is not random. It carries the fingerprints of evolutionary processes — mutation rates, cell birth and death rates, selection pressures.
Imagine a library where every book is a mutation, and the number of copies of each book tells you how widespread that mutation became. The SFS is the catalog that counts how many books have 1 copy, how many have 10, how many have 1,000. A library that grew quickly will have many rare books (mutations present in few cells). A library shaped by strong selection will have a distinctive pattern of common books, the ones that were "selected for."
Mathematically, the SFS at time t is denoted S_j(t), where j represents a frequency and S_j(t) counts how many distinct mutations exist at that frequency. If j = 10 (meaning 10% of the population), S_10(t) tells you how many different mutations each occupy exactly 10% of cells. The full spectrum is the collection of all these counts across all frequencies.
The Model: Birth, Death, and Advantage
To understand the SFS theoretically, the researchers built a mathematical model of population growth with two cell types. Type-0 cells are "wild-type" — the normal cells that make up most of the tissue. Type-1 cells carry a driver mutation that gives them a fitness advantage: when they divide, they do so slightly faster than their wild-type neighbors.
This is modeled as a continuous-time birth-death process, a workhorse of mathematical biology. In this framework, each cell lives for a random time, then either divides (producing an offspring) or dies. The key parameter is the net growth rate: how much faster the population grows than it shrinks. Wild-type cells have one growth rate; mutant cells have a higher one. The difference — the "selection coefficient," often denoted s — measures how much fitter the mutants are.
"Supercritical" means the population is growing overall. This is realistic for many biological contexts: early tumor development, stem cell populations, bacterial colonies. In a supercritical process, clones have room to expand, and selection has room to act.
Mutation flows one direction: wild-type cells can give rise to mutant cells (with some probability at each division), but the researchers assume mutant cells don't revert to wild-type — a reasonable assumption for most cancer driver mutations. Neutral "passenger" mutations accumulate in both cell types as they divide, marking the lineages like genetic barcodes.
Figure 1 from the paper illustrates this process. Each color represents a distinct mutation, and arrows trace how mutations propagate through the population. Wild-type cells (shown in one color) divide and give rise to mutant cells (shown in another). Passenger mutations hitchhike along with driver mutations, spreading through the population as the clone expands. This is the basic architecture of clonal evolution in cancer.
Mathematical Approach: Moments and Strong Laws
The researchers' primary mathematical tool is the analysis of moments — expected values, variances, and higher-order statistics of the SFS. If you could measure the SFS in thousands of identical tumors, the moments would describe the average pattern, the spread around that average, and more subtle features of the distribution.
They prove "strong laws of large numbers" for the driver SFS — mathematical results showing that as time goes on, the random, noisy SFS converges to its expected value. This is important because it means their formulas aren't just averages over hypothetical ensembles; they describe what actually happens in any sufficiently large population over sufficient time.
The proof technique involves constructing L²-approximations — mathematical objects that approximate the true SFS closely enough that convergence can be rigorously established. This is technically demanding work, requiring sophisticated probability theory.
One key extension: the researchers don't assume every driver mutation confers the same fitness advantage. In reality, different mutations have different effects — some barely help, some dramatically accelerate growth. So they extend their results to a setting where the selection coefficient s is itself a random variable, drawn from some distribution. This makes the model more realistic and more useful for real applications.
What They Found
The Mean Spectrum at Large Times
The first major result is an exact characterization of the mean driver SFS at large times. The mean, written E[S_j(t)] for the expected number of clones at frequency j, takes on a specific mathematical form involving the selection coefficient and the mutation rate.
As time becomes large, the mean SFS follows a power law. Roughly, the expected number of clones at frequency j scales like a constant divided by j. This means clones at low frequencies are numerous (many rare mutations), while clones at high frequencies are rare (few mutations have swept to dominance). This is intuitive — in a growing population, most mutations arise early when the population is small, so they start rare and stay rare. Only a few lucky or advantageous mutations manage to expand substantially.
More precisely, for large t, the mean SFS behaves like:
where μ is the mutation rate and s is the selection coefficient. The exponential term e^{-μj/s} creates a cutoff: mutations at frequencies larger than roughly s/μ are exponentially rare. This is the mathematical embodiment of the intuition that selection can only efficiently expand clones with sufficient fitness advantage.
Cutoff Frequency: The Phase Transition
Here is the paper's most striking result: there exists a specific frequency — the "cutoff frequency" — at which the number of clones transitions from many to few.
Below this cutoff, the expected number of clones at frequency j grows with time (as t increases, you get more and more rare clones accumulating). Above the cutoff, the expected number actually decreases with time — clones at high frequency are being removed by the stochastic dynamics of birth and death, while new ones are rarely created at such high frequencies.
The cutoff frequency itself is approximately s/μ. This makes intuitive sense: the selection coefficient s measures how much advantage a mutation gives; the mutation rate μ measures how often new mutations arise. If mutations are common relative to their advantage (high μ, low s), you get many clones at low frequencies. If mutations are rare but powerful (low μ, high s), you get fewer clones but they expand further.
This cutoff has profound practical implications. It partitions the SFS into two regimes: the "bulk" of rare clones at low frequencies, where the total number of clones grows with time, and the "tail" of dominant clones at high frequencies, where the number of clones is roughly constant. A tumor biologist who sequences a sample and counts clones at different frequencies could, in principle, use this framework to estimate s and μ — the fundamental parameters of tumor evolution.
Strong Laws and Convergence
The strong law results establish that, with probability 1, the SFS converges to its mean as time goes to infinity. In mathematical terms:
This convergence holds for all frequencies j, not just some. It means that for any given frequency, if you wait long enough, the actual random number of clones at that frequency will be arbitrarily close to its expected value. The noise washes out; the signal emerges.
The researchers prove this by constructing an L²-martingale — a mathematical object that tracks the deviation of S_j from its mean — and showing this martingale converges to zero. The technical conditions require the selection coefficient to be positive (supercritical) and the mutation rate to be finite.
Random Fitness Effects
Real mutations don't all have the same effect. Some driver mutations might increase fitness by 0.1%, others by 10%. To model this, the researchers introduce a distribution of selection coefficients: when a new driver mutation arises, its s is drawn from some probability distribution G(s). This makes the mathematics considerably more complex but considerably more realistic.
With random fitness, the mean SFS becomes an integral over the distribution of s:
The shape of this integral depends sensitively on G(s). If there's a lot of heterogeneity — many different fitness effects — the SFS will look different than if all mutations have similar effects. This offers a way to infer the distribution of fitness effects from observed SFS data, a question of great interest in evolutionary biology.
The strong law results also extend to this random-fitness setting, with additional technical conditions on G(s).
Intermediate and Large Clones
Finally, the researchers examine how the number of clones changes across frequency space. For "intermediate" frequencies — neither rare nor dominant — and "large" frequencies — clones approaching fixation — they derive asymptotic results.
The number of intermediate-frequency clones grows like log(t) — slowly, but measurably. The number of large clones is of order 1 — roughly constant over time. These are different scaling regimes with different evolutionary dynamics: intermediate clones are being constantly created by mutation and removed by competition, while large clones are the survivors, the champions of the selective process.
Mean Driver SFS: Clone Count vs Frequency
Shows exponential decay of mean clone count with increasing frequency, with the cutoff visible at intermediate frequencies.
| Label | Value |
|---|---|
| 0.01 | 100 |
| 0.05 | 18 |
| 0.10 | 4 |
| 0.20 | 0.5 |
| 0.30 | 0.08 |
| 0.40 | 0.01 |
Effect of Fitness Heterogeneity on SFS
Random fitness distribution produces a more slowly decaying SFS tail, reflecting heterogeneity in selection coefficients.
| Label | Value |
|---|---|
| 0.01 | 100 |
| 0.05 | 12 |
| 0.10 | 2 |
| 0.20 | 0.15 |
| 0.30 | 0.02 |
| 0.40 | 0.003 |
These charts illustrate the theoretical predictions. The first shows the mean number of clones as a function of frequency for different times, illustrating the growth of the spectrum and the emergence of the cutoff. The second compares the fixed-s model with the random-s model, showing how fitness heterogeneity changes the SFS shape.
Why This Changes Things
From Sequences to Parameters
The site frequency spectrum is widely used in population genetics, but almost exclusively for neutral mutations. When mutations don't affect fitness, the SFS is well understood mathematically, and researchers can use it to estimate parameters like mutation rates and effective population sizes. The classic example is human population genetics: by sequencing many individuals and counting rare variants, demographers infer the history of human population growth.
Selection complicates everything. With selection, the SFS is no longer independent of the fitness effects of mutations. The shape of the spectrum carries information about selection coefficients, not just mutation rates. This paper provides the mathematical framework to extract that information.
For cancer biology, this is transformative. Right now, when oncologists sequence a tumor, they can identify mutations and estimate their frequencies, but interpreting those frequencies is difficult. Why are there so many low-frequency mutations? What does it mean if a mutation is at 40% frequency — did it arise early, or did it have a huge selective advantage? The SFS framework, enriched by this paper's results, provides a principled way to answer these questions.
Estimating Fitness From Frequency Data
The cutoff frequency s/μ is particularly valuable because it's measurable from data. If you sequence a tumor and count the number of clones at different frequencies, you can look for where the number of clones stops growing and starts shrinking. That transition point estimates s/μ.
If you additionally have some independent estimate of the mutation rate μ — which cancer genomic projects are increasingly providing — you can solve for s. This gives you the average selective advantage of driver mutations in that tumor. Different tumors might have different s values, reflecting different evolutionary landscapes. A tumor with high s is one where selection is strong, where advantageous clones expand rapidly. A tumor with low s is one where evolution is more neutral, more of a random walk.
This is a genuinely new capability. Existing methods for estimating selection in tumors typically require phylogenetic trees reconstructed from multiple samples, or time-series data tracking clone frequencies over months or years. The SFS approach needs only a single sample, a single snapshot. It's mathematically harder, but operationally simpler.
Understanding Tumor Heterogeneity
Tumor heterogeneity — the presence of many distinct genetic subclones within a single tumor — is a major challenge in cancer treatment. A therapy that targets one clone may leave others untouched. A tumor that appears homogeneous under the microscope may contain vast genetic diversity beneath the surface.
The SFS quantifies this heterogeneity. The total number of distinct clones, integrated across all frequencies, tells you how genetically diverse the tumor is. The distribution across frequencies tells you about the evolutionary process that generated that diversity. A tumor with many rare clones and few common ones has experienced different evolutionary pressures than one with many intermediate-frequency clones.
"The collection of all these various mutations precisely define the genetic tumor heterogeneity and contribute towards the total tumor heterogeneity," the researchers note. Their mathematical results give us tools to measure and compare heterogeneity across tumors, potentially identifying which tumors are likely to be more aggressive, more treatment-resistant, more likely to evolve.
The Mathematics of Clonal Sweeps
Beyond cancer, the results apply to any exponentially-growing population under selection. Bacterial populations evolve antibiotic resistance; viral populations adapt to new hosts; stem cell populations regenerate tissue. In all these contexts, driver mutations arise and expand, shaping the genetic composition of the population.
The "clonal sweep" is the process by which an advantageous mutation rises from rarity to dominance, dragging along neutral passengers in its wake. The SFS captures the aftermath of many overlapping sweeps, each mutation marking its frequency. Understanding the statistics of this process has been a central problem in population genetics for decades.
This paper's contribution is to provide exact, rigorous results for the supercritical birth-death process with selection. Previous work had studied related models — the Wright-Fisher process, the Moran process, branching processes — but none captured the specific combination of features relevant to cancer evolution: continuous time, linear birth-death dynamics, two types with different fitness, ongoing mutation.
The strong law results are particularly noteworthy. Proving that the SFS converges almost surely to its mean required constructing novel mathematical objects and establishing new convergence theorems. These results tell us that the stochastic noise in the SFS is not irreducible; there is signal beneath the noise, and with enough data, it emerges.
What's Next
From Theory to Application
The most pressing next step is applying these theoretical results to real tumor data. The mathematical framework exists; what remains is to test whether it accurately describes actual tumors. This will require:
High-resolution single-cell sequencing. Current bulk sequencing methods average over millions of cells, making it difficult to measure the SFS precisely, especially at low frequencies where clones contain few cells. Single-cell sequencing can, in principle, count every mutation in every cell, giving a complete SFS. Practical challenges remain — amplification bias, dropout, cost — but the technology is advancing rapidly.
Benchmarking against known tumors. The theory could be tested in model systems where the truth is known: tumors in mice with controlled mutation rates, or cancer cell lines with engineered fitness effects. If the theory correctly predicts the SFS in these controlled settings, it gains credibility for real applications.
Developing inference software. The formulas in this paper are mathematical expressions involving integrals and special functions. Turning them into practical estimators requires numerical methods and software implementation. This is a non-trivial engineering challenge.
Extensions and Generalizations
The current model makes several simplifying assumptions that could be relaxed in future work:
Spatial structure. Real tissues are not well-mixed populations. Cells interact with their neighbors, forming spatial patterns of clones. Spatial models are mathematically much harder but biologically more realistic. Some clones may expand locally without ever achieving global dominance.
Multiple driver mutations. The model considers a single driver mutation (type-1). But real tumors accumulate multiple drivers, each conferring additional fitness gains. The evolutionary dynamics become a "march of the mutants," with each new driver starting its own clonal expansion on top of previous ones. The SFS in this setting is more complex, with nested subclones and composite frequencies.
Time-varying environments. Fitness is not static. A mutation that helps in one microenvironment might hinder in another (for example, a mutation that accelerates growth in the core of a tumor might reduce it at the invasive edge). The SFS under fluctuating selection is an open problem.
Non-exponential growth. The model assumes exponential growth of the overall population. In reality, tumors may experience logistic growth (slowing as they approach carrying capacity), or boom-bust cycles (treatment response and relapse). Extending the results to these settings would broaden the applicability.
Clinical Implications
The long-term goal is clinical utility. If the SFS framework can reliably estimate selection coefficients from tumor sequencing data, oncologists gain a new tool for prognostication and treatment planning.
A tumor with strong selection (high s) may be more likely to evolve resistance to therapy. Resistant clones, if they exist at low frequency at diagnosis, will expand under drug pressure. Identifying strong selection early might prompt more aggressive or combination therapy.
A tumor with weak selection (low s) may be more genetically stable, its evolution driven more by drift than by selective sweeps. Such tumors might respond differently to chemotherapy, which often works by exploiting rapid division (and thus selective pressures).
The researchers are careful to note that "our results allow for estimation of relevant evolutionary parameters, such as the fitness increase of mutant versus wild-type cells." This estimation is the bridge from theory to practice.
The Broader Vision
At its core, this paper is about the geometry of evolution. Mutation creates diversity. Selection prunes it. The SFS is the fingerprint left by this dual process, encoding information about both mutation rates and selection strengths.
Understanding this geometry matters because evolution is the fundamental force shaping living systems. Cancer is evolution happening in real time, in a tissue, with deadly consequences. Bacterial resistance is evolution in a hospital, with global implications. Stem cell dynamics is evolution in a bone marrow, determining blood cell production.
The mathematics developed here — the moment formulas, the strong laws, the cutoff frequency — provides a language for talking about these processes precisely. It transforms qualitative intuitions ("selection is strong in this tumor") into quantitative claims ("the selection coefficient is approximately 0.05 per generation").
"The SFS can be used to estimate evolutionary parameters of a phylogeny (mutation, birth, death rates) and also provide evidence for either neutral or selective evolution," the researchers note. Their work makes this estimation possible in a wider range of biological contexts than ever before.
Caveats and Honest Uncertainty
Any honest assessment must acknowledge what remains unknown.
First, the model assumes mutation rates and selection coefficients are constant over time. Real tumors may experience fluctuating mutation rates (exposure to mutagens, DNA repair defects), and selection coefficients may change as the microenvironment changes.
Second, the mathematical results are asymptotic — they describe the behavior of the SFS as time goes to infinity. Real tumors exist for finite times. The gap between finite-time behavior and asymptotic behavior is not fully characterized.
Third, the theory requires knowing which mutations are drivers and which are passengers. In practice, this is difficult. Most mutations are passengers; only a handful are drivers. Distinguishing them requires functional experiments or large-scale comparative genomics, not just frequency data.
Fourth, the clinical utility is hypothetical. The theory provides estimators, but whether those estimators perform well on real data — with all its noise, biases, and complications — remains to be seen.
These caveats are not reasons to dismiss the work. They are reasons to test it.
The Takeaway
Every cell in your body carries the mutations of its ancestors — a record of every division, every DNA copy, every error. Most of these errors are neutral, invisible to selection. But a few are drivers, tiny improvements that give a cell's lineage a slight edge. Over a lifetime, those edges compound. Clones expand. Some become large enough to detect. A rare few, in the wrong context, become cancer.
The mathematics of this process is what the researchers have formalized. Their site frequency spectrum is not just a statistical tool; it's a window into evolutionary history. By counting mutations at each frequency, by tracking the shape of that distribution, they can infer the forces that shaped it.
The cutoff frequency — that sharp threshold where the number of clones transitions from many to few — is the paper's most striking finding. It says that evolutionary dynamics are not smooth and gradual but have sharp transitions, critical points, phase changes. Below the threshold, diversity accumulates exponentially. Above it, dominance crystallizes.
This is the geometry of cancer's evolution: a landscape of rare clones at the bottom, a few dominant ones at the top, and in between, the constant churn of competition and selection.
Understanding that landscape is the first step toward navigating it — and, perhaps, toward intervening before a precancerous clone becomes a lethal tumor.
Sign in to join the conversation.
Comments (0)
No comments yet. Be the first to share your thoughts.