When Branches Fuse: Mathematicians Prove Evolution's Hidden Networks Are Detectable

The Science
To understand what Brits, Holtgrefe, van Iersel, and Martin have accomplished, you need to understand what phylogenetic networks actually are—and why biologists have needed them for decades.
A phylogenetic tree is a familiar object: a branching diagram showing how species or genes evolved from common ancestors. The trunk splits into branches, which split again, and so on, until you reach the leaves—which represent the living organisms we're studying. For much of the 20th century, this was the dominant framework for understanding evolutionary history.
But evolution isn't always a clean branching process. Sometimes, two lineages merge. A virus might swap genetic material with a completely different strain. Two species might hybridize, producing offspring that carries DNA from both parents. A bacterium might pick up genes from a neighbor through horizontal transfer. These events—which biologists call reticulate (from the Latin for "net-like")—create evolutionary histories that look less like trees and more like webs.
Phylogenetic networks are the mathematical framework scientists built to represent these more complicated histories. Instead of a simple branching structure, a network contains nodes where multiple lineages come together—nodes that don't exist in standard tree diagrams. When you draw one of these networks, it looks, appropriately, like a net: a tangled web of ancestry rather than a clean genealogical tree.
The question Brits and colleagues address is deceptively simple: if you could observe the genetic patterns in living organisms, could you figure out whether their evolutionary history was a tree or a network? And if it was a network, could you recover its exact structure?
This is what mathematicians call an identifiability problem. In statistics, identifiability asks whether the thing you're trying to measure can actually be recovered from the data you have. You can't estimate a parameter that isn't identifiable—if the data doesn't contain enough information to distinguish between two different values of that parameter, then no amount of computation will tease them apart.
To make this concrete, the researchers worked within a standard framework for modeling genetic evolution. When DNA is passed down through generations, letters change due to mutations. Biologists use probabilistic models—specifically the Jukes-Cantor, Kimura 2-parameter, and Kimura 3-parameter models—to describe how likely different types of mutations are. These models make assumptions about which substitutions are more or less probable. The Jukes-Cantor model treats all mutations as equally likely—a simplification that's useful for theoretical work. The Kimura models introduce more realism by distinguishing between different types of nucleotide changes.
Under any of these models, a phylogenetic network determines something called a leaf-pattern distribution. This is the probability distribution over the genetic patterns you would expect to see at the leaves of the network, given the model's assumptions about mutation. The leaf-pattern distribution is, in a sense, everything you could ever know about the network from genetic data—it's the complete statistical fingerprint that the evolutionary history leaves behind.
The question becomes: given this distribution, can you uniquely recover the network that produced it?
Prior work had established what mathematicians call generic identifiability—roughly, that you can recover the network structure for "most" choices of parameters. But "most" in mathematics is a slippery concept. A measure-zero subset of pathological cases could still cause problems in practice. If you're a biologist trying to analyze real data, you don't know whether your particular parameter values fall in the well-behaved region or the problematic one. You need to know that identifiability holds everywhere, not just generically.
Brits and colleagues focused on a specific class of networks called level-1 phylogenetic networks. Think of these as networks where the reticulate cycles don't interact with each other—the network has a clean structure that makes the math tractable. In a level-1 network, any reticulate event (a node where lineages merge) is isolated from the others. This is a biologically reasonable restriction: many real hybridizations and recombination events are indeed independent of each other.
Within this framework, the researchers proved two main results.
What They Found
The first result concerns what the researchers call the semi-directed network parameter. To understand what this means, consider that phylogenetic networks can be described at different levels of precision. A fully directed network specifies exactly which lineage contributed what to each reticulate event. A semi-directed network loses some of this directional information—it's less precise but easier to identify from data. The "modulo redirecting triangles" part refers to a technical simplification: in some network structures, you can't tell which way certain triangles are oriented, so the researchers collapse these equivalent structures together.
The key finding is that the semi-directed network parameter of a level-1 network is fully identifiable under all three models studied—JC, K2P, and K3P. This means that for every single choice of parameters (every possible pattern of mutation rates, every possible pattern of reticulate contributions), you can recover the network structure from the leaf-pattern distribution. Not almost all choices. Not generically. Every choice.
This is a strong result. In mathematical terms, full identifiability means that if two level-1 networks produce the same leaf-pattern distribution, they must be the same network (up to the triangle equivalence). The mapping from network structure to statistical distribution is injective—each network maps to a unique distribution, with no two networks colliding onto the same fingerprint.
The second result addresses a more fundamental question: can you even tell whether evolution was reticulate? If a network and a tree produce the same leaf-pattern distribution, then no amount of genetic data can distinguish them—you simply cannot tell from the patterns whether the history involved merging lineages or just branching ones.
The researchers proved that this cannot happen—except in specific, well-characterized cases. Specifically, a level-1 network and a phylogenetic tree can induce the same leaf-pattern distribution only if the network is actually a tree (no reticulation), or if the network is a tree augmented with certain substructures called 2-blobs. A 2-blob is a piece of the network where two lineages merge but where the statistical signature of that merger gets absorbed into the mutation process in a way that makes it invisible.
This is the result that has the most practical implications. It means that, in the vast majority of cases, reticulate evolution leaves a detectable signature in the leaf-pattern distribution. The presence of hybridization, recombination, or horizontal gene transfer isn't just noise that obscures the true history—it's information that the statistical models can pick up on.
The researchers proved this result for the JC and K2P models; they note that extending it to K3P remains an open problem. The K3P model is more complex because it distinguishes between different types of nucleotide changes in more detail, which creates additional mathematical complications.
They also extended these results beyond the Markov models to several coalescent-based models. The coalescent is a mathematical framework for modeling how genetic lineages merge backward in time—it's the probabilistic foundation for many population genetics methods. The researchers showed that their identifiability results have consequences here too: if a network and tree can't be distinguished under the Markov models, they generally can't be distinguished under these coalescent models either.
Why This Changes Things
Let's step back from the mathematics and ask: why should anyone who isn't a phylogeneticist care about these results?
The answer lies in how modern biology reconstructs evolutionary history. When you hear about studies that trace the origins of SARS-CoV-2, or that map how antibiotic resistance genes spread through bacterial populations, or that reconstruct the relationships between early human populations—all of these rely on phylogenetic methods. And increasingly, those methods have to account for reticulate evolution.
Horizontal gene transfer, for instance, is rampant in bacteria. When a bacterium picks up genes from its environment or from other bacteria, that's not a tree-like process—it's a network. Ignoring this network structure and force-fitting the data onto a tree can lead to serious errors. You might conclude that two species are closely related because they share similar genes, when in fact one borrowed those genes from the other.
Similar issues arise with viral evolution. RNA viruses like influenza and coronavirus recombine freely. Plant hybridization is common in agriculture and natural ecosystems. Even in human evolution, there's increasing evidence of interbreeding between modern humans and archaic groups like Neanderthals and Denisovans.
For decades, statisticians and bioinformaticians have developed methods to infer phylogenetic networks from genetic data. But these methods all depend on assumptions—about the mutation process, about the network structure, about which models are "good enough" to capture the relevant biology. If the network isn't identifiable from the data, then no inference method, however clever, can recover it. The signal simply isn't there.
The Brits et al. result changes this calculus. By proving that level-1 networks are fully identifiable under standard mutation models, they give practitioners a solid theoretical foundation. They can trust that, at least in principle, the network structure is recoverable—that the problem is solvable, even if the practical algorithms still need work.
The second result is equally important. It says that reticulation leaves detectable signatures—that the evidence for network-like evolution is preserved in the data, not washed out by mutation and drift. This isn't guaranteed a priori. You might imagine that if lineages merge and then evolve for many generations, the evidence of that merger could be overwritten, making the history look tree-like. The researchers' result says this doesn't happen in general. The "signature" of reticulation survives.
This matters for study design. If reticulate evolution is detectable, then researchers should actively look for it rather than defaulting to tree-based methods. It argues for the broader adoption of network-based approaches in phylogenetics, and for the development of better statistical methods to detect and characterize reticulation.
It also matters for interpreting existing results. Many published phylogenetic trees were constructed under the assumption of strict branching evolution. If the true history involved reticulation, those trees might be misleading—and we'd have no way of knowing without reanalysis with network methods. The Brits et al. result suggests that, in principle, the data needed to detect these cases exists—it just needs to be analyzed correctly.
The coalescent connections are important too. The coalescent is the workhorse model for population genetics—it's used in everything from estimating demographic history to detecting selection. If coalescent-based methods can't distinguish certain networks from trees, that's a limitation that practitioners need to be aware of. The Brits et al. result formalizes this limitation, giving researchers a clear understanding of when tree-based methods are adequate and when network methods are necessary.
What's Next
The paper opens several doors, and closing them will occupy researchers for years.
The most immediate extension is to the Kimura 3-parameter model. The researchers proved full identifiability under JC and K2P, and they proved network-tree distinguishability under JC and K2P—but K3P remains open. The K3P model is more realistic than the others (it distinguishes between transitions and transversions, which have different biological rates), so results that don't extend to it are incomplete. Proving K3P identifiability will likely require new mathematical techniques.
Beyond K3P lies the question of level-2 networks and beyond. Level-1 networks have isolated reticulation events; level-2 networks allow some interaction between them. Biologically, this corresponds to scenarios where multiple hybridization or recombination events are correlated—where what's happening in one part of the network affects what's happening in another. These cases are more realistic, but the mathematics is harder. Identifiability may fail in some configurations, or it may require additional assumptions.
There's also the question of parameter estimation. Brits et al. proved that the network is identifiable—that the mapping from network to distribution is injective. But they didn't prove that the network can be efficiently computed from the distribution. Identifiability is a necessary condition for valid inference, but it's not sufficient. You also need algorithms that can actually do the inference, and that can handle real data with finite sample sizes, missing data, and model violation.
Practical implementation will require numerical methods for computing the likelihood of a network given observed patterns, optimization methods for searching over the space of possible networks, and statistical tests for deciding whether the data support a network over a tree. Each of these is a research program in its own right.
The connection to coalescent models also needs further development. Brits et al. showed that their results have consequences for coalescent-based models, but they didn't fully characterize what those consequences are. Understanding exactly how their results translate to population genetics settings—with all the complications of recombination, population structure, and demographic change—will be important for applying these theoretical results to real biological problems.
Finally, there's the question of robustness. The proofs in this paper rely on exact mathematical models of mutation. Real genetic data violates these assumptions in various ways—mutation rates vary across sites, substitution patterns deviate from simple models, and sequencing errors introduce noise. An important direction for future work is understanding how well the identifiability results hold up under model violation—whether the network structure remains recoverable even when the assumptions are only approximately satisfied.
Twenty years ago, most evolutionary biologists would have said that reconstructing network-like evolutionary histories was impossibly hard—that the mathematical and statistical challenges were insurmountable. Brits, Holtgrefe, van Iersel, and Martin haven't solved all those challenges. But they've answered a fundamental question: whether the problem is solvable in principle. The answer is yes, at least for a reasonable class of networks and models. Now the field knows where to focus its efforts—on algorithms and computation, on relaxing assumptions, on extending to more complex network structures. The path forward is clearer than it was before.
For anyone studying evolution—whether of viruses, bacteria, plants, or humans—this work is a reminder that the tree of life has more connections than a tree. The branches don't just split; sometimes they fuse. And the mathematical tools to detect those fusions are finally catching up to the biological reality.
The full paper is available on arXiv at https://doi.org/10.48550/arXiv.2607.12919.