Meridia Insight Science Breakthroughs Knowledge

The Social Structures Epidemics Cannot See

Two populations could have radically different social structures—one where most transmission happens in small clusters, another where it happens in large gather

Two populations could have completely different social structures but generate identical epidemics—and we'd never know

In the summer of 2020, as COVID-19 surged around the world, epidemiologists faced a puzzle that haunted their models: How do people actually catch this disease? The obvious answer—that infection spreads through pairwise contacts—was true but incomplete. People get sick in households, at choir practice, in meatpacking plants, at family gatherings. They catch viruses in groups. And yet the mathematics of epidemic spread had long ago settled the question of how to model such settings. Or so it seemed.

A new theoretical result overturns that certainty in a way that reaches deep into what we can and cannot know about how diseases spread through populations. Roni Muslim, a physicist at the Asia Pacific Center for Theoretical Physics in South Korea, has proven that when epidemics spread through temporary groups—as they do in the real world—aggregate epidemic data can never uniquely reconstruct the true group-size distribution. Instead, the data can only identify an entire equivalence class of distributions that happen to produce identical disease dynamics. Two populations could have radically different social structures—one where most transmission happens in small clusters, another where it happens in large gatherings—and still produce mathematically indistinguishable epidemics (Muslim, 2026).

This isn't a limitation of our measurement tools or our data collection. It's woven into the mathematical structure of epidemic dynamics itself. No matter how precise our surveillance, no matter how comprehensive our case tracking, there are features of human social organization that the disease itself cannot reveal.

The Science

Epidemiologists have long known that who interacts with whom matters enormously for how a disease spreads. The seminal SIR and SIS models, developed in 1927 by Kermack and McKendrick, provided a mathematical framework for tracking how susceptible individuals become infected and then either recover with immunity (SIR) or return to susceptibility (SIS). These compartmental models reduced the complexity of human contact patterns to a single parameter: the basic reproduction number, ℛ₀—the average number of secondary infections produced by a single infected person in a fully susceptible population.

But ℛ₀ is a summary statistic, not a mechanism. It collapses all the messy reality of human social behavior into one number. To understand what happens when diseases spread through groups rather than pairs, Muslim introduced a model where transmission occurs through temporary groups that are independently reassembled at every exposure event. No persistent hyperedges, no permanent social structures retained from one interaction to the next. This "annealed" approximation serves as a reference model that isolates the pure effect of group-size distribution from topological correlations and persistent group structures.

The model works as follows: A focal susceptible individual encounters a group of size n, drawn from a participant-view distribution Q(n). The remaining n−1 partners are sampled uniformly without replacement from the other N−1 individuals in the population. If ℓ of these partners happen to be infected, the focal individual becomes infected at a rate λ_{n,ℓ} = β ℓ^q / (n−1)^η, where β is the baseline transmission scale, q is the order of the nonlinearity, and η controls how exposure depends on group size.

When q = 1, the contributions of infected partners add linearly—one infected person in a group contributes the same infection risk as any other. But when q > 1, something more interesting happens. The infection rate grows nonlinearly with the number of infected partners, capturing synergistic effects where simultaneous exposure to several infectious sources produces more than the sum of independent contact contributions. This could represent superspreading events, collective immune evasion, or any mechanism where having multiple infected contacts is more dangerous than merely having more contacts.

The critical question Muslim asked was this: After averaging over all possible group compositions—accounting for the fact that a group of five might contain zero, one, two, or more infected individuals depending on the current state of the epidemic—how much information about Q(n) survives in the aggregate infection rate?

The answer, proved exactly for finite populations, is unsettlingly little.

What They Found

The proof centers on what Muslim calls the "factorial-moment sufficiency theorem." After averaging the infection rate over all possible group compositions using the hypergeometric distribution, the aggregate Markov generator depends on the group-size distribution only through a specific vector of weighted factorial moments.

The moment-weights are defined as:

where (x)_{\underline{k}} = x(x−1)⋯(x−k+1) is the falling factorial and the sum runs from groups large enough to contain k distinct positions.

The theorem states that if two group-size distributions share the same first q moments—M₁^{(η)} through M_q^{(η)}—then they generate mathematically identical epidemic dynamics. Not approximately identical. Exactly identical. The probability distributions of infected counts over time are identical, including mean prevalence, fluctuations, survival probability, extinction-time distributions, and final outbreak sizes.

This equivalence is not a mean-field approximation or a large-population limit. It holds exactly for finite N.

To make this concrete, Muslim constructed two specific distributions with radically different shapes that nonetheless share the first two factorial moments. Distribution A places equal weight on groups of size 3 and 7 (Q_A(3) = Q_A(7) = 0.5), while distribution B concentrates equally on groups of size 2 and 12 (Q_B(2) = Q_B(12) = 0.5). These distributions have completely disjoint supports—one never produces groups of size 2 or 12, the other never produces groups of 3 or 7. Yet both satisfy M₁^{(0)} = 5 and M₂^{(0)} = 29, and both generate identical stochastic dynamics for a quadratic (q = 2) transmission kernel.

Figure 1: Schematic illustration of the epidemic model with temporary groups.
(a) The event-level group-size distribution, P​(n)P(n), is converted into the
participant-view distribution,
Q​(n)=n​P​(n)/⟨n⟩PQ(n)=nP(n)/\langle n\rangle_{P}. (b) A focal individual interacts with
n−1n-1 partners sampled without replacement. If ℓ\ell of these partners are
infected, the focal individual becomes infected at rate
λn,ℓ\lambda_{n,\ell}, whereas recovery occurs independently at rate μ\mu.
The group is resampled after every update.
Figure 1: Schematic illustration of the epidemic model with temporary groups. (a) The event-level group-size distribution, P​(n)P(n), is converted into the participant-view distribution, Q​(n)=n​P​(n)/⟨n⟩PQ(n)=nP(n)/\langle n\rangle_{P}. (b) A focal individual interacts with n−1n-1 partners sampled without replacement. If ℓ\ell of these partners are infected, the focal individual becomes infected at rate λn,ℓ\lambda_{n,\ell}, whereas recovery occurs independently at rate μ\mu. The group is resampled after every update. Source: Roni Muslim

Figure 1: The epidemic model with temporary groups. (a) The event-level group-size distribution P(n) is converted into the participant-view distribution Q(n) = nP(n)/⟨n⟩P. (b) A focal individual interacts with n−1 partners sampled without replacement. If ℓ of these partners are infected, the focal becomes infected at rate λ{n,ℓ} = β ℓ^q/(n−1)^η.

The numerical simulations confirm the theorem precisely. When Muslim solved the master equation and ran Gillespie stochastic simulations for both distributions with N = 200, ℛ₀ = 0.48, and initial infected fraction i(0) = 0.30, the mean prevalence trajectories were indistinguishable. The survival probability curves—measuring the probability that the epidemic persists beyond time τ—overlapped exactly. Even the final outbreak-size distributions in the SIR model, computed exactly for N = 200, were identical.

The only way to distinguish these distributions is to change the rules of transmission. When Muslim introduced a weak higher-order channel—adding a cubic term ε ℓ̄₃ to the infection kernel to activate the third moment—the curves immediately separated. For distribution B, this perturbation crossed the bistability threshold at ε_{c,B} = 0.0216, while distribution A didn't become bistable until ε_{c,A} = 0.0258. Between these values, the two systems were genuinely in different dynamical regimes despite sharing identical behavior without the perturbation.

Figure 4: Moment hierarchy and nonidentifiability in SIS dynamics. All curves
are obtained from numerical solutions of the finite-population master
equation for N=200N=200, η=0\eta=0, and rescaled time τ=μ​t\tau=\mu t.
(a) The distributions Q0Q_{0} and QAQ_{A} generate identical mean-prevalence
trajectories for q=1q=1, ℛ0=0.58\mathcal{R}_{0}=0.58, and i​(0)=0.30i(0)=0.30 because
M1,0(0)=M1,A(0)=5M_{1,0}^{(0)}=M_{1,A}^{(0)}=5. (b) For the same parameters with q=2q=2,
the difference between M2,0(0)=20M_{2,0}^{(0)}=20 and M2,A(0)=29M_{2,A}^{(0)}=29
separates the curves. (c) The distributions QAQ_{A} and QBQ_{B} remain
dynamically identical for q=2q=2, ℛ0=0.101\mathcal{R}_{0}=0.101, and i​(0)=0.65i(0)=0.65
because both have M1(0)=5M_{1}^{(0)}=5 and M2(0)=29M_{2}^{(0)}=29. (d) For the same
parameters with q=3q=3, the curves separate because
M3,A(0)=168M_{3,A}^{(0)}=168, whereas M3,B(0)=201M_{3,B}^{(0)}=201. Solid and dashed curves
distinguish the distributions being compared. Open symbols serve only as
visual markers and do not represent separate simulation data.
Figure 4: Moment hierarchy and nonidentifiability in SIS dynamics. All curves are obtained from numerical solutions of the finite-population master equation for N=200N=200, η=0\eta=0, and rescaled time τ=μ​t\tau=\mu t. (a) The distributions Q0Q_{0} and QAQ_{A} generate identical mean-prevalence trajectories for q=1q=1, ℛ0=0.58\mathcal{R}_{0}=0.58, and i​(0)=0.30i(0)=0.30 because M1,0(0)=M1,A(0)=5M_{1,0}^{(0)}=M_{1,A}^{(0)}=5. (b) For the same parameters with q=2q=2, the difference between M2,0(0)=20M_{2,0}^{(0)}=20 and M2,A(0)=29M_{2,A}^{(0)}=29 separates the curves. (c) The distributions QAQ_{A} and QBQ_{B} remain dynamically identical for q=2q=2, ℛ0=0.101\mathcal{R}_{0}=0.101, and i​(0)=0.65i(0)=0.65 because both have M1(0)=5M_{1}^{(0)}=5 and M2(0)=29M_{2}^{(0)}=29. (d) For the same parameters with q=3q=3, the curves separate because M3,A(0)=168M_{3,A}^{(0)}=168, whereas M3,B(0)=201M_{3,B}^{(0)}=201. Solid and dashed curves distinguish the distributions being compared. Open symbols serve only as visual markers and do not represent separate simulation data. Source: Roni Muslim

Figure 3: Moment hierarchy and nonidentifiability. (a) For q = 1, distributions Q₀ and Q_A generate identical trajectories because they share M₁ = 5. (b) For q = 2, they separate because Q₀ has M₂ = 20 while Q_A has M₂ = 29. (c) Q_A and Q_B remain identical for q = 2 because both have M₁ = 5 and M₂ = 29. (d) For q = 3, they separate because M₃,A = 168 while M₃,B = 201.

The deterministic limit reveals why the moments matter so much. As the population size grows large, the system transitions from a stochastic Markov process to a set of ordinary differential equations. In this limit, the first moment M₁^{(η)} determines the invasion threshold: the disease can only invade if ℛ₀ M₁^{(η)} > 1, where ℛ₀ = β/μ is the basic reproduction number scaled by the recovery rate.

But the second moment controls something qualitatively different—the nature of the epidemic transition itself. For a linear (q = 1) transmission kernel, the disease-free state loses stability continuously as ℛ₀ increases past 1, and the endemic prevalence rises smoothly from zero. For a quadratic (q = 2) kernel, the second moment χ = M₂^{(η)}/M₁^{(η)} determines whether the transition remains continuous or becomes discontinuous.

When χ = 1, the system has a tricritical point where the transition changes character. For χ > 1, a bistable regime emerges: for a range of ℛ₀ values near the invasion threshold, both the disease-free state and an endemic state are locally stable, separated by an unstable intermediate fixed point. The system exhibits hysteresis—depending on initial conditions, the same ℛ₀ can produce either disease extinction or sustained endemic infection.

The boundaries are sharp. The invasion threshold occurs at ℛ₀ = 1. The saddle-node bifurcation that closes the bistable region occurs at:

These two boundaries meet precisely at the tricritical point (χ, ℛ₀) = (1, 1).

Figure 2: Deterministic phase structure of the SIS model with the quadratic
infection kernel q=2q=2. (a) Stationary infected fraction i∗i^{*} as a
function of ℛ0\mathcal{R}_{0} for χ=0.5\chi=0.5, 11, and 44. Solid and dashed
curves denote stable and unstable fixed points, respectively. The open
circle marks the saddle-node bifurcation for χ=4\chi=4, and the
vertical dotted line indicates ℛ0=1\mathcal{R}_{0}=1. (b) Phase diagram in the
(χ,ℛ0)(\chi,\mathcal{R}_{0}) plane. The black line marks the linear stability
boundary of the disease-free state, ℛ0=1\mathcal{R}_{0}=1, while the red curve
shows the saddle-node boundary
ℛSN=4​χ/(1+χ)2\mathcal{R}_{\mathrm{SN}}=4\chi/(1+\chi)^{2}. The two boundaries meet at
the tricritical point (χ,ℛ0)=(1,1)(\chi,\mathcal{R}_{0})=(1,1). The blue, orange, and
green regions correspond to the disease-free, bistable, and
single-endemic-attractor regimes, respectively. (c) Evolution of the
infected fraction in the bistable regime for χ=4\chi=4 and
ℛ0=0.8\mathcal{R}_{0}=0.8. The horizontal red dashed line marks the unstable
fixed point i−∗=0.095i_{-}^{*}=0.095, whereas the black dash-dotted line marks the
stable endemic fixed point i+∗=0.655i_{+}^{*}=0.655. Initial conditions below i−∗i_{-}^{*}
approach the disease-free state, whereas those above i−∗i_{-}^{*} approach the
endemic state.
Figure 2: Deterministic phase structure of the SIS model with the quadratic infection kernel q=2q=2. (a) Stationary infected fraction i∗i^{*} as a function of ℛ0\mathcal{R}_{0} for χ=0.5\chi=0.5, 11, and 44. Solid and dashed curves denote stable and unstable fixed points, respectively. The open circle marks the saddle-node bifurcation for χ=4\chi=4, and the vertical dotted line indicates ℛ0=1\mathcal{R}_{0}=1. (b) Phase diagram in the (χ,ℛ0)(\chi,\mathcal{R}_{0}) plane. The black line marks the linear stability boundary of the disease-free state, ℛ0=1\mathcal{R}_{0}=1, while the red curve shows the saddle-node boundary ℛSN=4​χ/(1+χ)2\mathcal{R}_{\mathrm{SN}}=4\chi/(1+\chi)^{2}. The two boundaries meet at the tricritical point (χ,ℛ0)=(1,1)(\chi,\mathcal{R}_{0})=(1,1). The blue, orange, and green regions correspond to the disease-free, bistable, and single-endemic-attractor regimes, respectively. (c) Evolution of the infected fraction in the bistable regime for χ=4\chi=4 and ℛ0=0.8\mathcal{R}_{0}=0.8. The horizontal red dashed line marks the unstable fixed point i−∗=0.095i_{-}^{*}=0.095, whereas the black dash-dotted line marks the stable endemic fixed point i+∗=0.655i_{+}^{*}=0.655. Initial conditions below i−∗i_{-}^{*} approach the disease-free state, whereas those above i−∗i_{-}^{*} approach the endemic state. Source: Roni Muslim

Figure 2: Deterministic phase structure for q = 2. (a) Stationary infected fraction i* as a function of ℛ₀ for χ = 0.5, 1, and 4. The bistable regime (χ = 4) shows a discontinuous transition with an unstable intermediate fixed point. (b) Phase diagram showing the disease-free, bistable, and single-endemic regimes. The saddle-node boundary meets the ℛ₀ = 1 line at the tricritical point (1, 1).

Why This Changes Things

The implications reach far beyond the mathematics of moment equivalence. This result strikes at the heart of what epidemic models can tell us about the social structures that underlie disease transmission.

Public health officials have long struggled to infer "who infected whom" from case data. Contact tracing attempts to reconstruct transmission chains, but many infections go undetected, and even perfect contact data cannot directly observe the group contexts in which transmission occurred. The standard assumption has been that with sufficiently comprehensive data, it should be possible to reverse-engineer the underlying contact patterns.

Muslim's theorem shows that this assumption is fundamentally wrong—not because our data are insufficient, but because the disease dynamics themselves don't contain enough information. The aggregate epidemic process is blind to most features of the group-size distribution. It sees only a few factorial moments; the rest is mathematically invisible.

This structural nonidentifiability differs from practical nonidentifiability, which arises from measurement noise, limited observations, or weak output sensitivity. Additional experiments or observables might reduce practical degeneracy at the output level, but they cannot distinguish two models with identical stochastic generators. No amount of careful data collection can recover information that the mathematics has already discarded.

For epidemic modeling, this creates a fundamental tension. On one hand, we want models that capture the full complexity of human contact patterns—households, workplaces, schools, public transit, social gatherings, each with different size distributions and transmission dynamics. On the other hand, the aggregate epidemic data can only constrain a few summary statistics of these distributions. The rest is free parameters that cannot be identified from epidemic dynamics alone.

The moment hierarchy also explains why some features of group organization matter more than others. The first moment—the average group size experienced by a participant—sets the invasion threshold. If you're trying to predict whether a disease can establish itself in a population, you only need to know the average group size (adjusted for the size-bias of participation). The second moment controls whether you'll see bistability and hysteresis. If you want to predict whether small perturbations near the epidemic threshold will die out or trigger sustained outbreaks, you need to know the variance of group sizes. But the finer details—the exact shape of the distribution, its skewness, its tail—cannot be recovered from epidemic data alone.

This has practical consequences for model selection and parameter estimation. When fitting models to data, different group-size distributions that share the same factorial moments are genuinely indistinguishable. Bayesian inference will spread probability across all of them, potentially understating uncertainty by treating each as a separate hypothesis when they're actually one equivalence class.

It also matters for evaluating interventions. If you can't identify the true group-size distribution from epidemic data, you can't definitively attribute changes in epidemic dynamics to interventions targeting specific group sizes. A lockdown that closes large venues might look identical to an intervention that reduces participation in medium-sized groups, if both produce the same change in the first two factorial moments. The observed epidemic dynamics can't tell you which intervention actually occurred.

What's Next

The theorem leaves several important questions open. Muslim's model assumes independently reassembled temporary groups with no correlations between groups or participation patterns. Real populations have persistent social structures—household members interact repeatedly, individuals participate in multiple overlapping groups, and group compositions are not independent samples from a fixed distribution. These correlations could, in principle, carry additional information that breaks the moment equivalence.

The model also assumes that transmission occurs through a focal-center update rule where only one susceptible in a group can be infected per event. Other formulations allow multiple simultaneous infections within a group, which would change the structure of the Markov generator and potentially the moment hierarchy.

The higher-order channel that breaks the equivalence—adding a cubic term to a quadratic kernel—suggests one way forward. If real epidemics involve multiple transmission mechanisms operating simultaneously, each mechanism's order determines which moments become active. Observing a system transition between dynamical regimes as transmission channels activate or deactivate could, in principle, probe different moments. But this requires knowing the full transmission kernel, which is itself nonidentifiable from aggregate data.

Muslim's result could also be extended to other epidemic models. The analysis covers SIS and SIR dynamics, but similar reasoning might apply to SEIR models (with exposed but not yet infectious individuals), multi-strain competition, or network-based epidemic processes. The key is whether the infection hazard can be expanded in a finite factorial basis; if so, the moment sufficiency theorem suggests that a finite set of weighted moments will determine the aggregate dynamics.

The mathematical structure Muslim uncovered connects higher-order epidemic interactions to bifurcation theory in a new way. The moment hierarchy M₁, M₂, ... doesn't just organize what we can infer about contact patterns—it also organizes the phase structure of epidemic dynamics. The first moment controls the existence of endemic states; the second controls their stability and the nature of the transition. This suggests a systematic approach to classifying epidemic phase behavior based on moment constraints rather than detailed structural assumptions.

For public health, the message is both sobering and clarifying. We cannot reconstruct the full social architecture of transmission from epidemic data alone. But we can identify which features of that architecture actually matter for specific questions. If you want to predict invasion thresholds, know the mean group size. If you want to predict bistability, know the variance. The rest requires other data sources—social surveys, contact diaries, mobility data, genomic sequencing—each of which captures different aspects of human interaction that the epidemic dynamics alone cannot reveal.

The 2020 pandemic revealed how much we still don't understand about how respiratory diseases spread through populations. Muslim's theorem doesn't solve that problem, but it defines its boundaries. There are things the disease can tell us, and things it cannot. The moments tell their story; the rest remains hidden.

The dynamics contain information about the full distribution or only about a finite set of its moment combinations.

Comments (0)

No comments yet. Be the first to share your thoughts.