← News
Science Breakthroughs Science Breakthroughs Knowledge

The Fairness Blink Spot: Why Equity Tools in Science Miss So Many Women

The Fairness Blink Spot: Why Equity Tools in Science Miss So Many Women
2,250 Preregistered experiments participants
Citation Diversity Statements Papers studied

Chunhua Li. That's the name of a physicist who has published dozens of papers in high-energy physics. To an English-speaking reader, it gives away almost nothing. Is Chunhua a woman or a man? The given name could be either. Now try Sarah Johnson. Same science, same track record—and the name does the work of self-identification for her in a single glance.

This one small difference in how a name travels across language is doing quiet, structural damage to the very systems built to fix gender inequity in science. New research from a team led by Binglu Wang, Jose Cervantez, Jiahui Xue, Katherine Milkman, and Dashun Wang identifies what they call the "legibility gap"—the systematic way that equity interventions, when they rely on inference from names, reward women whose names signal gender in English while bypassing women whose names lose those cues in transliteration. It is, in essence, a fairness blind spot hiding inside fairness itself.

The consequences are concrete and measurable. Papers that adopt a new practice called citation diversity statements—where authors disclose the algorithmically estimated gender mix of their reference lists—do cite women more often. But the researchers found that this gain accrues almost entirely to women with gender-signaling Western names. Women whose names lose their gender cues in English transliteration, predominantly East Asian women, see fewer citations in these very same papers. The tool designed to boost representation is quietly redistributing recognition toward the women who are already easiest to recognize.

The Science

The researchers began with a suspicion that the machinery of modern equity work has a built-in cultural bias. When funders, journals, and departments want to measure gender representation, they rarely ask people directly. Instead, they feed name lists through algorithms that infer gender from a name's statistical association—"Sarah" as female, "Michael" as male. These tools are trained overwhelmingly on Western name data, and they work perfectly well for names that already carry clear gender signals in English.

But science is global. Names arrive in the English-language ecosystem already transformed. An East Asian name written in Chinese, Korean, or Japanese script carries its own cues about the person who bears it, including sometimes the gendered characters that make up a given name. When that name is transliterated into the Latin alphabet for an English-language paper, those cues frequently vanish. "Li" gives away nothing about gender. Neither does "Wang," "Chen," or "Kim." The name becomes, in the authors' term, illegible.

The study proceeds in two parts. First, the observational analysis. The team looked at citation diversity statements, an emerging norm in which authors run their reference lists through gender-inference algorithms and disclose the estimated composition. The researchers assembled a dataset of papers that included these statements and compared their citation behavior with matched papers that did not. They then broke down the cited women by whether their names would be legible as female in English.

The second part is experimental. The researchers ran two preregistered experiments with a combined sample of 2,250 participants. The design was built to isolate the mechanism: is it linguistic legibility that drives who gets recognized as a woman, or is it a broader cultural unfamiliarity with East Asian names? By carefully controlling what subjects saw, the experiments could separate these two explanations.

What They Found

The headline result is stark. Papers that include citation diversity statements cite more women overall—but the increase is not shared across all women. The gains flow almost entirely to women with gender-signaling Western names. Women with names that lose gender cues upon transliteration receive fewer citations in these same papers, even as the stated intention of the practice is precisely to boost underrepresented scholars.

The experiments pin down the mechanism. The researchers found that it is not cultural unfamiliarity driving the effect—it is legibility itself. When participants were given the same East Asian names but the gender cue was rendered legible, recognition of those people as women rose. When the cue was present but the name came from an unfamiliar culture, recognition still worked. The binding constraint was whether the name itself could communicate gender in the linguistic context being used. Legibility, not familiarity, determined who was counted as a woman—and therefore who benefited from the equity intervention.

This is the crucial nuance that changes how we should think about the problem. It would be easy to assume the gap is a simple prejudice against non-Western names. The data suggest something more subtle and more insidious: the algorithmic infrastructure of equity work is legible to some names and illegible to others, and it distributes recognition accordingly.

The authors frame this as a failure of global equity infrastructure. The tools we built to measure fairness were themselves built on a partial view of the world. They recognize the women they were trained to recognize. The women their training neglected do not merely go unhelped—in this accounting, they are simply invisible, and the interventions pass over them entirely.

Why This Changes Things

The deeper problem the researchers identify is that averages hide distributions. On average, citation diversity statements appear to work: authors who adopt them cite more women. A journal editor glancing at aggregate numbers might conclude the policy is functioning well. But the aggregate conceals a redistribution of recognition away from the very women who are hardest to see, and the "success" of the policy is built on a form of preferential recognition for those whose names already cooperate.

This illuminates a general principle that extends far beyond citation statements. Any equity intervention that relies on name-based inference—diversity metrics for hiring, authorship attribution, grant portfolio analysis, even the algorithms behind "gender balance" dashboards—inherits the same blind spot. The legibility gap is not a feature of one practice but of a whole class of algorithmic equity infrastructures being deployed across science and beyond.

The authors are careful to note the direction of the effect is not simply benign neglect. It is an active redistribution: women whose names are illegible receive fewer citations specifically in papers that adopt the diversity statements. The policy does not just fail to help them; within the ecosystem of such papers, they are recognized less. Equity infrastructure, in other words, does not merely reproduce existing inequity—it can actively concentrate the benefits of its own interventions onto an already-privileged subset.

There is a broader resonance here for anyone who has watched the rise of algorithmic fairness tools in hiring, lending, or education. The lesson is that a fairness tool is only as fair as the model of the world it carries inside it. When that model is built on data that overrepresents one culture, the tool exports that culture's assumptions to every context in which it is deployed. The legibility gap is a specific name for a general phenomenon: the way our tools of recognition encode the biases of their training data and then certify those biases as objective measurement.

What's Next

The authors do not leave us without directions forward. Their findings imply a set of practical fixes, starting with the design of gender-inference tools themselves. If the binding constraint is legibility, then the solution is to build recognition that does not depend on transliterated names carrying gender cues in English. That could mean multilingual inference systems, handling names in their original script, or combining name data with biographical signals. It could also mean moving away from name-based inference altogether where direct self-identification is possible.

There is also a design principle for the equity interventions themselves. Rather than optimizing for average improvement in recognition, the authors argue, we need to attend to whether policies are equitable across cultures—distributing their benefits where the need is greatest rather than where recognition is easiest. The abstract puts it succinctly: in global systems of recognition, equity depends not only on whether policies are effective on average, but on whether they are equitable across cultures.

Several open questions remain. How widespread is the legibility gap across disciplines and regions where name-transliteration practices differ? Can existing gender-inference algorithms be retrained to close the gap, or will the bias be intractable at the level of the training data? And crucially, how should journals and funders weigh the decision to adopt citation diversity statements given the documented redistribution they induce until the tools improve?

What makes this finding matter is not that it reveals a malicious actor—no one is deliberately excluding East Asian women from citation diversity benefits. It matters because the inequity is structural, embedded in the mundane machinery of recognition that well-meaning people deploy every day. The legibility gap is a reminder that fairness is not a switch you flip but a system you must continuously inspect from every culture's vantage point. As science becomes more global and equity efforts become more algorithmic, the question of who our tools can see is not academic. It is the quiet filter determining whose work gets counted, whose contributions get recognized, and whose names—however excellent the science behind them—never quite register as belonging to a woman at all.