Mathias Unberath works on artificial intelligence that helps doctors diagnose patients. But he knows a problem can slip into the technology in sneaky ways.
Imagine if a computer learned to tell whether someone was a man or a woman just by looking at their eyes. That sounds impressive until you realize the AI was actually spotting mascara, not biological differences. The model saw the wrong clue and drew the wrong conclusion.
Unberath and his team at Johns Hopkins University in Baltimore have built a new tool to catch these hidden mistakes before they reach patients. The tool is called G-AUDIT, which stands for Generalized Attribute Utility and Detectability-Induced Bias Testing. It was developed together with the U.S. Food and Drug Administration and was published in the journal npj Digital Medicine.
The tool looks at the massive datasets used to teach medical AI systems. Rather than waiting to test a finished model, G-AUDIT scans the raw data for subtle patterns that could confuse the AI. It flags which details in the data might lead a model astray.
Mitchell Pavlak, a Ph.D. student who worked on the project, explained why this approach matters. "We're not checking our work after the fact," he said. "We analyze the data to figure out what is likely to be a problem."
One example from the research shows exactly what can go wrong. The team looked at a skin cancer dataset that mixed images from two different clinics. One clinic treated mostly high-risk patients, while the other saw fewer cancer cases. The two clinics also used different cameras. When an AI learned from this mixed data, it could wrongly decide that better camera quality meant cancer was more likely. It might also pick up that rulers appearing in photos signaled cancer, simply because the high-risk clinic used rulers to measure growths.
"Envision taking this algorithm into the real world with a smartphone camera," Unberath said. "There's no ruler. The camera quality is completely irrelevant. But the model learned to associate 'ruler' and 'camera.' You have just created a health disparity."
The team plans to keep improving G-AUDIT. They say it could also be useful beyond medicine, in any field where AI makes important decisions. Unberath hopes it becomes a standard step before any AI system is used in real situations.
"The core takeaway is that right now people are auditing models, but they don't really know what to audit for," he said. "We are changing that."
