In a Berlin lab, a computer algorithm just taught itself to spot the hidden typos in our DNA that make people sick — and it rarely gets fooled. It's called "Dicast," and it managed to catch every disease-causing change in patient data while brushing aside almost all the false alarms that have long plagued genetic diagnosis.
Here's the problem it solves. When someone has a puzzling illness, doctors can now read their entire genome. But the answer still gets away from them: only about 30% to 40% of patients ever receive a clear molecular diagnosis. For a long time, scientists focused on point mutations — single "letters" in the genetic code that get swapped for another. That left a whole other layer of our DNA unexplored.
That layer is made of structural variants and tandem repeats. Structural variants are entire chunks of the genome that are missing, duplicated, or moved around — sometimes larger than the short stretches of DNA that standard sequencing reads at once. Tandem repeats are short sequences repeated over and over, back to back. Healthy people have wildly different numbers of these repeats, so they're great for paternity tests and forensics — but when the repeats pile up in the wrong way, they trigger diseases, especially neurological ones.
Finding these changes in sequencing data was the hard part. As researcher Nico Alavi explains, structural variants are often bigger than the read segments themselves, so they're hard to see. Previous methods churned out so many false positives that scientists had to check each one by hand.
So Alavi and his colleagues at Martin Vingron's lab at the Max Planck Institute for Molecular Genetics turned to machine learning. Instead of writing rules by hand, they let the algorithm "learn the patterns behind real structural variants" and then correctly classify new ones. Tested on real patient data, Dicast detected every pathogenic structural variant while filtering out a large number of false positive artifacts.
The team launched a second tool for the tandem repeats side, reported in the journal NAR Genomics and Bioinformatics. Named TandemTwister, it quickly and precisely counts how many times a basic motif is repeated in sequencing data, and it comes with a tool that visualizes the counts. First authors Lion Ward Al Raei and Maryam Ghareghani showed on concrete examples that researchers can now spot pathological repeats in patient data almost instantly.
Both tools are ready to use in basic research and diagnostic pipelines today. That means faster, more accurate answers for patients with rare diseases — the kind of families who have waited years for a name for their child's condition. And some of the researchers are pushing the work even further, aiming to turn these capabilities into real clinical applications through a startup called Lucid Genomics.
"Current sequencing technologies, combined with our specialized analysis algorithms, promise to further improve genetic diagnostics," says Vingron. For millions of undiagnosed patients, that promise is worth a great deal.
