AI tools uncover long-standing mistakes in chemistry data and conference papers
An AI model flagged incorrect boiling-point values in a 75-year-old chemistry reference, and separate AI agents found numerous errors in recent machine-learning conference papers.
Sebastian Pios, a theoretical chemist at Zhejiang Lab in Hangzhou, China, employed an artificial-intelligence model to estimate boiling points and found values that conflicted with entries in a 75-year-old reference, later verifying that the source data were wrong. In parallel, SAI Labs used AI agents to review 168 submissions for the 2026 International Conference on Machine Learning, extracting claims, rerunning experiments where possible, and confirming only a limited number of them.
Of the 92 papers with enough assessable claims, the agents reproduced more than two claims for just 34 papers and succeeded on over 80% of claims for only eight papers. Experts such as Odd Erik Gundersen and James Zou note that AI fact-checkers can process vast amounts of information quickly but still make human-like errors, necessitating manual review. A separate study by Zou and colleagues scanning NeurIPS papers reported a rise in objective errors from 3.8 in 2021 to 5.9 in 2025, highlighting the risk of propagating flawed foundations in future research.
Why it matters
AI can rapidly spot hidden errors in scientific work, but human oversight remains essential to ensure research reliability.
In this story