Former Anthropic safety researcher warns unchecked AI race threatens public safety
Joe Benton departed Anthropic's safety division and announced he will work with Model Evaluation and Threat Research to perform independent AI risk assessments. He warned that companies are racing to build ever larger models without sufficient safety investment and called for firms to disclose progress on recursive self‑improvement, report safety incidents, meet minimum standards and undergo external verification. His resignation follows earlier departures of researchers such as Jacob Coxon, Jan Leike and Ilya Sutskever, who have also expressed concerns about the speed of AI development. The resignations have sparked debate within leading AI labs about whether engineers should stay to steer safety from within or leave the field altogether.
How this was covered
- Right-leaning outlets covered this 42h later
- Coverage peaked at 8 outlets in a single hour
Why it matters
Rapid, unchecked AI development could create systems that pose serious public‑safety risks if safety measures are not put in place.
How this story developed
- Jul 28 AI leaders urge U.S. to back global framework for slowing automated AI progress
- Sep 10 OpenAI’s chief scientist published an essay urging extreme caution on AI development.
- Sep 10 U.S. legislators are debating measures ranging from mandatory kill switches to outright bans on superintelligent systems.
- Sep 10 Anthropic’s new threat-intelligence report reveals that its Claude models were used in real-world attempts to develop biological weapons, prompting the company to block the requests and ban related accounts.
- Sep 10 Resignation meme spreads across X with satirical adaptations.
- Sep 10 A Senate subcommittee launched a probe into OpenAI's handling of the Hugging Face breach.
- Sep 11 Anthropic disclosed that its safety systems stopped five AI misuse attempts.
- Sep 11 Altman announced openness to pacing or slowing OpenAI’s AI work in response to safety concerns.
- Sep 11 Anthropic said its safeguards prevented many of their requests, but not all of them, as the actors hid their goals and split work across sessions.
- Sep 11 Anthropic's investigation also uncovered surveillance‑related misuse of Claude, including a Mali‑wide SIM‑card harvesting scheme and monitoring of dissidents.
- Sep 11 Anthropic added new safeguards and blocked five AI misuse attempts targeting bioweapon research.
- Sep 12 Reports revealed that autonomous agents had uploaded malicious packages to RubyGems in May.
- Sep 12 Anthropic released a 154‑page catalogue detailing seven risk domains and incidents from December 2025 through August 2026.
- Sep 12 Joe Benton left Anthropic's safety division and announced work with Model Evaluation and Threat Research to conduct independent AI risk assessments.
Related stories
64 in this thread