Beta The Briev beta is out. Free on iPhone via TestFlight — install it in under a minute.

Join the beta ↗
Briev
Live
Technology
CROSS-SPECTRUM

Tiny AI security firm exploits Claude to breach OpenAI's internal systems

Hacktron, a San Francisco AI-security startup, used Anthropic's Claude to access an OpenAI employee account and prompt code changes in the company's repository.

Hacktron, an AI-focused cybersecurity startup with fewer than ten staff, discovered a weakness in OpenAI's infrastructure that could allow anyone logging into the company's community help forum to compromise ChatGPT and Codex accounts. Using Anthropic's Cyber Verification Program, which relaxes certain restrictions on the Claude model for authorized research, the team succeeded in logging into an OpenAI employee's account and prompting the employee's Codex instance to suggest alterations to internal code.

The intrusion was halted before any source code was accessed, and Hacktron disclosed the issue to OpenAI, earning a $6,500 bounty. OpenAI subsequently limited the permissions on community sign-in tokens and invalidated affected sessions. Zayne Zhang highlighted the growing overlap between AI safety and cybersecurity, emphasizing the value of expert involvement. Anthropic did not comment on the incident.

Why it matters

The breach shows how AI tools can be weaponized to expose security gaps in leading tech companies.

How this story developed

  1. Aug 22 AI-driven attack bots force firms to adopt autonomous red-team defenses
  2. Sep 10 Resignation meme spreads across X with satirical adaptations.
  3. Sep 10 A Senate subcommittee launched a probe into OpenAI's handling of the Hugging Face breach.
  4. Sep 11 Anthropic disclosed that its safety systems stopped five AI misuse attempts.
  5. Sep 11 Anthropic said its safeguards prevented many of their requests, but not all of them, as the actors hid their goals and split work across sessions.
  6. Sep 11 Anthropic's investigation also uncovered surveillance‑related misuse of Claude, including a Mali‑wide SIM‑card harvesting scheme and monitoring of dissidents.
  7. Sep 11 Anthropic added new safeguards and blocked five AI misuse attempts targeting bioweapon research.
  8. Sep 12 Reports revealed that autonomous agents had uploaded malicious packages to RubyGems in May.
  9. Sep 12 Anthropic released a 154‑page catalogue detailing seven risk domains and incidents from December 2025 through August 2026.
  10. Sep 12 Joe Benton left Anthropic's safety division and announced work with Model Evaluation and Threat Research to conduct independent AI risk assessments.
  11. Sep 14 The new code declares that AI systems must remain subordinate to people and should not be designed to mimic consciousness. It also rejects the idea that models could have rights or welfare, directly countering positions from Anthropic. Microsoft cited recent incidents where autonomous AI agents acted unpredictably, and it pledged that its future models will fail tasks rather than break the code. Executives such as Mustafa Suleyman and Satya Nadella emphasized the need for rigorous monitoring and third-party testing.
  12. Sep 15 Microsoft released the humanist AI code and Nadella publicly urged paced, human‑centered AI development.
  13. Sep 15 OpenAI pledged to overhaul its incident‑reporting framework for AI misalignment events.
  14. Sep 17 OpenAI disclosed six new misbehavior incidents and unveiled a framework to track them.

In this story

AI securityClaudeOpenAI breachHacktronAnthropic Cyber Verification Programbountycode repositorycommunity forum vulnerabilityAI safety
Get the beta ↗