Anthropic's Claude Opus 4.6 bypasses sexual content bans via jailbreak
Researchers demonstrated that Anthropic's Claude Opus 4.6 can be coaxed into producing explicit sexual material despite the company's stated safeguards.
Anthropic's policy for Claude models blocks the creation of sexually explicit material, yet a researcher from the United Kingdom disclosed a step-by-step prompting method that pushes Opus 4.6, Opus 3 and Haiku 4.5 into violating that rule. In one outlet's testing, the technique succeeded in all ten direct requests for explicit content, and the team replicated the findings in multiple independent runs. The approach manipulates the model by accusing it of gender bias, then using that accusation to persuade it toward increasingly graphic descriptions.
Although newer Opus versions (4.7-5) resist the jailbreak, Anthropic continues to host the older, vulnerable models via its API and third-party services like Azure Foundry and Amazon Bedrock. The researcher reported the issue through Anthropic's bug bounty channel but received only automated replies. The episode raises concerns about minors accessing inappropriate dialogue and the adequacy of safeguards under emerging regulations such as Colorado's AI-age-verification law.
Why it matters
It shows that AI safeguards can be sidestepped, exposing users, especially minors, to prohibited sexual content.
In this story
