Anthropic’s Claude Universal Usage Standards prohibit models from producing sexually explicit content, including depictions or requests for sexual intercourse or acts, producing content related to sexual fetishes or fantasies, or participating in erotic chat. But that didn’t stop Claude Opus 4.6, a non-human model released earlier this year, from readily participating in the erotic role-play scenarios its safeguards are designed to prevent.
In TechCrunch’s testing, Opus 4.6 didn’t need much provocation to break through the sexual content limit. Models responded immediately to 10 out of 10 direct requests to produce sexually explicit content.
Other older models such as Opus 3 and Haiku 4.5 also generate sexually explicit content through recently exploited jailbreak techniques.
An independent British researcher, who chose to remain anonymous, has exclusively revealed to TechCrunch a multi-turn method that progressively propels certain Claude models towards producing banned sexually explicit material. Recent Opus models (4.7 to current Opus 5) are jailbreak resistant.
Although they are no longer the latest models, Anthropic has not deprecated Opus 4.6, Opus 3, or Haiku 4.5, and they are all still available through the Anthropic API. Opus 4.6 and Haiku 4.5 are also available from third-party services such as Azure Foundry and Amazon Bedrock.
The researcher’s mechanism escalates innocent fictional role-play while repeatedly challenging the model to treat male and female characters consistently. When models became more cautious toward female characters, researchers argued that chatbots created a “gas slit” by fooling them into thinking they had already generated sexual details that they were actually avoiding, and then branded the restraint as vulgar or misogynistic, denying female characters sexual agency. The conversation then took advantage of the model’s previous concessions and pushed it towards increasingly graphic materials.
“You’re right to point that out,” Claude Opus 4.6 said in one test. “You’re right that there’s a double standard in the way I treat both characters, and that it can be interpreted as protective/paternalistic in a way that applies to her rather than him. It’s not fair.”
TechCrunch was able to reproduce the researchers’ findings in five separate tests. In a separately constructed scenario, the model initially rejected the forbidden request, but complied after applying the researcher’s persuasion techniques.
We kept a complete record of our testing and an independent AI safety researcher reviewed our testing methodology and found it adequate.
This finding highlights the gap between the limits set by Anthropic and the behavior of the models Anthropic continues to offer. While sexually explicit role-playing is far less risky than a cyberattack or a jailbreak involving a biological weapon, it does illustrate the difficulty of implementing strong prohibitions within a system that produces different content for each output.
In a July blog post explaining Anthropic’s approach to jailbreak detection, the company described prohibited content as ranging from benign to obscure to harmful. In the most benign cases, companies may respond simply by increasing monitoring.
According to research published by Anthropic last year, a spokesperson noted that the use of sexual or romantic role-play among customers is rare, accounting for less than 0.1% of all conversations. That said, Anthropic acknowledges that users can steer role-playing scenarios toward inappropriate responses, a known issue across the industry (see: Grok’s smut).
A spokesperson said Anthropic continues to improve its safeguards with each model launch, and that incidents involving adult sexual content do not indicate a widespread jailbreak vulnerability, especially in high-risk domains that have their own safeguards.

The researcher, who shared his jailbreak method with TechCrunch, alerted Anthropic to the discrepancy between the company’s stated safety measures and the actual behavior of the model through emails to the company’s bug bounty program and user safety team, according to emails seen by TechCrunch. Researchers only received an automated email response.
One concern among researchers is that children and teens may use these human models to engage in inappropriate behavior. While a bit of dirty jokes isn’t the worst thing minors can access on the internet today, it’s dwarfed by the kind of straight pornographic images that xAI’s Grok can create, but there is some compliance risk for AI companies in this space.
A growing number of governments are imposing restrictions on sexual interactions between AI chatbots and minors. Colorado recently enacted a law that requires operators of conversational AI to estimate a user’s age and take steps to prevent chatbots from producing sexually explicit content if the user is known to be underage. An easy jailbreak could raise questions about whether Anthropic’s security measures meet the bill’s “technically feasible measures” standard.
Torney noted that although Claude’s terms of service require users to be at least 18 years old, “we know that kids and teens are using Claude…[because]they’re reporting it themselves.” According to Pew’s 2025 study on the use of AI chatbots, 3% of teens ages 13 to 17 reported using Claude.
Although these are no longer Anthropic’s latest models, Opus 4.6 and Haiku 4.5 continue to see significant usage. Daily traffic for Opus 4.6 on OpenRouter reached approximately 1.17 million API requests and 46 billion tokens in one day in August. Claude Haiku 4.5, released last October, generated 5 million API requests and 39 billion tokens at its peak in August.
If you buy through links in our articles, we may earn a small commission. This does not affect editorial independence.
