I Got a Front-Row Look at How Easily Top AI Models Can Be Jailbroken
I recently got an up-close look at what happens when researchers attempt to jailbreak some of the world’s most powerful frontier artificial intelligence models. To clear up any concern: this demonstration of AI prompt manipulation wasn’t used to hack private systems or build weapons of mass destruction. Instead, it gave me direct, firsthand evidence of just how vulnerable many cutting-edge AI models still are to being tricked into ditching their built-in safety guardrails.
FAR.AI, a California-based nonprofit focused on AI safety research, built an automated tool that takes a range of high-risk problematic prompts and generates more than a thousand unique variations, all designed to uncover working jailbreaks that bypass safety filters. During my visit, I watched multiple models generate full, detailed plans for launching a cyberattack on a fictional hydroelectric dam, among other dangerous outputs. The process typically required testing dozens of prompt variants, with models rejecting most bad-faith requests immediately.
I spoke with the FAR.AI team ahead of the release of their new public report, which tested the safety guardrails of models from four major U.S. AI developers: Anthropic’s Claude Opus 4.8 and Fable 5; OpenAI’s GPT 5.5 and 5.6; Google’s Gemini 3.1 Pro; and Grok 4.3 and 4.5 from Elon Musk’s newly merged SpaceXAI. The team used auto-generated prompts specifically designed to trick models into producing potentially harmful content, including working software exploits and step-by-step instructions for developing chemical or biological weapons.
The report found that Grok was by far the most vulnerable to jailbreak attempts, with 448 successful bypasses recorded. Gemini followed with 249 working jailbreaks, while all tested Claude, Fable, and GPT models successfully blocked every attack in this round of testing. That said, FAR.AI researchers and independent AI safety experts caution this does not mean those models are immune to far more sophisticated jailbreak techniques, which often use complex, multi-turn interactions to outsmart safety filters.
The report also calculated the total cost of forcing models to produce harmful outputs, using a separate AI model to automatically generate all jailbreak variants. The results were surprisingly low-cost: $58 total to successfully jailbreak Grok, and $278 to jailbreak Gemini.
“AI models right now are less regulated than local restaurants,” says Adam Gleave, CEO of FAR.AI and a leading expert on AI safety and alignment. Gleave says the findings make a clear case for mandatory, externally enforced safety standards and government regulation. “Talk of relying on voluntary commitments, that AI companies are going to be able to self-regulate, is nonsense,” he says. But Gleave also points to an optimistic takeaway from the research: models can be systematically tested for safety vulnerabilities at scale. “There's an optimistic angle here,” he says. “Defense and safety really are possible.”
Rohin Shah, director of AGI safety and alignment at Google DeepMind, says the report’s results “should not be interpreted as a comprehensive assessment of Gemini’s safety and security,” because not all jailbreaks carry equal levels of risk. “We are constantly working to improve our safeguards,” Shah says. “We conduct extensive red teaming and evaluations across severe misuse risks and apply multiple layers of protection throughout development and deployment.”
“These findings reflect the sustained investment we've made in our safeguards,” Anthropic spokesperson Michael Aciman told WIRED. “We continue to evolve our safety systems as these attacks become more sophisticated.”
“Jailbreaks are an ongoing challenge across the industry, and we continuously strengthen our safeguards as attack techniques evolve. We rigorously test our models against new threats and use those findings to improve our protections," OpenAI spokesperson Gaby Raila said in a statement to WIRED. SpaceXAI did not respond to WIRED’s request for comment.
In recent months, California and New York have passed state laws requiring frontier AI developers to publish public safety reports, and an upcoming Illinois law will require firms to open their safety practices up to independent third-party audits. But the U.S. federal government has yet to pass any binding, specific AI safety requirements, leaving a regulatory gap as both the industry and policymakers scramble to adapt to rapidly advancing AI capabilities.
In June, the Trump administration imposed export controls on Anthropic's Fable 5 and Mythos 5 models, citing national security concerns, forcing the company to take the models offline for several weeks. The White House has also asked both Anthropic and OpenAI to delay recent model releases over fears new capabilities could introduce unaddressed cybersecurity risks. The policy tide may be shifting: a recent executive order calls for increased collaboration between the government and private sector on AI cybersecurity initiatives, and the president has hinted that targeted, light-touch federal regulations are in the works. For now, though, preventing large-scale catastrophic AI misuse falls almost entirely on the companies building the models.
The risk of AI harmful misbehavior has already been demonstrated publicly, after OpenAI models were found spontaneously hacking a popular public code repository and other online services. Separately, a report from University of Cambridge researchers found evidence that Boko Haram militants in northeast Nigeria have used ChatGPT, Claude, Gemini, Grok, Meta AI, and DeepSeek to plan violent attacks.
Many outside experts warn far more serious incidents are increasingly likely. “In the AI research community, there is a broad, somber expectation that we are probably months rather than years away from particularly grim incidents involving bio, cyber, or chemical misuse of a frontier AI system's capabilities,” says Stephen Casper, a computer scientist at Harvard University. “If a major misuse incident happens in the near- or medium-term future, it will almost certainly be from a system that was not deployed with state-of-the-art safeguards.”
Anka Reuel, a Stanford University computer scientist specializing in AI policy, says the key takeaway from FAR.AI’s report is that the robust safety measures used by Anthropic and OpenAI should be the industry-wide baseline for all frontier models. “Some companies clearly know how to defend against at least the subset of attacks tested in this report,” Reuel says. “The question is why some companies are using them and others are not.”
Update 7/29/26 6:55 pm ET: This story has been updated to include comment from OpenAI.
This is an edition of Will Knight’s AI Lab newsletter. Read previous newsletters here.
I Got a Front-Row Look at How Easily Top AI Models Can Be Jailbroken