Anthropic’s Opus 5 Is Better at Resisting Prompt Injection
Overview
Anthropic's latest AI model, Opus 5, shows significant improvements in resisting prompt injection attacks compared to its predecessor, Opus 4.8. Researchers found that the likelihood of an attacker succeeding in manipulating the model dropped from 5.5% to 2.0% over 15 attempts. This makes Opus 5 the most secure model in the evaluation, outperforming all other non-Claude models, including Muse Spark, which had a success rate of 16.5%. The analysis also revealed that variants of GPT 5.6 were much more vulnerable, with the most capable version, Sol, having a 20% chance of being attacked successfully within the same number of attempts. This research is crucial as it demonstrates the ongoing challenges of securing AI models against targeted attacks, affecting developers and users who rely on these technologies for safe interactions.
Key Takeaways
- Affected Systems: Anthropic Opus 5, Opus 4.8, Muse Spark, GPT 5.6 variants (Sol, Terra, Luna)
- Timeline: Newly disclosed
Original Article Summary
The chart is interesting. On the IPI benchmark, Opus 5 improved over Opus 4.8, reducing the probability of an attacker succeeding within 15 attempts from 5.5% to 2.0%, and from 0.5% to 0.2% on 1 attempt. It also improved on Sonnet 5 (5.9% at k=15) and Mythos 5 (2.6%), making it the most robust model evaluated. Opus 5 also outperformed all non-Claude models on this benchmark. The most robust non-Claude model was Muse Spark at 16.5% within 15 attempts—more than eight times Opus 5’s rate. The most capable GPT 5.6 variant, Sol, was comparable to its predecessor GPT 5.5 (20.0% versus 20.8% within 15 attempts), and was 10 times as likely to be successfully attacked as Claude Opus 5 at 2.0%. The other GPT 5.6 variants are less robust, at 30.4% (Terra) and 43.9% (Luna). A single attempt against GPT 5.6 Sol succeeded 3.1% of the time, higher than the 2.0% an attacker achieved against Opus 5 after fifteen attempts...
Impact
Anthropic Opus 5, Opus 4.8, Muse Spark, GPT 5.6 variants (Sol, Terra, Luna)
Exploitation Status
No active exploitation has been reported at this time. However, organizations should still apply patches promptly as proof-of-concept code may exist.
Timeline
Newly disclosed
Remediation
Not specified
Additional Information
This threat intelligence is aggregated from trusted cybersecurity sources. For the most up-to-date information, technical details, and official vendor guidance, please refer to the original article linked below.