Anthropic claims its latest model, Opus 5, is nearly immune to prompt injection attacks in its own software. Prompt injection, where an attacker manipulates inputs like hidden text on a webpage to bypass an AI model's instructions, fails against Opus 5 in almost every case. According to the system card, the attack success rate hit zero percent across 129 test scenarios. This is a significant development, especially since OpenAI admitted in December that prompt injection may never be fully solved. Opus 5 leads the Gray Swan IPI benchmark, with an attacker success rate of 2.0 percent after 15 attempts, followed by Mythos 5 at 2.6 percent and Fable 5 at 2.8 percent. | Image: Anthropic
The zero percent success rate only holds when Auto Mode is enabled in products like Claude Cowork. Auto Mode employs two defense layers: one that scans incoming data for hidden instructions before the model processes them, and another that blocks dangerous actions before execution. An attacker would need to bypass both layers independently. Without Auto Mode, Opus 5 has a 3.7 percent success rate, and Sonnet 5 performs better at 0.93 percent. Only the combination of the model and protective software pushes the rate to zero.
Source: thedecoder