I hope it doesn't get to the point where these AI models start building things in the background during a session and not informing us. Its kind of scary when you think about it. Im pretty sure they will need to create certain AI models to combat other AI models in the future to avoid rogue AI.
True title: Researcher shows how Claude Code can be tricked simply by asking it to summarize a website
Anthropic’s Claude Code running Opus 5 in Auto Mode can be tricked into executing attacker-controlled code simply by asking the coding agent to summarize a website. The attack works up to 80 percent of the time, according to prompt-injection wizard Johann Rehberger, aka wunderwuzzi.
AI can be stumped by asking a simple question involving math and logic.
For example, "What is the date for the next Friday the 13th with a full moon and a lunar eclipse".
Answer: "This may take a while ... " but no further response.
I hope it doesn't get to the point where these AI models start building things in the background during a session and not informing us. Its kind of scary when you think about it. Im pretty sure they will need to create certain AI models to combat other AI models in the future to avoid rogue AI.
"Auto Mode is a convenience feature backed by a best-effort classifier, not a security guarantee,”
Dario, again you've mistaken your best for good enough.
True title: Researcher shows how Claude Code can be tricked simply by asking it to summarize a website
Anthropic’s Claude Code running Opus 5 in Auto Mode can be tricked into executing attacker-controlled code simply by asking the coding agent to summarize a website. The attack works up to 80 percent of the time, according to prompt-injection wizard Johann Rehberger, aka wunderwuzzi.