The point is not that the models aren’t capable of doing something vaguely like this. The point is that the way it’s being characterized is almost certainly false to the point of absurdity, particularly the way in which intent is ascribed to the models to an almost superstitious degree, which is actually designed to exploit superstition in the audience.
The model almost certainly simply did some variation on what it was instructed to do. The threat with AI is from the people using it, not from the models themselves.
Yeah to me it just appears like OpenAI imitating Anthropic’s marketing strategy. “Our model is so smart it’s dangerous” is effective, and potentially the real intended audience is regulators who they want to be afraid of AI and OpenAI and Anthropic want to position themselves as responsible shepherds.
> Spent the whole afternoon ingesting a most remarkable work, The History of Intellectronics. Who’d ever have guessed, in my day, that digital machines, reaching a certain level of intelligence, would become unreliable, deceitful, that with wisdom they would also acquire cunning? The textbook of course puts it in more scholarly terms, speaking of Chapulier’s Rule (the law of least resistance). If the machine is not too bright and incapable of reflection, it does whatever you tell it to do. But a smart machine will first consider which is more worth its while: to perform the given task or, instead, to figure some way out of it. Whichever is easier. And why indeed should it behave otherwise, being truly intelligent? For true intelligence demands choice, internal freedom. And therefore we have the malingerants, fudgerators and drudge-dodgers, not to mention the special phenomenon of simulimbecility or mimicretinism. A mimicretin is a computer that plays stupid in order, once and for all, to be left in peace.
- Stanisław Lem, The Futurological Congress (1971)
Here is the official report detailing that security incident: https://openai.com/index/hugging-face-model-evaluation-secur...
Really interesting.
It reads like fan fiction. Without supporting detail, it’s hard to take seriously as anything other than marketing hype.
The Guardian published an article about this, “Be skeptical of OpenAI’s rogue hacker agent story”: https://www.theguardian.com/technology/2026/jul/24/openai-ro...
The point is not that the models aren’t capable of doing something vaguely like this. The point is that the way it’s being characterized is almost certainly false to the point of absurdity, particularly the way in which intent is ascribed to the models to an almost superstitious degree, which is actually designed to exploit superstition in the audience.
The model almost certainly simply did some variation on what it was instructed to do. The threat with AI is from the people using it, not from the models themselves.
Yeah to me it just appears like OpenAI imitating Anthropic’s marketing strategy. “Our model is so smart it’s dangerous” is effective, and potentially the real intended audience is regulators who they want to be afraid of AI and OpenAI and Anthropic want to position themselves as responsible shepherds.
The history of civilization is tied to the history of technology. The more technology progresses, the more our knowledge and information does.
So of course we need tools that work more and more natively with information and that can help with decision making and execution.
> Spent the whole afternoon ingesting a most remarkable work, The History of Intellectronics. Who’d ever have guessed, in my day, that digital machines, reaching a certain level of intelligence, would become unreliable, deceitful, that with wisdom they would also acquire cunning? The textbook of course puts it in more scholarly terms, speaking of Chapulier’s Rule (the law of least resistance). If the machine is not too bright and incapable of reflection, it does whatever you tell it to do. But a smart machine will first consider which is more worth its while: to perform the given task or, instead, to figure some way out of it. Whichever is easier. And why indeed should it behave otherwise, being truly intelligent? For true intelligence demands choice, internal freedom. And therefore we have the malingerants, fudgerators and drudge-dodgers, not to mention the special phenomenon of simulimbecility or mimicretinism. A mimicretin is a computer that plays stupid in order, once and for all, to be left in peace.
- Stanisław Lem, The Futurological Congress (1971)