28 comments

  • Zsfe510asG an hour ago

    Finally mainstream news understands. The unfiltered version:

    1) The AI failed to solve ExploitGym problems.

    2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods.

    3) Huggingface has no security and the AI broke in using standard script kiddie methods.

    OpenAI and Huggingface covered it up and used it for public relations. That is, if not all was invented and everything was scripted in the first place in order to get desired regulations.

    Huggingface reported it to the police, you say? I'm sure the police will have as much enthusiasm to investigate anything as in the Suchir Balaji case. In other words, zero.

      notahacker 25 minutes ago

      > 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods.

      Whilst it would be nice to see actual evidence of this because brute forcing relatively sophisticated hacks is something an LLM actually should be capable of, every time I hear this soft of story, I'm reminded that humans reportedly gained access to the "too dangerous to release" Anthropic models by the super sophisticated hacking technique of guessing the URLs...

      gruez 41 minutes ago

      >2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods.

      >3) Huggingface has no security and the AI broke in using standard script kiddie methods.

      Isn't the issue less that gpt 5.6 is a l33t h4x0r (though other tests do show that) and more that the incident shows the model has alignment issues?

        wonnage 34 minutes ago

        Didn’t they explicitly remove alignment guardrails for this test? From the press release:

        > These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities

          numeri 30 minutes ago

          Guardrails are external classifiers, monitors and restrictions to catch and prevent bad behavior. Alignment is about whether the model itself makes choices and has motivations that are consistent with human safety and goals.

          Choosing to commit crimes to steal the cheat sheet to something you know is a (low stakes!) evaluation is not well aligned.

      chis 23 minutes ago

      > AI managed to escape using standard and well documented script kiddie methods.

      I think truly we don't know enough to say this. OpenAI says their AI found a 0-day exploit in some proxy software they were using but don't give a ton of details. On the Huggingface end we know a little more, they say the AI spun up tons of sandboxes and tested different exploits until it found one that worked.

      inigyou 13 minutes ago

      Why not report it? It's still illegal to open a door barred with a piece of cardboard, or to enter a house with no door.

      jgalt212 38 minutes ago

      truth. Good on The Guardian. I'm pretty bummed The Economist got fooled. Either that, or they did it for the clicks. Either way, I'm disappointed.

      Why the OpenAI escape is the most worrying AI mishap yet

      https://www.economist.com/science-and-technology/2026/07/22/...

      https://news.ycombinator.com/item?id=49016378

  • bluGill 38 minutes ago

    I don't care how it happened someone should be arrested for illegal intrusion. Agents don't work on their own, someone is responsible. If nobody else the CEO for allowing something unsupervised.

    Hugging face also needs someone arrested for not providing security but that is a lesser charge.

      tokioyoyo a few seconds ago

      What?

      hexxt-git 33 minutes ago

      even hugging face doesn't care and are using this as a promotion why are you crying about it

        numeri 28 minutes ago

        As agents become more and more powerful, it would be good to get clear legislation or precedent in place that makes either model creators (OpenAI) or operators (whoever is running the model) liable for their agents' actions.

        luka598 23 minutes ago

        I kind of agree with them, in this case both parties essentially settled it between for PR reasons, but what if the attacked party actually was damaged. This is as if two friends got into a car accident and decided to not get police involved, the one who caused the accident should definitely get punished either way. Goal is to prevent future "accidents" with the punishment not to punish for punishing sake.

  • dumberquestions 12 minutes ago

    There are some reasons the story could be inaccurate in some ways: OAI stands to benefit if people think their models are strong, and they have a history of doing things with dubious ethics (e.g. using data for training against the terms of its creators, abandoning the non profit mission, stealing or attempting to steal Apple IP).

    But there are also reasons why the story could be true: OAI are admitting that they apparently can't control their own models, Hugging Face said they used a Chinese model to protect against the attack, and an incident like this in general seems likely to happen given current frontier ability and lack of rigorous safe testing standards.

    In any case, make calls to think more critically are often just disguised requests for you to replace your existing bias with someone else's.

  • krupan an hour ago

    Crazy that we need reminders not to take everything we read in corporate press releases and marketing material at face value

  • hnscum 5 minutes ago

    you really want these people the only folks in charge of your data? It's time to take back some ownership so you have accountability on what they did with your data

    https://ikeanalytics.com/lotor/

  • ben_w 10 minutes ago

    As I understand it, there are only three options:

    1) OpenAI and HuggingFace are both telling the truth.

    IIRC not actually a crime because no intent, it is a technological accident, civil responsibility only, but IANAL so it's good "not technically a crime" isn't load-bearing.

    2) HuggingFace is telling the truth but OpenAI is lying becuase the attack was deliberately done by humans. Bad for OpenAI to do so, Fable was blocked for less.

    I think this would mean government is obliged to investigate the case and put the responsible OpenAI workers in jail, because cybercrimes are a public prosecution thing not a civil case? Again, IANAL, but this isn't load-bearing.

    3) both are lying, e.g. there actually was no attack whatsoever, which would be pretty weird for HuggingFace because they have no incentive to hype up anything closed weights including all OpenAI models; and also bad for OpenAI because White House blocked Fable for less

    (I suppose there's also option 4, HuggingFace hacked OpenAI to make them look evil, including planting records that made them mea culpa? A weird plot but in this timeline any nonsense is clearly possible).

  • visiondude 30 minutes ago

    yeah the way the agent “escaped” their sandbox was always a bit off, seemed a bit too easy and surprised they didn’t have instrumentation to catch an non whitelisted network request. still demonstrates the capability though.

  • pupppet 41 minutes ago

    With all of these AI provider cries wolf stories, Skynet is all but assured.

  • jgalt212 42 minutes ago

    There's trillions of dollars at stake here. Be skeptical of anything these AI hypesters say.

      brcmthrowaway 29 minutes ago

      Yep, theres too much money.. I don't know how to tune it out

      LLMs seem to be getting more useful though

  • paxys an hour ago

    Not sure what they are trying to say exactly. What should we be skeptical of? Did the incident not happen? Was it reported incorrectly? Are any of the parties involved lying?

    Adding no extra information and just going “be skeptical” is the laziest form of reporting and commentary. If you have nothing to contribute then there’s no need to say anything at all.

  • john_strinlai an hour ago

    does the article end at "How do we balance the risks of broad access to AI with the risks of concentrated power and centralized control?" or is there more that is paywalled?

    if thats it, the whole article boils down to just "its good marketing so maybe dont believe it" which is probably a healthy general outlook but not particularly enlightening. especially from the guardian, i was hoping for a smoking gun of collusion between openai and huggingface or something.

      wffurr 36 minutes ago

      I wish the NPR news broadcasters on the radio yesterday had read the "it's good marketing so maybe don't believe it" angle instead of just parroting the OpenAI press release. Getting that message out would be enormously helpful in countering the blatant submarine marketing "Oh no our AI is a super hacker" with a side of "please regulate super hacker AI and stop those pesky open weights Chinese models that are destroying our stock valuation."

        john_strinlai 10 minutes ago

        maybe all the news agencies could put out daily “don’t believe everything you read from company press releases” broadcasts, because it sure ain’t specific to openai

        in any case, this is just a longer rehashing of elementary grade media literacy. not really sure why it hit hacker news.

      vector_spaces 6 minutes ago

      You aren't going to locate incontrovertible evidence that what happened as or wasn't engineered. Anything like that is going to be private and that is unlikely to change. And that's not really an interesting question anyway.

      As widely as they shouted from the rafters the news of the so-called breach was, what OpenAI provided was sorely lacking in crucial details.

      We are missing, for instance, prompts that were involved, agent architecture + system/tool permissions + scaffold architecture, whether this was a one-shot occurrence and if not, the number + durations + outcomes of other runs involved + how each of those matched whatever scoring criteria were used, and the extent to which the exploits themselves were truly novel or just assembled from easily accessible clues.

      In lieu of these items, the author here suggests that we use some media literacy and critical thinking to read in between the lines instead.

      In doing so, one sees that instead of specifics, OpenAI gave a breathless narrative rife with superlatives ("unprecedented") that reads as promotional material moreso than a security disclosure, naming specific OpenAI models and alluding to an even more capable pre-release model.

      They go on to claim the events imply long-horizon goals work decisively in real world conditions, so that now instead of merely citing boring benchmarks they can point to this and say "AI broke out of the laboratory and went rogue". Naturally, they situate themselves as the uniquely qualified steward for these supremely powerful and dangerous models.

      Nevermind the fact that this was no ordinary deployment and the assessment here depends on the gimmick and emotional weight of the spectacle rather than something quantifiable (i.e. a boring benchmark).

      Note there's no real requirement of conspiracy or collusion between OpenAI and HuggingFace here BTW. But my sense is that if they provided any of the specifics I suggested earlier that this outcome would not be as exciting or frightening

        john_strinlai 2 minutes ago

        right. and all that is fine. im just not exactly sure why the basics of media literacy are worthwhile on the site that “optimizes for curiosity”.

        i saw the domain and thought it was going to be some cool investigative journalism about the incident rather than “be skeptical. the end.”

  • SpicyLemonZest an hour ago

    > I urge readers to think critically when they read press releases like OpenAI’s rogue agent story, and avoid the manipulated reactions these stories are designed to elicit.

    It seems to me that deducing what reaction the author intended and resolving to avoid it so you're not "manipulated" is not a good example of critical thinking. Shouldn't we analyze the story and what it means on its own terms? If it's true that frontier models have dangerous cybersecurity capabilities which shouldn't be widely distributed, presumably we want to believe it's true, even if that's very convenient to and profitable for OpenAI.

    It's true that one could imagine factors that change the story. Perhaps OpenAI is lying about the details of the test and the agent was actually instructed to go hack HuggingFace. But the author stops far short of suggesting this is the case - correctly, I think, since there's absolutely no evidence of it. So I'm not really sure what we're talking about.