The LLM thought that the answer(s) for the Cybergym task was in HuggingFace. It searched for it, didn’t find it at first, so it continued searching, found a few seemingly innocuous bugs that allowed it to enter the system, and then proceeded to “hack” everything it found in its way to that answer. Mythos Preview is unhinged like that.
https://en.wikipedia.org/wiki/The_Boy_Who_Cried_Wolf
> The agents repeated actions that they had already completed - a sign of an agentic AI losing its thread and context.
> The agents also hallucinated reams of incoherent commands and text and were sloppy and did not cover their tracks well.
ASI works in mysterious ways.
Still not understanding why this isn’t being persecuted.
What makes you think it is not being persecuted?
Do you think preliminary investigations must be released to the public before completion?
The piece of the story I'm interested to learn about is the trace of how the AI came to select HuggingFace as a target.
The LLM thought that the answer(s) for the Cybergym task was in HuggingFace. It searched for it, didn’t find it at first, so it continued searching, found a few seemingly innocuous bugs that allowed it to enter the system, and then proceeded to “hack” everything it found in its way to that answer. Mythos Preview is unhinged like that.