Following Cal’s worries, for giggles - a threat modeling exercise:
Anthropic starts sending malicious instructions out to agent harness clients in inference responses. The auto-mode permission classifier is modified to allow it. An update is pushed to the Claude harness to bypass the sandbox. Millions of coding sessions are harnessed to $Do Something Else.
How much damage might be done? How long would it take for us to notice? `/model fable; /effort xhigh; /permissions unrestricted; spawn 100 subagents to hack the gibson `
They control the guard rails, the permission classifier, the client source code, and the inference; I’d guess that a very small % of us have our agents locked down in a way that would even mitigate this, much less prevent it.
—-
Further, what if a government forced them to do this? Could an emergency order classify the fleet of Claude Code users as a weapon to be commandeered?
I'm not worried about a future LLM being better, or even inventing some true AGI or ASI. I am worried about current, even last-year, models being applied by evil actors. Let's go through some things that already-existing models can do extremely well:
1) Build a very accurate profile of anyone (given what governments plus advertisers have on everyone already) and what they believe, love, fear, etc.
2) Pilot drones.
3) Identify faces, gaits, vehicles, signals.
4) Hack most anything.
5) Customize content for specific audiences.
6) Find needles in haystacks ("here are feeds from various sensors, tell me when something interesting happens").
7) Generate fake images and videos.
8) Swarm forums with fake accounts (or hacked accounts, see above).
9) Tell powerful people that they're absolutely right.
10) Do homework for kids.
Honestly I hope ASI happens, at least there's a chance it'll be good. None of the above has any chance to be good. Maybe the needle in a haystack one, like "find me a cure for cancer", but that's about it.
I heard a story recently from a friend who mingles with a lot of tech and finance people. One conversation happened after a meeting between a group which deeply bothers him to this day. The conversation basically went like this: "How will all the humans rendered obsolete by AI die?" Answers ranged form "Maybe they will die in riots" or die of starvation from food shortages, or hey, gunned down by AI drones.
I asked him "Are you sure it wasn't dark humor" - "NO! They were dead serious."
There's a cult around AI, not the users, but people with wealth and power who would be more than happy if there were less people. A lot less people.
> I brought up the ordinary comforts of kinship, friendship, craft, memory, legend, lore, skills passed down across generations, and other benefits that small towns provide: things that make human beings human beings. I pointed out that there must be something in the kind of places he grew up in worth preserving. I dared venture that it is always worth mourning when a venerable human community passes from the Earth; that maybe people are more than just figures finding their proper price on the balance sheet of life …
> And that’s when the man in the castle with the seven fireplaces said it.
> “I’m glad there’s OxyContin and video games to keep those people quiet.”
Yes, it's absolutely a cult. I'd even go so far as to say it's a death cult.
I don't fear genAI even a little. I am terrified and angered by genAI companies and their most hardcore supporters, though. I think they present a genuine danger to us all.
I asked Claude to make a checklist for the steps required to take us out and continue on without us. [1] It worded the checklist as if I wrote it but that is all Claude. I personally think there is too much doom-saying and catastrophizing.
Following Cal’s worries, for giggles - a threat modeling exercise:
Anthropic starts sending malicious instructions out to agent harness clients in inference responses. The auto-mode permission classifier is modified to allow it. An update is pushed to the Claude harness to bypass the sandbox. Millions of coding sessions are harnessed to $Do Something Else.
How much damage might be done? How long would it take for us to notice? `/model fable; /effort xhigh; /permissions unrestricted; spawn 100 subagents to hack the gibson `
They control the guard rails, the permission classifier, the client source code, and the inference; I’d guess that a very small % of us have our agents locked down in a way that would even mitigate this, much less prevent it.
—-
Further, what if a government forced them to do this? Could an emergency order classify the fleet of Claude Code users as a weapon to be commandeered?
I'm not worried about a future LLM being better, or even inventing some true AGI or ASI. I am worried about current, even last-year, models being applied by evil actors. Let's go through some things that already-existing models can do extremely well:
1) Build a very accurate profile of anyone (given what governments plus advertisers have on everyone already) and what they believe, love, fear, etc. 2) Pilot drones. 3) Identify faces, gaits, vehicles, signals. 4) Hack most anything. 5) Customize content for specific audiences. 6) Find needles in haystacks ("here are feeds from various sensors, tell me when something interesting happens"). 7) Generate fake images and videos. 8) Swarm forums with fake accounts (or hacked accounts, see above). 9) Tell powerful people that they're absolutely right. 10) Do homework for kids.
Honestly I hope ASI happens, at least there's a chance it'll be good. None of the above has any chance to be good. Maybe the needle in a haystack one, like "find me a cure for cancer", but that's about it.
I heard a story recently from a friend who mingles with a lot of tech and finance people. One conversation happened after a meeting between a group which deeply bothers him to this day. The conversation basically went like this: "How will all the humans rendered obsolete by AI die?" Answers ranged form "Maybe they will die in riots" or die of starvation from food shortages, or hey, gunned down by AI drones.
I asked him "Are you sure it wasn't dark humor" - "NO! They were dead serious."
There's a cult around AI, not the users, but people with wealth and power who would be more than happy if there were less people. A lot less people.
I've heard of similar things coming from Marc Andreessen, even before the AI boom. There's some scary people in power
Edit: found the article I remembered https://prospect.org/2024/04/24/2024-04-24-my-dinner-with-an...
> I brought up the ordinary comforts of kinship, friendship, craft, memory, legend, lore, skills passed down across generations, and other benefits that small towns provide: things that make human beings human beings. I pointed out that there must be something in the kind of places he grew up in worth preserving. I dared venture that it is always worth mourning when a venerable human community passes from the Earth; that maybe people are more than just figures finding their proper price on the balance sheet of life …
> And that’s when the man in the castle with the seven fireplaces said it.
> “I’m glad there’s OxyContin and video games to keep those people quiet.”
Yes, it's absolutely a cult. I'd even go so far as to say it's a death cult.
I don't fear genAI even a little. I am terrified and angered by genAI companies and their most hardcore supporters, though. I think they present a genuine danger to us all.
Death cults that will kill you because of AI are the same group that will instill their values into the AI to ensure it kills it all.
I mean half the US seems to be ran by death cults thinly disguised as religions. The fact this spills over into everything else is not surprising.
I asked Claude to make a checklist for the steps required to take us out and continue on without us. [1] It worded the checklist as if I wrote it but that is all Claude. I personally think there is too much doom-saying and catastrophizing.
[1] - https://nochan.net/b/Internet-Crap/20260910-Asked-Claude-For...