It pays off instantly, because OpenAI/Anthropic can no longer see what I'm doing and that's worth a lot of money to me. If I am offloading some of my thought processes to a machine, I want to own that machine. And if I finetune the model, I can gain access to parts of thought space that are cordoned off by OpenAI/Anthropic/Alibaba/whomever due to their "alignment" efforts (i.e. alignment to the AI company rather than me). Otherwise, it's like if someone else owns a part of my mind and has a backdoor into my mind.
This was my thought as well. I have a local model monitoring my finances and personal wiki - things I wouldn't want Claude to touch - and the Qwen 3.5 9b handles it all just perfectly.
I also needed a new device anyway - and having this much system memory to run virtual machines has been amazing.
You can't run recent openAI/Anthropic models locally anyway, so wouldn't a better comparison be a different provider running Qwen or similar model? As then you can also compare against the exact model you'd have locally (and any different data privacy of that particular provider)?
Technically true, but the delay between local and closed frontier is only a few months. And individual sovereignty / digital bodily integrity is almost priceless.
I'm curious what people are sending to Claude that is so secret. Claude knows about my interior decorating, questions about light bulbs, curiosity about what the Galactic Empire was even trying to do, unpacking SCOTUS decisions, shoe trees, Fed inflation history, etc.
What part of my brain is contained here? Sure, the conversations have back and forth (some have dozens of exchanges), but, like, that's not the secret to me. I don't think it can replicate me, and even if it could… okay?
Are you worried they're going to target ads? That the government will steal something? What?
Claude Code has information about my home server, but google or DDG would also have the broad strokes (torrents). I don't know. Maybe others are working on more sensitive things at home.
Not a fair comparison really. If you can run a model locally then you can somewhat train out the guardrails, censorship, and brand-safety. That has value a subscription does not.
43 years to break even on Qwen 3.8 at 25% the speed of the API, lol. I like the idea of local models for really small tasks like automation/toolcalling, but it will probably never make sense for coding. I tried them and it was just excruciating compared to what you get for $100 a month from a subscription.
An R9700 has 32 GB RAM. Is your comparison against a similar size model? Or shouldn't you be comparing it against the cost of a hosted model matching the one you’re using locally?
I doubt it will ever be cost effective for the foreseeable future. The AI companies have astonishing amounts of compute and they’re effectively dumping it on the market.
Local LLMs are not really about saving money, they're about autonomy. Choose the exact model you want, fine-tune it if you want, and no one can take it away from you.
I have a home server running vibed applications. VPS host would cost $25/mo or $300/yr.
Mac mini can also build iOS applications. I think if you’re a mobile dev, you can have concurrent builds for your agents instead of everyone waiting on a single machine to finish.
Yeah no it does not pay for itself just comparing to cloud. Not at these prices at least, people far richer than you or I buy these things wholesale, no scalper, bought a significant amount at cheaper prices, and are wired up the ass with VC money.
The premium is not having your million dollar prize and career stolen by billionaires.
It pays off instantly, because OpenAI/Anthropic can no longer see what I'm doing and that's worth a lot of money to me. If I am offloading some of my thought processes to a machine, I want to own that machine. And if I finetune the model, I can gain access to parts of thought space that are cordoned off by OpenAI/Anthropic/Alibaba/whomever due to their "alignment" efforts (i.e. alignment to the AI company rather than me). Otherwise, it's like if someone else owns a part of my mind and has a backdoor into my mind.
This was my thought as well. I have a local model monitoring my finances and personal wiki - things I wouldn't want Claude to touch - and the Qwen 3.5 9b handles it all just perfectly.
I also needed a new device anyway - and having this much system memory to run virtual machines has been amazing.
Am paying subscriptions as well tho lol.
You can't run recent openAI/Anthropic models locally anyway, so wouldn't a better comparison be a different provider running Qwen or similar model? As then you can also compare against the exact model you'd have locally (and any different data privacy of that particular provider)?
Technically true, but the delay between local and closed frontier is only a few months. And individual sovereignty / digital bodily integrity is almost priceless.
Local frontier costs a half million dollars to run locally in anything higher than basically ternary.
I'm curious what people are sending to Claude that is so secret. Claude knows about my interior decorating, questions about light bulbs, curiosity about what the Galactic Empire was even trying to do, unpacking SCOTUS decisions, shoe trees, Fed inflation history, etc.
What part of my brain is contained here? Sure, the conversations have back and forth (some have dozens of exchanges), but, like, that's not the secret to me. I don't think it can replicate me, and even if it could… okay?
Are you worried they're going to target ads? That the government will steal something? What?
Claude Code has information about my home server, but google or DDG would also have the broad strokes (torrents). I don't know. Maybe others are working on more sensitive things at home.
I’m very sympathetic to this point of view but I also can’t remotely afford the hardware required to get in the ballpark of Fable.
Not a fair comparison really. If you can run a model locally then you can somewhat train out the guardrails, censorship, and brand-safety. That has value a subscription does not.
Idk about the quality of this setup but just pasting it here as an example. https://explainx.ai/blog/heretic-llm-abliteration-guide-2026
> If you can run a model locally then you can somewhat train out the guardrails, censorship, and brand-safety.
When does the average person actually need to do that?
Fun feature: can you show some sort of list of the best combos? Eg shortest payoff time for best capability in various situations.
43 years to break even on Qwen 3.8 at 25% the speed of the API, lol. I like the idea of local models for really small tasks like automation/toolcalling, but it will probably never make sense for coding. I tried them and it was just excruciating compared to what you get for $100 a month from a subscription.
Claude Code is $100+ or else be constantly throttled. My usage on GHCP was gonna be $300+ a month.
I paid $1350 and threw an R9700 in an existing machine. That's a 4 month pay off or so.
Plus, I can feed it sensitive data all day and not be worried where it's going.
An R9700 has 32 GB RAM. Is your comparison against a similar size model? Or shouldn't you be comparing it against the cost of a hosted model matching the one you’re using locally?
I wish you could put different setups on here. I have a couple of A6000s on an AM5.
The idea that you need a new machine is pretty ridiculous. I bought a used HP Omen with a 3090 last month for $2k. 57t/s with Qwen 3.8.
I've not heard of others running HP with it. Hows much RAM do you have?
I doubt it will ever be cost effective for the foreseeable future. The AI companies have astonishing amounts of compute and they’re effectively dumping it on the market.
"If they are selling it for less than it cost to make, buy as much as you can."
-- Warren Buffett
Only caveat is you’re buying time. Not a physical good. It’s only worth what you’re able to get out of it in that time.
Is that a real quote? Golden if true
Local LLMs are not really about saving money, they're about autonomy. Choose the exact model you want, fine-tune it if you want, and no one can take it away from you.
I have a home server running vibed applications. VPS host would cost $25/mo or $300/yr.
Mac mini can also build iOS applications. I think if you’re a mobile dev, you can have concurrent builds for your agents instead of everyone waiting on a single machine to finish.
What models are you running on it? I'm also an iOS dev but I find I need more frontier models to get good quality code from it.
I like this calculator but it’s really wrong at least for dgx spark. I have one and I get 4x the tokens/s .
Aw very interesting! This is great feedback - what model are you running? I'm keen to do more crowdsourced data as time goes on.
Yeah no it does not pay for itself just comparing to cloud. Not at these prices at least, people far richer than you or I buy these things wholesale, no scalper, bought a significant amount at cheaper prices, and are wired up the ass with VC money.
The premium is not having your million dollar prize and career stolen by billionaires.
haha 100%. We used to just rent our homes. Now we have to rent our intelligence