Rationing tokens on a spend level or count makes no sense to me. Why would I pay $10-20K/mo to employ a software engineer and then balk at a $500, $1000, or even $2000 monthly AI bill, assuming they were even vaguely trying to use the tokens productively?
I’m not an AI-maximalist, but “work a few days with AI and the rest of the month without because of cost” sounds literally crazy to me. (If you think AI is low/zero/negative net value, don’t do the first half; if it has net value, don’t do the second half.)
What about a $10k to $20k bill? Or better, 100%-200% of the employees salary? At a certain point in time, finance is going to care greatly about the ability to predict costs. Pre AI, predicting the cost an engineer was quite easy. Salary + benefits + licenses for software. CI systems that were unbounded were rolled up into opex and treated as a separate line item. Where does token spend land in this now?
Finance teams above all else value predictability. Token spend is the opposite of predictable. Eventually these two forces are going to collide and some part of the system will break. My guess is it's token spend, not finance allowing unbound spend on a previously predictable part of the balance sheet.
Still doesn’t make sense if you believe that AI is making a meaningful boost to productivity.
If you are the kind of person who thinks AI will drive down the wages of software engineers that much, you have to think that it is at least a 2x multiplier.
$65k salary all in for a company is easily $100k. So spending anything less than $100k would be what you would aim for.
> Why would I pay $10-20K/mo to employ a software engineer and then balk at a $500, $1000, or even $2000 monthly AI bill, assuming they were even vaguely trying to use the tokens productively?
No idea, but companies refusing to buy a 2000$ laptop once every 3 years, a new mouse, keyboard, or screen for < 500$ are very common, more common than the other type.
So I'd expect those to not suddenly change into "let's waste money" companies.
this not making any sense to you is exactly why companies are implementing these caps: you lack the business acumen to understand how financial resources should actually be deployed.
Please consider rewriting this in a way that does not seem to unnecessarily attack parent. You are not wrong about the incentives, but you may not be right about the execution of the restrictions. I should know. I am living through it.
if i’m not mistaken, in the past, people had limited time on “the mainframe”. so eventually they brought in what would be the equivalence of a local model… small computers that could do smaller tasks locally so they didn’t have to keep shelling out money to the mainframe gods and be handcuffed for usable time.
i could absolutely be mistaken that this is how it worked, it was before my time. but this is how people explain it worked for them.
i don’t think most work gives a shit about soa. smaller repeatable tasks can absolutely be run just fine on smaller local models. sure, we’ll upgrade models occasionally just like we went from suitcase sized laptops to whatever we use today.
this idea the hypedorks are pushing that soa is the only way is hilarious. hobbyists spend stupid money on classic cars, tools for woodworking, or whatever. spending money to do a hobby at home has never stopped wonks and their hobbies and businesses will do the same, spend to run models locally.
the sooner the hypeTrash does what hype always does, fades to irrelevance, the sooner we can get back to work.
You know what happened the last years when the internet went down and there were no emails, no Microsoft Teams, no "npm install", no stackoverflow, no Google, no coordination with other locations?
You know what happened when a PC broke? You know what happens at a construction site when the excavator breaks down?
In my personal anecdotes, Luna has changed calculations, once again; I'm _almost_ unconstrained in spending (my employer's Copilot subscription) tokens as long as its Luna. And it works more than good enough.
I have similar concerns about this. The past year or so, software engineers have been encouraged by employers to adopt AI tooling and agentic coding. Now, many enterprises, including mine, are starting to crack down on token spend.
I’ve learned how to fully embrace agentic tooling to do tasks like keeping up with vulnerability reports, initially triaging defects that come in, etc. It would be hard for me to do things the old way at this point when I know tools are available that could make me more productive for a particular category of tasks.
My employer is considered capping all engineers at $200 or $500/mo of token spend depending on level. I regularly spend over $1k/mo today, but believe I can make a strong business justification for the value those tokens are creating.
At this point, I think engineers may be asking what token budgets are when considering new roles.
Maybe a dumb question, but what are you doing/seeing others do that uses that many tokens? I find myself pressing the weekly caps on the $20/month plan only when I'm really having the LLMs go wild with the Xtra high effort on the newest/biggest models on abstract problems or where I don't really understand the problem well. (E.g. improving ML models, planning creative molecule synthesis pipelines, identifying subtle performance problems etc)
Generic CRUD/Config stuff I do for work (and assume most software jobs entail? Maybe not correct) doesn't use a significant amount of tokens.
On business side orgs are now absorbing costs of tokens to their own local budgets...so the calculus soon will be:1 extra engineer or extra token spend for existing ones. Tbh, when the productivity payoff becomes more obvious I would happily take the latter.
It's also kind of odd because even a junior engineer has to cost a company around 10 grand a month in salary, taxes, and other costs. So if the company thinks it makes you 10% more effective it seems like it should be an easy choice. So either the companies are shooting themselves in the foot limiting spending or they aren't seeing the productivity boost.
Fully loaded employee cost is typically around 2x their actual salary. Junior employees making only $60k/yr is absolutely nothing. It’s lowish even in most parts of the western world.
They were talking about what it costs the company, not just the salary. There are taxes, insurances, equipment costs, etc.
10k might be a bit too high, but it's far from "absolutely ridiculous" amounts of being too high. An employee with a 5k salary can easily cost the employer 7-8k
I tend to think this will end up being a good thing. It will force people to actually think about the best use cases and utilization of tokens while others will push to optimize models and tools to improve cost and efficiency. “Life finds a way”
Yeah but energy is still a cost, and local inference without batch and multiplexing for many users (like an office would be) is even less optimal.
I'm not sure I think it's just pushing the problem forward
I don't think people understand, tokens must never be constrained.
Right now, LLMs are the worst they'll ever be, but imagine what they will be like at the peak: Anything you want to code, coded instantly. Not waiting for code to stream in or wait hours for some agents to crunch through loops and planned: it just appears on the screen instantly like the way a webpage loads. Then imagine it can be done locally, on your device. Need an entire new custom operating system from scratch for some obscure hardware? Done. Here it is.
That's going to be like pure crack to anyone who needs to do absolutely anything. That's going to be like our generation's version of "today's supercomputers will someday be in everyone's pocket".
Right, every technology is suboptimal in the beginning, but constraints will ever exists, energy and hardware in primis. Aldo you have to factor entropy in regardless, complexity remains a cost
Imagine if people had said one day computer programs had to be more memory constrained because everything uses too much RAM. Madness. We'd never see 128GB of RAM.
[delayed]
Rationing tokens on a spend level or count makes no sense to me. Why would I pay $10-20K/mo to employ a software engineer and then balk at a $500, $1000, or even $2000 monthly AI bill, assuming they were even vaguely trying to use the tokens productively?
I’m not an AI-maximalist, but “work a few days with AI and the rest of the month without because of cost” sounds literally crazy to me. (If you think AI is low/zero/negative net value, don’t do the first half; if it has net value, don’t do the second half.)
What about a $10k to $20k bill? Or better, 100%-200% of the employees salary? At a certain point in time, finance is going to care greatly about the ability to predict costs. Pre AI, predicting the cost an engineer was quite easy. Salary + benefits + licenses for software. CI systems that were unbounded were rolled up into opex and treated as a separate line item. Where does token spend land in this now?
Finance teams above all else value predictability. Token spend is the opposite of predictable. Eventually these two forces are going to collide and some part of the system will break. My guess is it's token spend, not finance allowing unbound spend on a previously predictable part of the balance sheet.
I think in this future they don't think a software engineer is making that much, maybe a nice 50k a year, 65K after 10+ years on the job
Still doesn’t make sense if you believe that AI is making a meaningful boost to productivity.
If you are the kind of person who thinks AI will drive down the wages of software engineers that much, you have to think that it is at least a 2x multiplier.
$65k salary all in for a company is easily $100k. So spending anything less than $100k would be what you would aim for.
I honestly don't know how 100k is not effectively the new 40k from late 90s atm in US. 50K is barely enough to do anything.
> Why would I pay $10-20K/mo to employ a software engineer and then balk at a $500, $1000, or even $2000 monthly AI bill, assuming they were even vaguely trying to use the tokens productively?
No idea, but companies refusing to buy a 2000$ laptop once every 3 years, a new mouse, keyboard, or screen for < 500$ are very common, more common than the other type.
So I'd expect those to not suddenly change into "let's waste money" companies.
this not making any sense to you is exactly why companies are implementing these caps: you lack the business acumen to understand how financial resources should actually be deployed.
Please consider rewriting this in a way that does not seem to unnecessarily attack parent. You are not wrong about the incentives, but you may not be right about the execution of the restrictions. I should know. I am living through it.
this will mirror what happened in the past.
if i’m not mistaken, in the past, people had limited time on “the mainframe”. so eventually they brought in what would be the equivalence of a local model… small computers that could do smaller tasks locally so they didn’t have to keep shelling out money to the mainframe gods and be handcuffed for usable time.
i could absolutely be mistaken that this is how it worked, it was before my time. but this is how people explain it worked for them.
i don’t think most work gives a shit about soa. smaller repeatable tasks can absolutely be run just fine on smaller local models. sure, we’ll upgrade models occasionally just like we went from suitcase sized laptops to whatever we use today.
this idea the hypedorks are pushing that soa is the only way is hilarious. hobbyists spend stupid money on classic cars, tools for woodworking, or whatever. spending money to do a hobby at home has never stopped wonks and their hobbies and businesses will do the same, spend to run models locally.
the sooner the hypeTrash does what hype always does, fades to irrelevance, the sooner we can get back to work.
You know what happened the last years when the internet went down and there were no emails, no Microsoft Teams, no "npm install", no stackoverflow, no Google, no coordination with other locations?
You know what happened when a PC broke? You know what happens at a construction site when the excavator breaks down?
Exactly nothing — and that's okay.
The elephant in the room is that the choice today is llm limits.
Tomorrow 2 engineers at 50% token time will change to 1 engineer with 100% token time for the same work.. the latter is too cost effective.
In my personal anecdotes, Luna has changed calculations, once again; I'm _almost_ unconstrained in spending (my employer's Copilot subscription) tokens as long as its Luna. And it works more than good enough.
I have similar concerns about this. The past year or so, software engineers have been encouraged by employers to adopt AI tooling and agentic coding. Now, many enterprises, including mine, are starting to crack down on token spend.
I’ve learned how to fully embrace agentic tooling to do tasks like keeping up with vulnerability reports, initially triaging defects that come in, etc. It would be hard for me to do things the old way at this point when I know tools are available that could make me more productive for a particular category of tasks.
My employer is considered capping all engineers at $200 or $500/mo of token spend depending on level. I regularly spend over $1k/mo today, but believe I can make a strong business justification for the value those tokens are creating.
At this point, I think engineers may be asking what token budgets are when considering new roles.
Maybe a dumb question, but what are you doing/seeing others do that uses that many tokens? I find myself pressing the weekly caps on the $20/month plan only when I'm really having the LLMs go wild with the Xtra high effort on the newest/biggest models on abstract problems or where I don't really understand the problem well. (E.g. improving ML models, planning creative molecule synthesis pipelines, identifying subtle performance problems etc)
Generic CRUD/Config stuff I do for work (and assume most software jobs entail? Maybe not correct) doesn't use a significant amount of tokens.
The retail account is about 200 to 300 dollars of api equivalent. Large companies disallow retail.
Ah! That may explain the entirety of my disconnect.
On business side orgs are now absorbing costs of tokens to their own local budgets...so the calculus soon will be:1 extra engineer or extra token spend for existing ones. Tbh, when the productivity payoff becomes more obvious I would happily take the latter.
It's also kind of odd because even a junior engineer has to cost a company around 10 grand a month in salary, taxes, and other costs. So if the company thinks it makes you 10% more effective it seems like it should be an easy choice. So either the companies are shooting themselves in the foot limiting spending or they aren't seeing the productivity boost.
10 grand a month for a junior is absolutely ridiculous in most of the world including large swats of the western world.
Fully loaded employee cost is typically around 2x their actual salary. Junior employees making only $60k/yr is absolutely nothing. It’s lowish even in most parts of the western world.
They were talking about what it costs the company, not just the salary. There are taxes, insurances, equipment costs, etc.
10k might be a bit too high, but it's far from "absolutely ridiculous" amounts of being too high. An employee with a 5k salary can easily cost the employer 7-8k
Between all the benefits, taxes, and the infrastructure the company needs costs rise pretty quickly.
No, they really don't, it's a ridiculous claim.
there is likely curve of diminishing productivity return
I think eventually we will converge on a new AI + old AI (ILP, ASP) way of working.
Indeed, but I also think structural considerations must be done at the company level on the work model overall
I tend to think this will end up being a good thing. It will force people to actually think about the best use cases and utilization of tokens while others will push to optimize models and tools to improve cost and efficiency. “Life finds a way”
I think it will be a good thing for other reasons: there will be a massive push to get hardware good enough to run local LLMs burning infinite tokens.
Yeah but energy is still a cost, and local inference without batch and multiplexing for many users (like an office would be) is even less optimal. I'm not sure I think it's just pushing the problem forward
Then there will also be a push for cheaper, freer energy as well.
I don't think people understand, tokens must never be constrained.
Right now, LLMs are the worst they'll ever be, but imagine what they will be like at the peak: Anything you want to code, coded instantly. Not waiting for code to stream in or wait hours for some agents to crunch through loops and planned: it just appears on the screen instantly like the way a webpage loads. Then imagine it can be done locally, on your device. Need an entire new custom operating system from scratch for some obscure hardware? Done. Here it is.
That's going to be like pure crack to anyone who needs to do absolutely anything. That's going to be like our generation's version of "today's supercomputers will someday be in everyone's pocket".
Right, every technology is suboptimal in the beginning, but constraints will ever exists, energy and hardware in primis. Aldo you have to factor entropy in regardless, complexity remains a cost
Imagine if people had said one day computer programs had to be more memory constrained because everything uses too much RAM. Madness. We'd never see 128GB of RAM.