I think this is probably a good thing. It’s not just public services, but any company that consumers have to deal with (in ways they don’t want to). Historically, only those with the most time, money, connections, and education have been able to extract benefits from these systems. E.g., if someone wants to fight an insurance claim being denied, they might need to hire an expensive lawyer or spend hours performing their own research.
What I’m particularly curious about is what the “endgame” of this looks like. Initially, I imagine we’ll start with LLMs on the consumer side debating LLMs on the service side (this is arguably already the case for a few industries). But at some point, unless identical conditions lead to identical outcomes for the people involved, there will be lawsuits. So I suspect this will lead to a much more formally rigorous system with various “certificates” proving, e.g., that you lost your job or that a doctor has confirmed you have a disability.
On the one hand, that has the potential to make things much fairer across economic classes. You don’t get special benefits by having upper class resources. On the other hand, it could lead to increasing the actual requirements for certain benefits, which then reduces class mobility by locking in the benefits to those who already have them, with no clever way to wriggle out of your own level.
I’m just kind of thinking out loud. I don’t know what’s to going to happen long-term, but I think it’s important to monitor these changes carefully as they’re occurring.
> What I’m particularly curious about is what the “endgame” of this looks like.
A higher fraction of money going into the system is lost to navigating adversarial processes. And probably the occasional person getting in legal trouble for not validating their LLM output before feeding it to a formal process.
For customer-funded things (eg home insurance), this means higher prices for the same in-practice benefits (as distinct from nominal benefits; any gap between nominal and actual benefits will probably end up smaller).
For publicly-funded things (government benefits), this could show as increased budgets for the same result or decreased results on the same budget.
.
> So I suspect this will lead to a much more formally rigorous system with various “certificates” proving, e.g., that you lost your job or that a doctor has confirmed you have a disability.
Well, maybe standardization of how employers or doctors or whoever are required to report things to the government.
I disagree - this doesn’t make it more accessible for people who don’t have the means, it drowns the receivers in bureaucracy and forces them to do orders of magnitude more work out of a lack of understanding. It’s debatable whether or not these systems are built in good faith to allow benefit claimants to appeal or not, but what will happen here is the appeals process will be clamped down on, hard and heavy.
It’s the equivalent of downloading the entire source code of chrome to access a website. are you entitled to do it? Absolutely. Is it a gigantic pain in the ass for the person who has to serve you the 2GB of source code from their git mirror? Yep
Huh, the receivers in this case are the bureaucracy. The senders are the one that has to navigate said bureaucracy.
So, the entire thing is a wash. The systems are complex and confusing, many of them likely by design. Getting an appointment in some of them can take months. Knowing what you actually need can be daunting as quite often there is a byzantine set of requirements that depends on your exact situation and any missed information could take weeks or months to get another visit.
TBH, most government bureaucracy is just pure overhead connected to a jobs program.
A competent government, focused on delivering the services that citizens are actually entitled to, not on paying clerks or avoiding career-ending events at all costs, would do most of this digitally. And by digitally, I don't mean "you can fill the form online, maybe with some auto-fill and autocomplete suggestions for addresses", I mean just sending you the transfer if you're eligible, automatically fetching all the info they need to see if that is the case. If there even is a form, that's too much overhead.
Even Estonia (and other tech-forward EU governments) don't do this. Poland is considered pretty good on the digital front, but many digital processes still result in a (digitally-transmitted) form being printed in some obscure processing center and manually processed by a clerk during business hours.
In my understanding when it comes to doling out benefits, it's not an overhead connected to a jobs program. The less eligible people use benefits, the less the government has to pay out benefits. It's strictly better for the government to make getting benefits as hard as possible.
I hope it will force institutions to simplify entitlements. Right now the institutions can pretend like their kafkaesque system is fine because nobody manages to navigate the maze.
Maybe there is an economic equalizer, but I feel like Fable is going to win more appeals than GLM 5.2, so perhaps not truly an equalizer. If the LLM you're using to run your appeal is better than the one your insurance company is using to process your appeal, then you have a better chance of winning. So we have an arms race type situation, and money is always good in arms races.
i agree with most of this, but for appeals and thing where there is a clear correct and fair answer, i think the models converge on the same answer regardless of level of intelligence (e.g., if you ask fable 5 and sonnet 5 what 2+2 is, you’ll get the same output)
That's a good point. It will be interesting to see how it plays out.
In general, I do think this is positive. Companies love putting barriers in front of everything and maybe the AI arms race will result in "fine, you can have a button to cancel" or "fine, we'll approve your medically necessary claim without making you do an appeals dance for our entertainment".
> What I’m particularly curious about is what the “endgame” of this looks like. Initially, I imagine we’ll start with LLMs on the consumer side debating LLMs on the service side
More data centers, less low-end lawyers, probably more pressure on the judicial system.
> So I suspect this will lead to a much more formally rigorous system with various “certificates” proving, e.g., that you lost your job or that a doctor has confirmed you have a disability.
We're seeing how this plays out now in various states that have made their benefits systems nearly impossible to use, which helps them save money by not doling out benefits at all. Having helped some of my elder relatives try to apply for social security payments online makes me very, very concerned about how it's going to look when I'm that age.
The cost of an automated bot that will talk you endlessly in circles until you give up is negligible now.
I think people using LLMs the right way is great and will democratize information.
However, most of us have experienced people wasting our time, sending us AI slop. People need to learn how to use the tools properly and respect other people's time.
This is good I think. I have a friend that worked for a Law Firm / NGO / Charity kind of thing, whose primary purpose was writing appeals for benefits to the VA. These were benefits that the person in question was legally entitled to, but required outside legal assistance to actually attain. Regardless of your personal opinion on defense, VA fraud, etc. it would seem good that people who are legally entitled to a benefit are more easily able to attain the benefit.
I don't understand why this is often portrayed negatively. In general it's important that people can defend their rights, and I find it great that LLMs can help a lot with that. One could introduce page/word/character limits, and note that the other side can use LLMs too to summarize/organize/extract. Overall I think it's great that LLMs now help me to understand and rightfully claim things from insurances, companies, etc.
>I don't understand why this is often portrayed negatively. In general it's important that people can defend their rights, and I find it great that LLMs can help a lot with that.
Well I thought it was self evident. It’s portrayed negatively because it causes a denial of service. You should view this positively if you want worse access to public benefits.
Imagine if every time someone reached out to you to ask you a question they did a SAR for all of the information you had on them. They’re technically entitled to do that in many jurisdictions, but it probably won’t help with getting the answer to the question. It’s the same problem as drive by AI PRs except the people receiving them aren’t allowed to ignore them.
In my view there are three takeaways from this, all of which are correct:
1. Already strained government services need to bear increasing input, affecting waiting times and employee workloads, which is bad.
2. LLMs are empowering citizens to appeal in a social security landscape which often require expert-level knowledge to navigate, which is good.
3.This phenomenon says more about the complexity of our social security systems than it does about LLMs, hopefully prompting our governments to review and start thinking about social security reform.
2b. LLMs are empowering bad actors to exploit loopholes and collect benefits they're not entitled to.
The assumption from a lot of commenters seems to be that these are all legitimate claims and appeals, not a gold rush to get money from the government while the getting's good. In the end it will be the good faith actors that suffer - both those relying on these benefits to survive who will have to jump through more hoops to prove it and those funding the system with their taxes.
Having older parents and watching what my grandparents went thru long before LLMs I just don't worry about the poor old state that much.
The political zeitgeist is to always blame the "bad people", the "welfare queens" and the "undeserving" as the problem with the system. But never the complex system that's seemingly designed to benefit things like corporations extracting as much wealth from the system as they can. Or politicians ear marking the money for things they want to do, rather than the people that put the money in the system.
So, yea, your reasoning just really doesn't hit me in the heartstrings at all.
LLMs are just revealing how many systems rely on making applications artificially difficult, either to understand or to generate, to reduce volume. These systems have been failing silently for years, unmuting the alarm is the first step to fixing them.
It should be more about how Public Services are underfunded and the process for applying and appeals for benefits is designed to be slow to account for this underfunding, but now something has bypassed one part of the design. It should be an argument as to why we need more funding for public services.
To me this is one of the reasons that UBI might not be most horrible idea. Sufficiently calibrated UBI would remove this useless work on both sides. Thus making whole system more efficient and maybe even cheaper.
It's becoming obvious that essentially every public-facing organization needs to change. Short story, poetry, essay submissions, comments, reviews, conference paper and journal submissions, polling, job apps, etc.
At this point, I think it's an open question whether there's a scalable way to filter for human effort, that isn't just another layer of AI (e.g. Pangram). I'd love to see something practical come out of this, since the default is probably for public interfaces to just cease to exist, and convert to referral only systems or pure noise.
This is something I have been working on. One of the main issues my users have run into is navigation. The Gov. websites often are not very user friendly; so its hard and time consuming to navigate to and apply for the correct services.
To address this I have users just speak on their "Intent" for example: "I need help with housing and food assistance, my kid has a peanut allergy and I just lost my job so I don't know what to do."
Note: This is a real request
What happens here it a custom program is created, the program has a protocol that shows exactly how the agent navigates the Gov. system and applies for the benefits. Standard modules are applied to actually do the real work of applying, followups etc...
Despite having a lot of professional experience I have not been able to even get a follow-up interview for a job so I am just pouring my efforts into this.
I have been able to help people get housed, apply for benefits, get into shelters and most recently I started working on erpo.law to speed up and reduce the paperwork on all sides for red flag laws.
Lots of the people that have helped me have done so under the table and not at the approval of their agency heads.
Here is a short vid showing a request that locates an existing program, programs can be codified into re-useable units for people w/ common needs or created in real time... Staffed via O*NET roles that hit real providers I minded from other websites.
Note: This one matters to me a lot because my mom was in and out of mental institutions my whole life and I been committed before so I really like helping that population because I get it more than most do.
Im also doing city agencies NYC, Boston, Miami, LA, and Vegas to start because they all have their own tld so its easier for me to filter out the bs traffic and bots and stuff.
In terms of scaling human effort the HITL for each module/protocol-chain step can be delegated where needed and the program will persist that prefrence etc...
The people who need this additional help to access water/power/internet/etc often have the most barriers to being able to complete the necessary steps to access benefits. This is a great use of the technology.
"Flooding" implies some nefarious purpose. There's no evidence of that in this paper. Maybe these agents are finally giving people who could never find the time, resources, knowledge, etc. to engage with the system.
"Public services are increasingly becoming more accessible with LLMs" How about that as the title?
They also have zero evidence that this is because of LLMs. They just got some data related to volume, did a really crappy analysis that wouldn't even stand up as an intro to stats course project, and declared their biased opinion as likely true.
"Flooding" could be less nefarious as in the number of requests have risen dramatically in a short period of time quite possibly overwhelming those that are tasked to deal with the requests. Most places have peaks and valleys and can staff appropriately. Having a sudden rise from something unexpected will overwhelm the normal staff. We can debate whether the lack of recognition of the upticks being intentional or not, but possibly the ignoring the uptick is the nefarious bit
This is a bad thing. It means public services will ultimately deploy AI to respond to those requests, and people will be denied instantly with no recourse or human oversight, it will be tyranny of the machine.
I think this is probably a good thing. It’s not just public services, but any company that consumers have to deal with (in ways they don’t want to). Historically, only those with the most time, money, connections, and education have been able to extract benefits from these systems. E.g., if someone wants to fight an insurance claim being denied, they might need to hire an expensive lawyer or spend hours performing their own research.
What I’m particularly curious about is what the “endgame” of this looks like. Initially, I imagine we’ll start with LLMs on the consumer side debating LLMs on the service side (this is arguably already the case for a few industries). But at some point, unless identical conditions lead to identical outcomes for the people involved, there will be lawsuits. So I suspect this will lead to a much more formally rigorous system with various “certificates” proving, e.g., that you lost your job or that a doctor has confirmed you have a disability.
On the one hand, that has the potential to make things much fairer across economic classes. You don’t get special benefits by having upper class resources. On the other hand, it could lead to increasing the actual requirements for certain benefits, which then reduces class mobility by locking in the benefits to those who already have them, with no clever way to wriggle out of your own level.
I’m just kind of thinking out loud. I don’t know what’s to going to happen long-term, but I think it’s important to monitor these changes carefully as they’re occurring.
> What I’m particularly curious about is what the “endgame” of this looks like.
A higher fraction of money going into the system is lost to navigating adversarial processes. And probably the occasional person getting in legal trouble for not validating their LLM output before feeding it to a formal process.
For customer-funded things (eg home insurance), this means higher prices for the same in-practice benefits (as distinct from nominal benefits; any gap between nominal and actual benefits will probably end up smaller).
For publicly-funded things (government benefits), this could show as increased budgets for the same result or decreased results on the same budget.
.
> So I suspect this will lead to a much more formally rigorous system with various “certificates” proving, e.g., that you lost your job or that a doctor has confirmed you have a disability.
Well, maybe standardization of how employers or doctors or whoever are required to report things to the government.
I disagree - this doesn’t make it more accessible for people who don’t have the means, it drowns the receivers in bureaucracy and forces them to do orders of magnitude more work out of a lack of understanding. It’s debatable whether or not these systems are built in good faith to allow benefit claimants to appeal or not, but what will happen here is the appeals process will be clamped down on, hard and heavy.
It’s the equivalent of downloading the entire source code of chrome to access a website. are you entitled to do it? Absolutely. Is it a gigantic pain in the ass for the person who has to serve you the 2GB of source code from their git mirror? Yep
>it drowns the receivers in bureaucracy
Huh, the receivers in this case are the bureaucracy. The senders are the one that has to navigate said bureaucracy.
So, the entire thing is a wash. The systems are complex and confusing, many of them likely by design. Getting an appointment in some of them can take months. Knowing what you actually need can be daunting as quite often there is a byzantine set of requirements that depends on your exact situation and any missed information could take weeks or months to get another visit.
TBH, most government bureaucracy is just pure overhead connected to a jobs program.
A competent government, focused on delivering the services that citizens are actually entitled to, not on paying clerks or avoiding career-ending events at all costs, would do most of this digitally. And by digitally, I don't mean "you can fill the form online, maybe with some auto-fill and autocomplete suggestions for addresses", I mean just sending you the transfer if you're eligible, automatically fetching all the info they need to see if that is the case. If there even is a form, that's too much overhead.
Even Estonia (and other tech-forward EU governments) don't do this. Poland is considered pretty good on the digital front, but many digital processes still result in a (digitally-transmitted) form being printed in some obscure processing center and manually processed by a clerk during business hours.
In my understanding when it comes to doling out benefits, it's not an overhead connected to a jobs program. The less eligible people use benefits, the less the government has to pay out benefits. It's strictly better for the government to make getting benefits as hard as possible.
I hope it will force institutions to simplify entitlements. Right now the institutions can pretend like their kafkaesque system is fine because nobody manages to navigate the maze.
Maybe there is an economic equalizer, but I feel like Fable is going to win more appeals than GLM 5.2, so perhaps not truly an equalizer. If the LLM you're using to run your appeal is better than the one your insurance company is using to process your appeal, then you have a better chance of winning. So we have an arms race type situation, and money is always good in arms races.
i agree with most of this, but for appeals and thing where there is a clear correct and fair answer, i think the models converge on the same answer regardless of level of intelligence (e.g., if you ask fable 5 and sonnet 5 what 2+2 is, you’ll get the same output)
That's a good point. It will be interesting to see how it plays out.
In general, I do think this is positive. Companies love putting barriers in front of everything and maybe the AI arms race will result in "fine, you can have a button to cancel" or "fine, we'll approve your medically necessary claim without making you do an appeals dance for our entertainment".
> What I’m particularly curious about is what the “endgame” of this looks like. Initially, I imagine we’ll start with LLMs on the consumer side debating LLMs on the service side
More data centers, less low-end lawyers, probably more pressure on the judicial system.
> So I suspect this will lead to a much more formally rigorous system with various “certificates” proving, e.g., that you lost your job or that a doctor has confirmed you have a disability.
We're seeing how this plays out now in various states that have made their benefits systems nearly impossible to use, which helps them save money by not doling out benefits at all. Having helped some of my elder relatives try to apply for social security payments online makes me very, very concerned about how it's going to look when I'm that age.
The cost of an automated bot that will talk you endlessly in circles until you give up is negligible now.
In California I have seen this occur
> various states that have made their benefits systems nearly impossible to use, which helps them save money by not doling out benefits at all
I think people using LLMs the right way is great and will democratize information.
However, most of us have experienced people wasting our time, sending us AI slop. People need to learn how to use the tools properly and respect other people's time.
Now read the part where Voter Roll Purge request portal was flooded.
This is good I think. I have a friend that worked for a Law Firm / NGO / Charity kind of thing, whose primary purpose was writing appeals for benefits to the VA. These were benefits that the person in question was legally entitled to, but required outside legal assistance to actually attain. Regardless of your personal opinion on defense, VA fraud, etc. it would seem good that people who are legally entitled to a benefit are more easily able to attain the benefit.
I don't understand why this is often portrayed negatively. In general it's important that people can defend their rights, and I find it great that LLMs can help a lot with that. One could introduce page/word/character limits, and note that the other side can use LLMs too to summarize/organize/extract. Overall I think it's great that LLMs now help me to understand and rightfully claim things from insurances, companies, etc.
To be fair, the source mentions the upside ("Although improving Service accessibility is beneficial [...]")
Seems like bureaucracy was added as a subtle way to introduce rate limits, and AI is challenging that mechanism.
> Seems like bureaucracy was added as NOT a subtle way to introduce rate limits
some teams get incentives when the total awards go down in a stable, badly run bureaucracy.. this is real
LLMs are a magic machine to turn commons into tragedies.
Great quote.
Good luck trying to get your book published.
Or run an open source repo from community contributions.
>I don't understand why this is often portrayed negatively. In general it's important that people can defend their rights, and I find it great that LLMs can help a lot with that.
Well I thought it was self evident. It’s portrayed negatively because it causes a denial of service. You should view this positively if you want worse access to public benefits.
Imagine if every time someone reached out to you to ask you a question they did a SAR for all of the information you had on them. They’re technically entitled to do that in many jurisdictions, but it probably won’t help with getting the answer to the question. It’s the same problem as drive by AI PRs except the people receiving them aren’t allowed to ignore them.
In my view there are three takeaways from this, all of which are correct:
1. Already strained government services need to bear increasing input, affecting waiting times and employee workloads, which is bad.
2. LLMs are empowering citizens to appeal in a social security landscape which often require expert-level knowledge to navigate, which is good.
3.This phenomenon says more about the complexity of our social security systems than it does about LLMs, hopefully prompting our governments to review and start thinking about social security reform.
>2. LLMs are empowering citizens to appeal in a social security landscape which often require expert-level knowledge to navigate, which is good.
Including Voter Roll Purge requests? That’s in the paper as well.
Child protective services has a tip line.
You forgot:
2b. LLMs are empowering bad actors to exploit loopholes and collect benefits they're not entitled to.
The assumption from a lot of commenters seems to be that these are all legitimate claims and appeals, not a gold rush to get money from the government while the getting's good. In the end it will be the good faith actors that suffer - both those relying on these benefits to survive who will have to jump through more hoops to prove it and those funding the system with their taxes.
Having older parents and watching what my grandparents went thru long before LLMs I just don't worry about the poor old state that much.
The political zeitgeist is to always blame the "bad people", the "welfare queens" and the "undeserving" as the problem with the system. But never the complex system that's seemingly designed to benefit things like corporations extracting as much wealth from the system as they can. Or politicians ear marking the money for things they want to do, rather than the people that put the money in the system.
So, yea, your reasoning just really doesn't hit me in the heartstrings at all.
>Having older parents and watching what my grandparents went thru long before LLMs I just don't worry about the poor old state that much.
You should worry more about your parents whose benefits are jeopardized by stress to the benefits system.
Is there any evidence that this is actually a real worry?
Don’t audits usually find that vast majority of benefits are valid and correctly applied?
>But never the complex system that's seemingly designed to benefit things like corporations extracting as much wealth from the system as they can.
What “system”? Are you talking about? Social security? How is social security designed to maximize corporate wealth extraction?
If you mean the more nebulous “system” (“the system, man”), or capitalism, etc. then you’re equivocating.
is that actually happenng?
"No, but we're reporting it is" --southpark
LLMs are just revealing how many systems rely on making applications artificially difficult, either to understand or to generate, to reduce volume. These systems have been failing silently for years, unmuting the alarm is the first step to fixing them.
It should be more about how Public Services are underfunded and the process for applying and appeals for benefits is designed to be slow to account for this underfunding, but now something has bypassed one part of the design. It should be an argument as to why we need more funding for public services.
Or, "LLMs making it increasingly hard for government to categorically deny access to public services" depending on how you want to present the issue.
To me this is one of the reasons that UBI might not be most horrible idea. Sufficiently calibrated UBI would remove this useless work on both sides. Thus making whole system more efficient and maybe even cheaper.
It's becoming obvious that essentially every public-facing organization needs to change. Short story, poetry, essay submissions, comments, reviews, conference paper and journal submissions, polling, job apps, etc.
At this point, I think it's an open question whether there's a scalable way to filter for human effort, that isn't just another layer of AI (e.g. Pangram). I'd love to see something practical come out of this, since the default is probably for public interfaces to just cease to exist, and convert to referral only systems or pure noise.
This is something I have been working on. One of the main issues my users have run into is navigation. The Gov. websites often are not very user friendly; so its hard and time consuming to navigate to and apply for the correct services. To address this I have users just speak on their "Intent" for example: "I need help with housing and food assistance, my kid has a peanut allergy and I just lost my job so I don't know what to do."
Note: This is a real request
What happens here it a custom program is created, the program has a protocol that shows exactly how the agent navigates the Gov. system and applies for the benefits. Standard modules are applied to actually do the real work of applying, followups etc...
Despite having a lot of professional experience I have not been able to even get a follow-up interview for a job so I am just pouring my efforts into this.
I have been able to help people get housed, apply for benefits, get into shelters and most recently I started working on erpo.law to speed up and reduce the paperwork on all sides for red flag laws.
Lots of the people that have helped me have done so under the table and not at the approval of their agency heads.
Here is a short vid showing a request that locates an existing program, programs can be codified into re-useable units for people w/ common needs or created in real time... Staffed via O*NET roles that hit real providers I minded from other websites.
Video: https://youtu.be/JOjG71i2F5E
Note: This one matters to me a lot because my mom was in and out of mental institutions my whole life and I been committed before so I really like helping that population because I get it more than most do.
Im also doing city agencies NYC, Boston, Miami, LA, and Vegas to start because they all have their own tld so its easier for me to filter out the bs traffic and bots and stuff.
In terms of scaling human effort the HITL for each module/protocol-chain step can be delegated where needed and the program will persist that prefrence etc...
Seems plainly obvious to me as well
The people who need this additional help to access water/power/internet/etc often have the most barriers to being able to complete the necessary steps to access benefits. This is a great use of the technology.
When an immovable object meets an unstoppable force.
Hackernews wrestles with AI slop vs. public benefits.
"Flooding" implies some nefarious purpose. There's no evidence of that in this paper. Maybe these agents are finally giving people who could never find the time, resources, knowledge, etc. to engage with the system.
"Public services are increasingly becoming more accessible with LLMs" How about that as the title?
They also have zero evidence that this is because of LLMs. They just got some data related to volume, did a really crappy analysis that wouldn't even stand up as an intro to stats course project, and declared their biased opinion as likely true.
"Flooding" could be less nefarious as in the number of requests have risen dramatically in a short period of time quite possibly overwhelming those that are tasked to deal with the requests. Most places have peaks and valleys and can staff appropriately. Having a sudden rise from something unexpected will overwhelm the normal staff. We can debate whether the lack of recognition of the upticks being intentional or not, but possibly the ignoring the uptick is the nefarious bit
>Reduce the Expected Benefit.
This could be the chance to reverse America out of becoming a welfare state.
Stop corporate welfare first.
AI destroys everything. It's such spammy trash.
This is a bad thing. It means public services will ultimately deploy AI to respond to those requests, and people will be denied instantly with no recourse or human oversight, it will be tyranny of the machine.
Was it intentionally made difficult to communicate with the government and LLMs are just not sensitive to the artificial friction?