[{"data":1,"prerenderedAt":10},["ShallowReactive",2],{"article-ai-agents-broke-welfare-state-legitimate-claims":3},{"slug":4,"title":5,"summary":6,"date":7,"published":8,"content":9},"ai-agents-broke-welfare-state-legitimate-claims","AI agents just broke the welfare state, and the applicants are real people with real claims","Researcher Chris Schmitz tracked 84 cases across 11 jurisdictions where AI tools drove surges in public service submissions: UK housing ombudsman complaints more than doubled from 2,600 to 7,000+, the US CFPB saw 5x growth, and German social courts saw a 55% year-on-year rise. The counterintuitive finding: unlike bug-bounty programs flooded with AI slop, the vast majority of new public-service filings are legitimate claims from people who were entitled to benefits but couldn't navigate the administrative burden before AI lowered the cost of applying. The policy response isn't to throttle AI-assisted filings. It's to redesign government services for agent-native interaction, because administrative burden was never a feature. It was a rate limiter, and the rate limiter just broke.","2026-09-16",true,"\u003Ch1>AI agents just broke the welfare state, and the applicants are real people with real claims\u003C/h1>\n\u003Cp>On September 10, 2026, TechCrunch reported a dataset that should reshape how every government thinks about AI. Researcher Chris Schmitz, with Lewis Hammond and Alan Chan, tracked 84 cases across 11 jurisdictions where AI tools drove surges in public service submissions. UK housing ombudsman complaints more than doubled, from 2,600 in 2022 to just over 7,000 last year. The US Consumer Financial Protection Bureau saw 5x complaint growth over the same period. German social courts attributed a 55% year-on-year rise in caseload in 2025 to AI-generated claims. Brazilian judicial petitions and German parliamentary petitions showed similar jumps. In all 84 cases, submissions were roughly flat before 2022, then accelerated as AI diffused. Most have not slowed down.\u003C/p>\n\u003Cp>The paper, posted to arXiv on August 17, 2026, and set to be presented at the 9th AAAI Conference on AI, Ethics, and Society (AIES) on October 12, calls the phenomenon &quot;agentic flooding.&quot; But the name is misleading in one critical way. Unlike the bug-bounty programs that were flooded with worthless AI-generated security reports last year, the vast majority of new public-service filings are not spam. They are legitimate claims from real people who were entitled to benefits they couldn't reach before. Schmitz told TechCrunch: &quot;The vast majority of cases we find are people who are entitled to claim for something, claiming for that thing.&quot;\u003C/p>\n\u003Cp>This is not a story about AI breaking government. It's a story about AI exposing that government services were never scaled for the populations they were supposed to serve. Administrative burden (the learning costs, compliance costs, and psychological costs of navigating bureaucracy) was the de facto rate limiter. The rate limiter just broke.\u003C/p>\n\u003Ch2>The data: 84 cases, one mechanism, one pattern\u003C/h2>\n\u003Cp>The dataset is more rigorous than its headline suggests. Schmitz, Hammond, and Chan scanned 2,288 candidate government services across 12 countries (Australia, Brazil, Denmark, Estonia, France, Germany, Japan, the Netherlands, Singapore, South Korea, the United Kingdom, and the United States). Fewer than one in twenty met all three inclusion criteria: a plausible AI mechanism, evidence of changed demand patterns, and external attribution of that change to AI by either a government official or a reputable secondary source. The 84 cases that survived are conservative: the authors explicitly note their methodology &quot;likely undercounts&quot; real flooding because it requires public attribution, and many effects are invisible when submission formats are constrained.\u003C/p>\n\u003Cp>Of those 84 cases, government officials themselves assert AI involvement in 58 (69%); the remaining 26 are attributed by third-party sources. The most represented domain is Justice and Legal Services (23%, 19 cases), followed by Regulatory Complaints (12%, 10 cases) and Benefits and Social Protection (11%, 9 cases). The mechanism is overwhelmingly consistent: 87% of cases involve LLMs generating legally sophisticated text that users submit manually. Browser-navigation agents and fully autonomous submission pipelines are not yet evident. People are pasting a photo of a letter into Claude and getting a complaint back.\u003C/p>\n\u003Cp>The flooding has two dimensions. Quantitative flooding, more requests, appears in 60% of cases. Qualitative flooding, longer and more complex requests, appears in 90%. Half exhibit both. In one German social court case, some letters spanned over 4,000 pages. The BBC reported that UK complaints about trash collection can now run over 20 pages. The volume is the visible problem. The complexity is the operational one: a 20-page complaint takes longer to read than a 200-word one, regardless of whether the claim is valid.\u003C/p>\n\u003Ch2>The counterintuitive finding: legitimate demand, not spam\u003C/h2>\n\u003Cp>The sharpest finding in the paper is not the volume surge. It's who's behind it.\u003C/p>\n\u003Cp>When bug-bounty programs were flooded with AI-generated reports last year, the submissions were overwhelmingly worthless: low-quality text that rarely contained real security issues but still required vetting. The natural assumption is that public services would face the same pattern: AI slop, mass-produced nonsense, adversarial filings designed to overwhelm.\u003C/p>\n\u003Cp>Schmitz's data says the opposite. Most new filings are from people who were already entitled to what they're claiming. They just couldn't get through the process before. The paper frames this through the concept of &quot;administrative burden&quot;: the learning costs of understanding complex rules, the compliance costs of filling forms and assembling documents, and the psychological costs of repeated dehumanizing bureaucratic interactions. Moynihan et al. (2015) identified these three burden types in the public administration literature. AI agents reduce all three.\u003C/p>\n\u003Cp>This is the key insight: administrative burden was never a feature of government services. It was a bug. But it was a bug that served as an unintentional rate limiter, suppressing demand to a level the system could handle. The UK Department for Work and Pensions explicitly anchors benefit budgets to historical take-up rates: rates that were low because the process was forbidding. When AI removes the forbiddingness, the latent demand surfaces. The system was never designed for the population it was supposed to serve. It was designed for the fraction that could navigate it.\u003C/p>\n\u003Cp>The paper documents some adversarial cases: individuals attempting to get disability benefits with AI-generated fake medical certificates (Case D, Brazil), and organized campaigns mass-submitting FOI requests or voter roll challenges (Case A, Australia; Case J, United States). But these are the minority. The majority is a single mother who just found out she's been entitled to housing support for two years and can now write the complaint that gets her case reviewed.\u003C/p>\n\u003Ch2>The enterprise parallel: permissions wall, inverted\u003C/h2>\n\u003Cp>In June 2026, we wrote that AI agents aren't hitting a model wall. They're hitting a permissions wall. The enterprise problem was agents doing too much without oversight: pushing to production on automated evals, spending money without kill switches, sharing credentials that make forensics impossible. The enterprise risk was agents exceeding their authority.\u003C/p>\n\u003Cp>The public-sector problem is the inverse. The risk isn't agents doing too much. It's agents revealing that government services were doing too little. The permissions wall in enterprise is a control problem: how do you constrain agents who are too capable. The permissions wall in government is an access problem: how do you serve citizens who were always entitled but never reached.\u003C/p>\n\u003Cp>Both problems share a root cause: systems designed for a level of demand that assumed friction would persist. Enterprises assumed humans would review every agent action. Governments assumed administrative burden would suppress take-up. AI broke both assumptions simultaneously. The enterprise response is building control planes: kill switches, budget governors, scoped identities, sandboxing. The government response is harder, because the equivalent of a kill switch is a fee, and fees disproportionately deter the poorest and least digitally literate, the exact people the system was supposed to serve.\u003C/p>\n\u003Ch2>The policy trap: friction is fast, redesign is slow\u003C/h2>\n\u003Cp>The paper maps government responses into two categories: suppress demand or increase capacity. Both work. Both have trade-offs. The structural problem is timing.\u003C/p>\n\u003Cp>Demand suppression (fees, rate limits, in-person requirements, bot-blocking) is fast. Governments have already used friction in 17% of cases in the dataset. Australia considered reintroducing FOI fees. Japanese authorities blocked IP addresses from a public comment procedure. These measures deploy in days and can stop a surge. But they decrease the accessibility of government services, disproportionately deter vulnerable users, and can create procedural inequality. In many jurisdictions, charging fees for welfare applications is simply illegal. The paper's authors note that friction-inducing measures may also become less effective as agent capabilities improve: CAPTCHAs no longer reliably identify human visitors, and the arms race between bot-detection and bot-capability is running in the wrong direction.\u003C/p>\n\u003Cp>Capacity-building (service redesign, digital identity integration, AI-assisted processing) is slow but doesn't trade off access. Governments deployed AI tools to help with processing in 25% of cases: analyzing sentiment across public comments, detecting fake medical certificates, improving triage. But structural redesign (replacing PDF forms with structured APIs, making eligibility rules machine-readable, building digital identity into the most exposed services) requires cross-agency coordination, legislative change, and lead times measured in years. Many Western governments have taken decades to get basic digital infrastructure in place.\u003C/p>\n\u003Cp>The paper's warning is precise: &quot;One plausible trajectory is therefore that demand suppression again becomes a default response. Governments might reach for it because, when a surge occurs, it is the only intervention available on a short enough timeline.&quot; This pattern, what Lindblom (1959) called &quot;muddling through&quot;, keeps organizations operational but doesn't optimize outcomes for the public. The threat is not that governments will choose the wrong response. It's that they will choose the only response available in time, and the one available in time is the one that hurts the people the system exists to serve.\u003C/p>\n\u003Ch2>What governments should do\u003C/h2>\n\u003Cp>Schmitz, Hammond, and Chan propose three near-term actions. They are not revolutionary. They are the minimum viable preparation.\u003C/p>\n\u003Cp>\u003Cstrong>1. Audit service exposure.\u003C/strong> Governments should systematically map their services against the risk matrix the paper provides. Which services accept free-form text through open digital channels? Which services are financially attractive to claimants? Which services have historically relied on friction (complex forms, legal knowledge requirements, procedural difficulty) to suppress demand? Those are the services most exposed to flooding. The assessment can be done today. It does not require forecasting AI progress.\u003C/p>\n\u003Cp>\u003Cstrong>2. Develop a digital identity strategy.\u003C/strong> Strong identity verification allows per-claimant rate limits (stopping quantitative flooding) without blocking legitimate access. It also enables pre-population of known data, lowering the effort for entitled users to apply and for governments to process their cases. Over 100 jurisdictions already have digital identity infrastructure. Its maturity varies drastically. Integrating it into the most exposed services is the highest-leverage near-term action.\u003C/p>\n\u003Cp>\u003Cstrong>3. Establish legal certainty.\u003C/strong> Several response options (restricting digital submission channels, raising fees, deploying AI in automated decision-making) are either illegal or of unclear legality in many jurisdictions. This ambiguity is itself a barrier to preventative action. Governments should commission internal legal reviews now, before a surge forces a decision under pressure.\u003C/p>\n\u003Cp>These three actions share a property: they are all things governments can do before the surge, not during it. The paper's central structural argument is that the window for capacity-building is now, while the surge is moderate. Once the surge is acute, friction is the only tool available on the timeline the crisis demands, and friction is the tool that hurts the people the system serves.\u003C/p>\n\u003Ch2>The 18-month outlook\u003C/h2>\n\u003Cp>The bet is this. Agentic flooding is a leading indicator, not a final state. The paper finds that 87% of current flooding comes from LLM text generation, the most basic agent capability. Browser-navigation agents, autonomous submission pipelines, and multi-step tool-using agents are not yet driving the surge. When they do, the services that survived the text-generation wave on the strength of complex login flows or JavaScript-heavy portals will lose that defense. The paper's risk matrix identifies portal complexity as a factor that currently limits flooding (Factor 3). That factor is eroding.\u003C/p>\n\u003Cp>The governments that treat agentic flooding as an infrastructure problem (structured APIs instead of PDF forms, machine-readable eligibility rules instead of prose guidelines, digital identity baked into every exposed service) will be the ones that absorb the next wave without breaking. The governments that treat it as a spam problem will reach for friction, suppress the surge, and pat themselves on the back while the people they were supposed to serve go back to not claiming what they're owed.\u003C/p>\n\u003Cp>The paper stops short of saying AI is directly causing the surge, for methodological reasons. But the pattern is clear enough: flat before 2022, accelerating after, no slowdown in sight. The question is not whether the surge will continue. It's whether governments will use the current window, when the flooding is moderate, the mechanism is simple, and the legitimate claims are still recognizable as legitimate, to build the infrastructure that makes the next wave survivable. The window is open. It is closing.\u003C/p>\n\u003Chr>\n\u003Cp>\u003Cstrong>Sources:\u003C/strong> Chris Schmitz, Lewis Hammond, and Alan Chan, &quot;Characterizing Agentic Flooding of Government Services&quot; (arXiv:2608.16603, August 2026; to appear in proceedings of the 9th AAAI Conference on AI, Ethics, and Society, October 12-14, 2026); TechCrunch, &quot;AI agents are flooding public services with new requests&quot; (Russell Brandom, September 10, 2026); Cybernews, &quot;The public are flooding their local services with AI slop complaints&quot; (September 11, 2026); BBC, UK housing complaints coverage (2026); Moynihan et al., &quot;Administrative Burden: Rationale, Interpretation, and Applications&quot; (Journal of Public Administration Research and Theory, 2015); Lindblom, &quot;The Science of Muddling Through&quot; (Public Administration Review, 1959); dataset at github.com/CLSchmitz/flooding-dataset.\u003C/p>\n",1790928907315]