THE INDEPENDENT RECORD · AGENTIC AI AS A SERVICE AboutStandardsContact
GAASAGENTIC AI · AS A SERVICE
INDEPENDENT · SINCE 2026
UPDATED DAILY
NO HYPE · NO PAY-TO-PLAY
PER-TASK PRICING NOW STANDARD ● NEW BENCHMARK: 71% TASK COMPLETION ● ENTERPRISE PILOTS UP 4X ● RUNTIME FUNDING ACCELERATES ● "AGENTS ARE THE NEW SEATS" ● MARGINS UNDER PRESSURE ● THE INDEPENDENT RECORD ON GAAS
Verticals

Insurance-Claims Agents: How Adjudication Automation Actually Works (and Where It Breaks)

Insurance-claims agents are autonomous AI systems that read a claim, pull policy and medical or repair data, apply coverage rules, and either pay, deny, or route it to a human. The promising versions touch low-dollar, high-volume claims first, often priced per adjudicated claim or per dollar of leakage recovered. The hard part is not reading documents anymore, it's the legal and regulatory liability of a wrong denial, which is why the best deployments keep a human on the deny path and let the agent own only the clean approvals. Adoption is real but uneven: auto and dental are ahead, complex health and liability are years out.

By J. Okafor · Mar 12, 2026 · 13 min read

Table of Contents

What an Insurance-Claims Agent Actually Does

Strip away the marketing and a claims-adjudication agent is doing what a junior adjuster does on a slow Tuesday: it takes a submitted claim, figures out whether the policy covers it, checks the numbers, and decides what gets paid. The difference is that the agent does it in eleven seconds instead of nine days, and it does it the same way at 2 a.m. as it does at 2 p.m.

That sameness matters more than the speed. Anyone who has worked a claims floor knows that adjudication consistency is a quiet scandal. Two adjusters looking at the same dental claim, or the same fender-bender estimate, will reach different conclusions more often than carriers like to admit. An agent doesn't have a bad morning. It doesn't get a backlog and start rubber-stamping to clear the queue. Whatever bias it has, it has uniformly, which is its own problem, but a different one.

This is a vertical agent in the truest sense, and it belongs to the broader story this cluster keeps returning to: software that doesn't just assist a workflow but owns an outcome. A horizontal LLM wrapper can summarize a claim. A claims agent has to decide it, log why, and stand behind the decision when a regulator or a plaintiff's attorney comes asking. That gap, between summarizing and deciding, is the whole ballgame.

The Adjudication Pipeline, Step by Step

Adjudication isn't one task. It's a chain, and each link is a different kind of problem for an AI system.

Intake and extraction. The claim arrives, a faxed CMS-1500, a photo of a damaged bumper, a PDF estimate, an EDI 837 file. The agent has to turn that mess into structured data: procedure codes, diagnosis codes, dates of service, billed amounts, provider IDs. Two years ago this was the bottleneck. It mostly isn't anymore. Modern document models handle messy real-world claim documents well enough that extraction is no longer where projects die.

Policy and eligibility matching. Now the agent pulls the member's policy, confirms coverage was active on the date of service, and identifies which benefits apply. This sounds trivial and is not. Policy data lives in core systems built in the 1980s, and the agent's reliability is capped by how cleanly it can reach that system of record.

Rules application. This is the heart of adjudication, applying the coverage rules, fee schedules, bundling logic, medical-necessity criteria, and exclusions. A lot of this is deterministic and was already automated by old rules engines. The agent's value is handling the cases the rules engine kicks out as exceptions: the gray-area judgment calls that used to need a person.

Pricing and payment determination. The agent computes the allowed amount, applies the deductible and coinsurance, nets out coordination-of-benefits when there's a second payer, and lands on a number.

Disposition. Pay, deny, partially pay, or pend for review. The agent's confidence and the dollar amount jointly decide whether it acts alone or hands off. A clean $80 in-network claim it pays itself. A $140,000 inpatient stay with a medical-necessity question goes to a human every time, with the agent's analysis attached as a head start.

The smart architecture isn't "agent does everything." It's the agent owning the 60-70% of claims that are genuinely routine, doing them flawlessly and instantly, and surfacing the hard ones to humans who now have time to actually look at them.

Where Agents Are Winning First

Follow the money, but follow it to the low-stakes money. Adjudication automation lands first where claims are high-volume, low-dollar, and rule-dense.

Dental is arguably the cleanest beachhead. Procedures are highly standardized (the CDT code set is finite and well-behaved), claims are small, and the adjudication rules are explicit. A wrong call on a $120 filling is a fixable annoyance, not a lawsuit.

Auto physical damage is close behind. Photo-based estimating plus repair-cost databases let an agent look at images of a dented door and produce a defensible estimate. The auto carriers have invested heavily here because the volume is enormous and the per-claim economics are punishing when done by hand.

Pharmacy and vision claims share the dental profile: standardized, bounded, low-dollar.

Complex health and liability are the laggards, and rightly so. A disputed long-term-disability claim or a bodily-injury liability claim involves judgment, negotiation, and legal exposure that no carrier's general counsel is signing off on for full automation in 2026. McKinsey's work on claims describes a "touchless" future for the simple end of the book while the complex end stays firmly human-led, a split that maps almost exactly onto where agents work and where they don't. Their analysis of the future of claims in insurance is blunt that the value is concentrated in the high-frequency, low-severity segment.

This pattern, automate the boring 70%, escalate the consequential 30%, recurs across every vertical-agent category. It's the same shape you see in prior-authorization agents and in medical-coding agents, which is why those topics rhyme so closely with this one.

The Liability Wall: Why Denials Are Different

Here's the thing the vendor demos gloss over: approving a claim and denying a claim are not symmetric acts.

When an agent approves a clean claim, the worst case is the carrier paid something it might have contested. That's leakage, a cost-control problem, recoverable, unglamorous. When an agent denies a claim, the worst case is a sick person doesn't get care, a family doesn't get a payout after a fire, and a regulator or a class-action firm asks the carrier to explain, in writing, under oath, exactly how a machine decided.

Regulators have noticed. Several state insurance departments and the NAIC have moved toward requiring that AI-driven adverse determinations be explainable and that a qualified human remain accountable for the decision. The NAIC's model bulletin on the use of AI by insurers effectively tells carriers they own their algorithms' outcomes, vendor or not. In health specifically, the 2023-2024 litigation over algorithmic denials at large insurers put a spotlight on the danger of letting a model auto-deny at scale without meaningful human review.

So the durable design pattern is asymmetric autonomy: let the agent approve autonomously, but never let it deny autonomously. On the deny path, the agent does the analysis, drafts the denial rationale with citations to specific policy language, and hands a human a decision they can review in 90 seconds instead of researching for an hour. The human still pulls the trigger. The agent makes them ten times faster without making them legally invisible.

Carriers that ignore this and let agents auto-deny to juice their automation rate are building a liability time bomb. The economics of being wrong on a denial, bad-faith damages, regulatory fines, reputational ruin, dwarf the labor savings.

How These Agents Get Priced

The GaaS pricing question is sharper here than in most verticals, because in claims you can measure the outcome precisely.

Per-claim pricing is the common entry point, a few cents to a few dollars per adjudicated claim, depending on complexity. It's easy for a carrier to model against the loaded cost of a human adjuster's time.

Per-outcome / value-share pricing is where it gets interesting and where the leverage is. Claims have a built-in scoreboard: leakage. If an agent catches overpayments, duplicate billings, and miscoded claims that humans were missing, the vendor can charge a percentage of recovered dollars. This aligns incentives beautifully, and it's hard for a carrier to argue with a bill that's a slice of money it would otherwise have lost. Expect more vendors to push toward this model as they build confidence in their detection rates.

Per-seat licensing still shows up for the human-in-the-loop console, but it's the weakest model, it prices the tool, not the work, and carriers increasingly want to pay for the work.

The strategic tension: per-claim pricing is legible and easy to sell, but it caps the vendor's upside and invites commoditization. Value-share pricing captures more but requires the vendor to prove a clean counterfactual ("you'd have lost this money without us"), which is harder than it sounds when the baseline is a noisy human process. The pricing of vertical agents by industry-specific value capture is a recurring theme across this cluster, and claims is one of the few verticals where the value is genuinely countable.

The Build-vs-Buy Question for Carriers

Big carriers will be tempted to build. They have the data, the engineers, and the not-invented-here instinct. Most of them should buy, and here's the honest reasoning.

The hard part of a claims agent isn't the model, it's the thousands of edge cases, the connectors into ancient core systems, the regulatory documentation, and the accumulated knowledge of how this specific kind of claim actually adjudicates. A vendor selling to forty carriers sees forty times the edge cases and folds those lessons into the product. That proprietary workflow data is the moat, and a single carrier building in-house can't accumulate it as fast.

The counterargument: claims data is sensitive, core systems are bespoke, and a carrier's adjudication rules are arguably a competitive asset it doesn't want to hand to a vendor that also serves rivals. Andreessen Horowitz has argued that the defensibility of AI applications increasingly comes from owning the system of record and the proprietary workflow, not the model, which cuts both ways: it's exactly why vendors want this data, and exactly why some carriers will refuse to share it.

The likely equilibrium: small and mid-size carriers buy, the largest carriers build for their core lines and buy for the long tail, and a healthy vendor ecosystem grows in the middle. The services-to-software flip is happening here too, claims-processing BPO shops that used to sell labor are quietly turning themselves into agent companies that sell outcomes.

Reliability, Audit Trails, and Regulator Trust

No claims agent survives contact with a regulator unless it can answer one question: why did you decide this?

That makes the audit trail a first-class feature, not an afterthought. Every autonomous decision needs to log the inputs it saw, the policy provisions it applied, the rule it triggered, and the confidence it held, in a form a compliance officer or examiner can read. "The model said so" is not a defensible adjudication rationale. The agents that win in regulated insurance are the ones that produce a denial letter a human could have written, with citations to the actual policy language, every time.

Reliability targets are also unusually concrete. A 95%-accurate agent sounds great until you remember a mid-size health plan adjudicates millions of claims a year, 5% errors is a staggering absolute number of wrong payments and wrong denials. So the real metric isn't raw accuracy; it's calibrated confidence: the agent has to know what it doesn't know and escalate accordingly. An agent that's right 99% of the time on the claims it acts on, and honestly flags the rest, beats a cockier agent that's right 96% of the time across the board. Knowing when to ask for help is the entire safety story for vertical agents in regulated industries, and claims is the sharpest test of it.

Insights Most People Overlook

The approve path, not the deny path, is where the money quietly leaks. Everyone obsesses over wrongful denials because they make headlines and lawsuits. But the larger dollar drain at most carriers is overpayment on the approve side, claims paid that shouldn't have been, or paid too much. An agent that's conservative about denials but sharp about catching overpayments before they go out the door can deliver more financial value with far less regulatory risk. The PR-safe automation play is also the profitable one.

Consistency can be a liability, not just a feature. A human-staffed claims operation has random error scattered across thousands of adjusters. An agent has systematic error, if it's wrong about a coverage interpretation, it's wrong the same way on every single claim. That turns a scattered annoyance into a clean class-action pattern. Uniform decision-making is a selling point right up until the uniform decision is uniformly wrong, at which point it's plaintiff's-attorney gold.

The real bottleneck was never document-reading. Vendors still pitch extraction accuracy as the headline. That war is largely over. The actual constraint now is integration depth into the carrier's core system of record, the 1980s mainframe that holds the policy data. The agent that reads documents perfectly but can't reliably pull live eligibility is useless. Depth of integration, not model quality, decides who wins.

Per-claim pricing will commoditize; value-share pricing won't. Charging per adjudicated claim is a race to the bottom because any competent vendor can do it and the price floor is just compute. The defensible model is charging a slice of recovered leakage, but it requires proving a counterfactual against a noisy human baseline, which most vendors aren't yet rigorous enough to do. The first vendor that nails clean attribution of recovered dollars gets to charge like a partner instead of a tool.

"Touchless" is a vanity metric carriers will learn to distrust. Boards love a high straight-through-processing rate. But an automation rate optimized in isolation pushes operators to let agents auto-decide marginal claims they shouldn't. The healthier metric pairs automation rate with reversal rate on appeal, how often an agent's autonomous decision gets overturned when a human finally looks. A high automation rate with a creeping reversal rate isn't efficiency; it's deferred liability.

References

More in Verticals