THE INDEPENDENT RECORD · AGENTIC AI AS A SERVICE AboutStandardsContact
GAASAGENTIC AI · AS A SERVICE
INDEPENDENT · SINCE 2026
UPDATED DAILY
NO HYPE · NO PAY-TO-PLAY
PER-TASK PRICING NOW STANDARD ● NEW BENCHMARK: 71% TASK COMPLETION ● ENTERPRISE PILOTS UP 4X ● RUNTIME FUNDING ACCELERATES ● "AGENTS ARE THE NEW SEATS" ● MARGINS UNDER PRESSURE ● THE INDEPENDENT RECORD ON GAAS
Trust & Safety

Who's Liable When an AI Agent Makes a Costly Mistake?

When an autonomous agent wires money to the wrong vendor, deletes a production database, or promises a customer a refund it shouldn't have, the bill is real and someone has to pay it. The short answer: liability almost never lands where intuition expects. Contracts, not common sense, decide it, and most Agentic AI-as-a-Service (GaaS) agreements are deliberately written to push the loss onto the buyer. This piece breaks down who actually pays, why the "the agent did it" defense fails in court, and the specific contract terms and controls that shift the odds back in your favor.

By A. Reyes · Mar 7, 2026 · 12 min read

Table of Contents

The Question That Keeps Buyers Up at Night

A procurement agent at a mid-sized distributor was given authority to reorder stock when inventory dipped below threshold. It read a supplier email, parsed a price as $4.20 per unit instead of $420, and placed an order for 100,000 units it interpreted as a bargain. By the time a human noticed, three purchase orders had gone out and one supplier had already started a production run.

Who eats that cost?

This is no longer a hypothetical. As GaaS vendors sell agents that take real actions in real systems with real money attached, the gap between "the demo worked" and "the agent cost us $80,000" is where the lawyers live. And here's the uncomfortable truth that most vendor sales decks gloss over: the legal system has no special category for "an AI did it." Liability flows through existing doctrines of contract, agency, negligence, and product law, and those doctrines were built for a world of humans and tools, not autonomous software that decides on its own.

That mismatch is the whole story. Understanding it is the difference between a recoverable loss and an uninsured one.

There's a seductive intuition that autonomy creates a liability vacuum: if no human chose the action, no human is responsible. Courts have shown zero appetite for this.

In the law of agency, when you authorize someone (or something) to act on your behalf, you generally own the consequences of acts within that authority. A company that deploys an agent with the power to issue refunds has, in effect, told the world this thing speaks for us. When it issues a bad refund, the company is bound by it, the same way a business is bound by a junior employee who had apparent authority to cut a deal. The classic illustration came years before modern agents: in Moffatt v. Air Canada (2024), a Canadian tribunal held the airline responsible for a refund policy its customer-service chatbot invented out of thin air. The airline argued the bot was a "separate legal entity." The tribunal called that "remarkable" and made the company pay. The reasoning generalizes cleanly to far more capable agents.

The deeper principle is that liability attaches to the party who controls the deployment and benefits from it, not to the artifact. A hammer that slips doesn't owe you anything; the contractor swinging it might. An agent is a more sophisticated tool, but it's still, legally, a tool deployed by parties who chose to deploy it. The interesting fights are about which of those parties, and most GaaS arrangements involve at least three.

The Four Parties Who Could Be on the Hook

When an agent causes a costly mistake, point the finger at four candidates, and the real answer is usually a blend.

The deploying business (the buyer). This is the default loss-bearer. You configured the agent, granted it credentials and permissions, pointed it at your systems, and captured the upside when it worked. Tort and agency law tend to land here first. If you gave an agent write-access to your payment rails with no approval gate, a court will struggle to see the vendor as the proximate cause of a bad wire.

The GaaS vendor (the agent provider). The vendor built the orchestration, the guardrails, the tool-calling logic. If the harm traces to a defect in their product, a planner that ignored a hard constraint, a safety filter that silently failed, product-liability and breach-of-warranty theories come into play. But, as we'll see, vendors work very hard contractually to make sure this door stays shut.

The foundation-model provider. The LLM underneath (OpenAI, Anthropic, Google, etc.) is the reasoning engine, and its terms of service almost universally disclaim responsibility for outputs and bar high-stakes autonomous use without human oversight. Reaching this party is extremely hard; their contracts are with the GaaS vendor, not your business, so you usually have no direct claim against them at all.

The end user or third party. Sometimes the "mistake" is provoked, an adversarial user who jailbroke a customer-facing agent into giving away inventory, a topic explored more fully in #189: Adversarial users: jailbreaking customer-facing agents. Their conduct can shift or share blame, but it rarely lets the deploying business fully off the hook for putting an exploitable agent in front of the public.

The reason apportionment is so contested is that the chain of causation is genuinely murky. Did the model hallucinate, did the vendor's guardrail fail to catch it, or did the buyer over-permission the agent? Often all three. This is the same evidentiary fog that makes #208: The forensic challenge of investigating an agent's decision such a hard problem, without good logs, nobody can prove whose fault it was, and "can't prove it" usually defaults to "buyer pays."

What GaaS Contracts Actually Say About Liability

Read the fine print of nearly any GaaS agreement and a consistent architecture emerges. It is built to protect the vendor.

Liability caps. Almost every contract caps the vendor's total liability at the fees paid in the trailing 12 months, sometimes just 3 months. If you pay $2,000/month for an agent that causes a $200,000 loss, the vendor's maximum exposure under a 12-month cap is $24,000, and a per-month cap could limit it to $2,000. The economics are stark: vendors price the service assuming they will almost never pay for downstream harm.

Disclaimers of warranty. Output is provided "as is." The vendor warrants the service will function, not that the agent's decisions will be correct. This is a crucial distinction. "The agent ran" is the promise; "the agent was right" almost never is.

"Human in the loop" carve-outs. Many agreements condition any vendor responsibility on the buyer maintaining human review of consequential actions. The moment you let the agent act fully autonomously on something costly, you've often contractually assumed that risk yourself. This is why #195: The "human accountable owner" requirement for every agent is becoming a governance standard, it's also a liability-allocation mechanism hiding in plain sight.

Indemnification, and its limits. Vendors will often indemnify you for IP infringement (their training data, their model) but pointedly not for the consequences of the agent's actions in your business. Buyers routinely misread an IP indemnity as broad protection. It isn't.

For a grounded view of how enterprises are starting to push back on these terms, the analysis in Gartner's guidance on AI agent risk and governance is a useful baseline, and the legal community's framing in the Harvard Business Review's coverage of accountability in AI systems maps the doctrinal gaps cleanly.

The Per-Outcome Pricing Trap

Here's an angle the GaaS hype cycle conveniently underplays. The industry is moving toward per-outcome and per-task pricing, you pay only when the agent successfully closes a ticket, books a meeting, recovers a payment. It's marketed as aligned incentives: the vendor only wins when you win.

But pricing for outcomes does not mean owning outcomes. A vendor can happily charge you per resolved support ticket while its contract simultaneously disclaims all responsibility for tickets resolved wrongly. The pricing model implies the vendor is on the hook for results; the liability terms quietly say otherwise. That gap is where buyers get burned. If you're paying for outcomes, your contract should make the vendor share in the cost of bad outcomes too, and almost none do by default. The economics of this asymmetry connect directly to themes in the broader GaaS cluster around agent reliability and the true cost of autonomy.

How Courts Will Likely Apportion Blame

We're early, and case law is thin, but the direction is becoming legible. A few patterns are worth betting on.

Comparative fault will dominate. Rather than one party bearing 100%, expect courts to split liability by control and foreseeability. The buyer who removed every guardrail bears more. The vendor whose documented safety claim demonstrably failed bears more. The split will turn heavily on evidence, which is exactly why audit logs, the subject of #168: Audit logs regulators will demand from GaaS vendors, are becoming a frontline liability issue rather than a compliance afterthought.

Foreseeability is the hinge. If a harm was a known failure mode, agents are known to hallucinate, known to be vulnerable to prompt injection, then deploying one in a high-stakes path without controls looks like negligence on the buyer's side. "We didn't know it could do that" is a weak defense when the entire industry has been publicly documenting that exact failure.

Regulators will move faster than courts. The EU AI Act's risk-tiered obligations impose duties on providers and deployers of higher-risk systems, and they assign responsibilities by role regardless of how the private contract allocates them. You cannot fully contract out of a statutory duty. A liability waiver that's valid between two companies may be irrelevant to a regulator levying a fine.

The practical upshot: the party with the best records and the clearest division of responsibility wins the apportionment fight. The party that can't reconstruct what happened loses by default.

What Buyers Should Negotiate Before Signing

Liability is allocated at contract-signing, not at incident-time. By the time something breaks, your leverage is gone. So the work happens up front.

A few moves consistently change the risk picture:

None of this is exotic. It's the same diligence enterprises apply to any high-consequence vendor, applied to a technology whose failure modes are newer and weirder than most procurement teams are used to.

Insights Most People Overlook

The liability cap is usually smaller than a single bad action. Buyers anchor on monthly price and never multiply the worst-case action by the volume the agent operates at. An agent making 5,000 decisions a day at a $24,000 liability cap is, in expectation, almost entirely your risk to absorb. The math is rarely run before signing, and it should be the first thing you run.

Autonomy and liability are inversely correlated for the buyer's benefit only on paper. The more autonomous the agent, the more the vendor markets it as hands-off, and the more their contract treats your hands-off operation as your assumption of risk. The very feature you're paying a premium for (less human oversight) is the feature that contractually transfers the most liability to you. Read those two clauses side by side and the trade becomes obvious.

"Per-outcome pricing" is a liability tell, not just a pricing model. When a vendor is willing to be paid only on success, ask why they still won't share the cost of failure. A vendor genuinely confident in its agent should be willing to put at least some skin in the downside. Their unwillingness is the most honest signal you'll get about how reliable they actually think the agent is.

The foundation-model provider is effectively judgment-proof to you. Because you have no contract with the LLM vendor, their disclaimers don't even have to win an argument, there's no argument to be had. This collapses the realistic field of defendants down to you and the GaaS vendor, which is precisely why the GaaS contract terms matter so much more than the model's terms.

Good logging is a financial instrument. Most teams file audit logging under "compliance" or "debugging." In a comparative-fault world, your logs are the asset that determines whether you recover 80% of a loss or 0%. Underfunding observability is, quite literally, declining to insure yourself.

References

#autonomous agent accountability

More in Trust & Safety