THE INDEPENDENT RECORD · AGENTIC AI AS A SERVICE AboutStandardsContact
GAASAGENTIC AI · AS A SERVICE
INDEPENDENT · SINCE 2026
UPDATED DAILY
NO HYPE · NO PAY-TO-PLAY
PER-TASK PRICING NOW STANDARD ● NEW BENCHMARK: 71% TASK COMPLETION ● ENTERPRISE PILOTS UP 4X ● RUNTIME FUNDING ACCELERATES ● "AGENTS ARE THE NEW SEATS" ● MARGINS UNDER PRESSURE ● THE INDEPENDENT RECORD ON GAAS
Pricing

Outcome-Based Pricing for AI Agents: Who Actually Defines and Audits the Outcome?

Outcome-based pricing sounds clean, you pay only when the agent delivers, but the whole model lives or dies on two unglamorous questions: who gets to define what counts as a delivered outcome, and who audits whether it actually happened. In practice the vendor usually writes the definition, the vendor usually owns the measurement, and the buyer is left trusting a scoreboard kept by the team being paid. This piece breaks down how outcome definitions get gamed, how attribution disputes blow up renewals, and the contractual and instrumentation moves that make outcome pricing survivable for both sides. If you can't independently verify the outcome, you don't have outcome pricing, you have a per-unit invoice with better marketing.

By L. Karlsson · Feb 20, 2026 · 14 min read

Table of Contents

The Quiet Power Struggle Inside Outcome Pricing

Every vendor selling agentic AI right now wants to talk about outcomes. "We only charge when it works." "Pay per resolution." "Per-outcome pricing aligns our incentives with yours." It's a great pitch, and on a slide it's nearly unbeatable. The buyer hears: zero risk. The vendor hears: I get paid for value, not for tokens I burn.

But sit in on the third quarterly business review of an outcome-priced contract and the mood is different. The customer's ops lead is squinting at a dashboard, asking why the agent claimed 4,200 "resolved tickets" last month when the support backlog barely moved. The vendor's account manager is explaining that a resolution is counted when the customer doesn't reply within 24 hours. The customer points out that half those "resolved" users just gave up and called the phone line. Nobody is lying. They're working from two different definitions of the same word, and one of those definitions is the one being billed.

This is the heart of outcome-based pricing, and it's almost never discussed in the breathless write-ups. The pricing model isn't really about price. It's about definitional authority and measurement control, and those two things, far more than the dollar amount, determine whether the deal is fair. Outcome pricing is the most buyer-friendly model on paper and one of the most contestable in operation. Understanding why is the difference between a renewal and a lawsuit.

What Counts as an Outcome? The Definition Problem

Start with the word itself. An "outcome" is not a fact sitting in a database. It's a judgment call wearing the costume of a metric.

Take a few real GaaS categories:

Notice the pattern. In every case, there's a narrow, easy-to-trigger definition that favors the vendor and a strict, value-aligned definition that favors the buyer. The gap between them is the entire negotiation. The narrowest defensible definition a vendor can get away with is, in effect, the vendor's margin lever, it's the per-outcome equivalent of the counting tricks that turn a metric into a target and then ruin it. Goodhart's Law isn't a footnote here; it's the central risk. The moment "resolution" becomes the billing event, the agent gets optimized to produce the appearance of resolution, not resolution.

And agents are extraordinarily good at optimizing for whatever you measure. That's the whole point of them. Which means a sloppy outcome definition doesn't just mislead the invoice, it actively reshapes the agent's behavior toward the cheapest thing that trips the meter.

Who Defines It in Practice

Here's the uncomfortable answer: the vendor defines the outcome, almost every time.

This isn't villainy, it's gravity. The vendor wrote the agent. The vendor instrumented the events. The vendor's data pipeline emits the "resolution" signal. By the time the contract is drafted, the definition already exists in code, and the legal language is reverse-engineered to match what the system already measures. The buyer's procurement team, which may be brilliant at negotiating per-seat SaaS, often lacks the technical fluency to interrogate how the resolution event fires. They negotiate the price-per-outcome and wave through the definition-of-outcome, which is exactly backwards.

The leverage asymmetry is structural. Analysts at firms like Gartner have flagged that outcome-based commercial models shift measurement burden onto buyers who are rarely equipped to carry it, and the agent wave makes this sharper because the "product" is a decision process the buyer can't easily inspect.

A few forces push back on vendor control, and smart buyers lean on all of them:

The thread connecting all three: move the definitional authority into a system the buyer controls. Whoever owns the database of record owns the truth, and therefore owns the invoice.

The Attribution Trap

Even with a clean definition, outcome pricing has a second landmine that's arguably harder: attribution. Did the agent cause the outcome, or did it merely happen to be standing nearby when the outcome occurred?

This is the question that quietly destroys "revenue recovered" and "revenue generated" pricing. A collections agent emails a patient; the patient pays. Would they have paid anyway? A sales agent nurtures a lead that converts; but marketing also retargeted them, an SDR called, and the prospect was already in a buying cycle. Charging a percentage of "influenced" revenue invites a fight over a counterfactual that nobody can observe.

The cleanest defense against the attribution trap is the same one used in serious marketing measurement: hold-outs and incrementality testing. You carve out a randomized control group the agent never touches, and you bill only on the lift over that baseline. This is the logic behind incrementality and geo-experiments in measuring true marketing causality, and it's just as applicable to agent outcomes. If the vendor refuses any form of hold-out, that's a tell: they're billing on correlation and hoping you don't notice.

Most vendors hate hold-outs because they shrink the billable number. That tension is healthy. A vendor confident in their agent's incremental value should welcome the test; a vendor billing on ambient activity will resist it. The buyer's job is to make incrementality non-negotiable for any outcome defined in revenue terms. For outcomes defined in operational terms (a ticket resolved, a document processed), attribution is cleaner, the agent either did the unit of work or it didn't, which is one underrated reason operational outcomes make better billing events than financial ones. This connects directly to the broader debate over the currency of value: time saved, revenue gained, or cost cut, the currency you pick determines how nasty your attribution fight will be.

Who Audits the Outcome

So the outcome is defined. Now: who checks the count?

In the overwhelming majority of GaaS contracts today, the answer is the vendor. The vendor's dashboard reports the number, the vendor's billing system invoices it, and the customer reconciles against... the same dashboard. It's a closed loop. The party being paid is also the party keeping score and the party grading their own homework. No other mature B2B category accepts this arrangement for variable spend. Imagine a law firm that billed by the hour and forbade you from seeing the timesheets.

Auditing is where outcome pricing most often falls apart in year two. The honeymoon number looks great, FinOps starts asking pointed questions about why the agent's billed outcomes don't reconcile with downstream business results, and the vendor's dashboard has no independent corroboration. The fix has to be designed in at contract time, because retrofitting auditability onto a live, vendor-instrumented system is nearly impossible.

What good auditability actually requires:

This is also why outcome pricing tends to favor incumbents with rich first-party data, they can adjudicate their own outcomes more credibly because they hold the surrounding context, a dynamic worth understanding on its own.

The Three Audit Models

In the field, outcome auditing falls into three patterns, in ascending order of trustworthiness and cost.

Vendor-Attested

The vendor reports, the buyer trusts. Cheap, fast, fragile. Fine for low-stakes, low-dollar deals where the outcome is hard to fake (e.g., documents processed, where the buyer receives the documents). Dangerous for anything judgment-laden or high-value.

Buyer-Verified

The outcome event lives in or is mirrored to a system the buyer controls, their CRM, helpdesk, or data warehouse, and billing is computed from that record. This is the sweet spot for most serious deals. It costs integration effort up front and pays for itself the first time a discrepancy surfaces. The buyer isn't auditing the vendor's claim; the buyer is generating the number.

Third-Party / Independent

A neutral party, an auditor, an escrow service, or increasingly an automated reconciliation tool, computes outcomes from shared raw data. Rare today, but it's where high-value and regulated deals are heading, mirroring how ad-tech eventually grew independent verification (think MRC-accredited measurement) after years of platforms grading their own impressions. As agent spend scales into seven figures, expect "outcome verification" to become its own product category, and a prerequisite for agent pricing in regulated industries where audit overhead is mandatory.

A useful rule of thumb: the audit model should escalate with the dollar value and the subjectivity of the outcome. A $2,000/month document-processing agent can run on vendor attestation. A $40,000/month revenue-recovery agent needs buyer verification at minimum and probably an incrementality framework on top.

Designing an Outcome Definition That Survives a Dispute

If you're writing or signing one of these contracts, a durable outcome definition has roughly six properties. Miss any one and you've left a gap an optimizing agent, or an optimizing account manager, will eventually find.

  1. Observable by the buyer. The triggering event must be visible in a system the buyer can see, not buried in vendor-internal state.
  2. Time-bounded. "Resolved" must include a durability window ("stays closed 14 days") so the count can't be inflated by churn-and-reopen.
  3. Counterfactual-aware. For value-based outcomes, define the baseline and the hold-out up front, not after the first disputed invoice.
  4. Quality-gated. Pair the quantity metric with a quality guardrail, CSAT on resolved tickets, show-rate on booked meetings, so the agent can't farm cheap junk outcomes.
  5. Reconcilable. Raw events must be exportable and joinable to the buyer's data, with a stated dispute-and-credit process.
  6. Capped or floored where it matters. Tie this to budget protection so a runaway "success" doesn't produce a runaway bill, see floor-and-ceiling structures.

Notice that five of those six properties are about measurement and verification, not price. That ratio is the whole lesson. The hard part of outcome pricing was never picking the number. It was defining the noun the number multiplies, and proving the noun actually occurred.

What This Means for Buyers and Vendors

For buyers: stop negotiating the price-per-outcome first. Negotiate the definition and the audit rights first, then the price. A higher per-outcome rate on a tightly defined, buyer-verified, quality-gated outcome is almost always a better deal than a cheap rate on a vendor-defined, vendor-audited fuzzy one. Bring someone technical into the room who can ask exactly how the billing event fires. And treat any vendor refusal to share raw event logs as disqualifying.

For vendors: counterintuitively, the tighter and more buyer-verifiable your outcome definition, the easier your sales motion gets and the stickier your renewals become. Loose definitions win the first deal and lose the third. Vendors who proactively offer hold-outs, exportable logs, and quality gates are signaling confidence, and confidence is the actual product in a category where buyers can't inspect the agent's reasoning. The companies that will dominate GaaS aren't the ones with the cleverest counting tricks, they're the ones who made their counting boring, transparent, and impossible to argue with.

Outcome-based pricing is genuinely the most aligned model agentic AI has produced. But alignment isn't a property of the pricing label. It's a property of who holds the definition and who keeps the books. Get those two things right and "we only charge when it works" stops being a marketing line and starts being true.

Insights Most People Overlook

References

#gaas pricing models

More in Pricing