THE INDEPENDENT RECORD · AGENTIC AI AS A SERVICE AboutStandardsContact
GAASAGENTIC AI · AS A SERVICE
INDEPENDENT · SINCE 2026
UPDATED DAILY
NO HYPE · NO PAY-TO-PLAY
PER-TASK PRICING NOW STANDARD ● NEW BENCHMARK: 71% TASK COMPLETION ● ENTERPRISE PILOTS UP 4X ● RUNTIME FUNDING ACCELERATES ● "AGENTS ARE THE NEW SEATS" ● MARGINS UNDER PRESSURE ● THE INDEPENDENT RECORD ON GAAS
Economics

Per-Outcome Pricing for AI Agents: Can You Actually Measure the Outcome?

Per-outcome pricing sounds like the holy grail of Agentic AI-as-a-Service: charge only when the agent delivers a result, not for the compute it burns. The catch is brutal and underdiscussed, most "outcomes" are genuinely hard to define, attribute, and verify. Before you sign an outcome-based contract (or write one), you need a defensible answer to three questions: what counts as the outcome, who decides it happened, and how do you stop both sides from gaming the definition. This piece is a working field guide to that measurement problem, and an honest accounting of when per-outcome pricing falls apart.

By N. Adeyemi · Jun 9, 2026 · 15 min read

Table of Contents

The Pitch and the Problem

The pitch writes itself. Software you pay for only when it works. No seat licenses for tools nobody opens, no consumption bills that spike when an engineer leaves a loop running over the weekend. You pay per resolved ticket, per qualified lead, per reconciled invoice, per closed-won deal. The vendor takes on the risk; you take on the upside. Marc Benioff has spent the better part of two years telling anyone who'll listen that the future is priced per outcome, and the framing has stuck because it maps so cleanly onto how buyers already think about ROI.

Then you try to write the contract.

The first sentence, "Customer shall pay $X per resolved support ticket", looks fine until your lawyer asks what "resolved" means, and your VP of Support points out that 40% of tickets are resolved by the customer giving up, and your data team mentions that the agent often deflects a ticket that a human would have escalated into a $12,000 expansion conversation. Suddenly the clean unit metric is a swamp. The outcome didn't get harder to deliver. It got harder to measure, and in per-outcome pricing, measurement is the product.

This is the part the keynote skips. Outcome-based pricing isn't a billing change. It's a measurement system you now have to build, operate, and defend, and the economics of the whole arrangement live or die on whether that system is trustworthy enough that neither party feels robbed.

What "Outcome" Actually Means

The word "outcome" is doing an enormous amount of work, and it hides at least three different things people mean when they use it.

An action completed. The agent did the thing: it sent the email, filed the ticket, updated the CRM field, generated the report. This is the easiest to bill and the least defensible as an "outcome," because completing an action is not the same as producing value. It's barely distinguishable from per-task pricing (see #2, Cost-per-completed-task: defining the category's core unit), just with prettier language.

A state change in the customer's world. A lead moved from "new" to "qualified." A ticket went from "open" to "closed." An invoice went from "received" to "paid." This is what most vendors mean when they say outcome, and it's the interesting middle ground, measurable in principle, because the state lives in a system of record, but contaminated by all the other forces that move that state.

A business result. Revenue booked, churn prevented, cost avoided, hours saved. This is what buyers mean when they say outcome, and it's almost never cleanly attributable to a single agent. Everyone wants to price on this layer and almost nobody can, because the causal chain from "agent did something" to "company made money" runs through a dozen other variables.

The entire negotiation in an outcome-based deal is really a fight about which of these three layers the price attaches to. Vendors drift toward "action completed" because it's countable and defensible. Buyers pull toward "business result" because that's what they actually care about. The contracts that work land deliberately in the middle, a state change that both sides agree is a fair proxy for value, and then spend most of their word count defining that state change with paranoid precision.

The Measurement Stack: Four Layers That Have to Hold

Think of outcome measurement as a stack. If any layer is weak, the pricing model leaks, either the vendor eats unbillable work or the buyer pays for value that didn't materialize. Four layers, each with its own failure mode.

Definition: Drawing the Line Around "Done"

Definition is where most outcome contracts are quietly lost. "Resolved ticket" needs an operational definition that a script could evaluate without human judgment: closed within the SLA window, no reopen within 7 days, customer satisfaction score not negative, no human agent touched it after the bot. The moment any of those conditions requires interpretation, you've introduced a place for disputes to live.

A good definition is exclusionary, not just inclusive. It says what doesn't count. A "qualified lead" that the agent generated by lowering the qualification bar isn't a real outcome, it's the metric eating itself. The best definitions I've seen include a quality gate that's checked after the fact: a sampled human audit, a downstream conversion threshold, a clawback window. Without that, you're paying for volume dressed up as value, and the agent will optimize for exactly the thing you measured rather than the thing you meant.

Attribution: Did the Agent Cause It?

Even with a crisp definition, you have to prove the agent caused the outcome rather than merely being present when it happened. This is the attribution problem, and it's the same beast that has tortured marketing analytics for twenty years. If a lead gets a nurture email from the agent, a call from a human SDR, and a retargeting ad, and then converts, who gets credit?

Honest attribution requires a counterfactual: what would have happened without the agent? The gold standard is a holdout group, route some share of traffic or tickets through the old process and compare. Almost nobody does this, because it means deliberately not using the thing you're paying for on a slice of your volume, and because it surfaces uncomfortable truths about how much lift the agent actually provides. McKinsey's work on capturing real value from AI keeps returning to this point: organizations that can't isolate AI's incremental contribution tend to overstate it, sometimes wildly. The agent that gets credit for every closed ticket in its queue is taking credit for the ones that would have closed themselves.

Verification: Who Confirms It Happened?

Verification is the question of who holds the source of truth and whether both parties trust it. If the vendor's own system counts the outcomes and sends the bill, the buyer is being asked to trust the dealer's scorecard. If the buyer's system of record (their CRM, their helpdesk, their ERP) is authoritative, the vendor has to trust data they don't control and can't fully see.

The cleanest arrangements use the buyer's system of record as the referee, with the vendor's agent writing structured, auditable events into it, a verifiable trail of what the agent did and when, tied to the state change being billed. This is where the reliability and observability sub-topics of the GaaS cluster stop being abstract: if you can't produce a tamper-evident log of agent actions, you can't run outcome pricing at all, because you can't survive the first billing dispute. a16z's analysis of pricing in the AI-agent era makes a related point, outcome models require an instrumentation layer that most early vendors simply haven't built, which is why so many "outcome-based" pitches quietly revert to usage billing once the contract is signed.

Settlement: Turning a Verified Outcome Into a Charge

Settlement is the boring layer that breaks in production. You've defined, attributed, and verified an outcome, now you have to bill it, handle disputes, process clawbacks when a "resolved" ticket reopens, and reconcile the vendor's count against the buyer's count at month-end. If those two numbers don't match, and they never match the first month, someone has to adjudicate, and the cost of that adjudication is a real, recurring expense that almost never makes it into the pricing model.

Settlement is also where the cash-flow mismatch bites the vendor. The agent burns inference cost in January (tokens, retries, tool calls, see #3, The hidden cost of retries: when one task becomes fifty model calls), but the outcome it's billing against might not verify until February, after the clawback window closes. The vendor financed the compute and waits to get paid for the result. That gap is survivable at small scale and lethal at large scale if the verified-outcome rate comes in below model.

Where Per-Outcome Pricing Works, and Where It Quietly Doesn't

Per-outcome pricing is not universally good or bad. It's a fit for a specific shape of problem, and a mismatch for others.

It works when the outcome is discrete, quickly observable, and lives in a system of record. Booking an appointment. Reconciling a transaction. Resolving a tier-1 password reset. These have a binary done/not-done state, a short verification lag, and an authoritative source of truth. The measurement stack is cheap to build because the world is already keeping score.

It struggles when the outcome is continuous, delayed, or judgment-laden. "Improved code quality." "Better customer sentiment." "Strategic research." There's no clean state change, the value shows up weeks later through a fog of other variables, and verification requires human judgment that reintroduces all the cost and subjectivity you were trying to escape. Here, outcome pricing becomes a negotiation about feelings, and feelings don't reconcile at month-end.

It actively misfires when the easy-to-measure proxy diverges from the real goal. A support agent paid per deflected ticket will deflect tickets that should have been escalated. A sales agent paid per meeting booked will book meetings with anyone who has a pulse. The cleaner and more automatic your outcome measurement, the more powerful the incentive to optimize the proxy at the expense of the intent, a dynamic close enough to Goodhart's Law that you should treat it as a design constraint, not an edge case. The Harvard Business Review's writing on the perils of measuring the wrong thing is two decades old and reads like it was written for this exact moment.

The Gaming Problem on Both Sides

Outcome pricing creates a measurement adversary on each side of the table, and pretending otherwise is how good deals go sour.

The vendor is incented to define outcomes loosely, count generously, and resist clawbacks. If the contract says "qualified lead" without teeth, the vendor's agent will hit the number by relaxing what "qualified" means. If "resolved ticket" doesn't exclude reopens, the vendor bills the same ticket twice.

The buyer is incented in the opposite direction: to dispute marginal outcomes, attribute every success to other causes, and stretch verification windows to delay payment. A buyer who's having a tight quarter can find reasons to contest the agent's count, and the vendor has limited recourse because the buyer holds the system of record.

This is why the contract's measurement language matters more than the price. A per-outcome deal with a fuzzy definition is just a lawsuit on a timer. The deals that hold up read like derivatives contracts, precise event definitions, named source of truth, audit rights, a dispute-resolution mechanism, and a clawback window with a hard expiry. That's a lot of machinery to charge for resolved tickets, which is exactly why pure per-outcome pricing remains rarer in practice than in pitch decks, and why most "outcome" pricing in the wild is really a hybrid: a usage floor that covers the vendor's compute, plus an outcome bonus on top. The hybrid isn't a cop-out. It's an honest acknowledgment that the measurement system isn't trustworthy enough yet to carry 100% of the revenue.

A Practical Test Before You Sign or Sell

Before you commit to per-outcome pricing, as buyer or vendor, run the proposed outcome through five questions. If you can't answer all five crisply, you're not ready to price on that outcome yet.

  1. Can a script decide it? If determining whether the outcome occurred requires human judgment, your verification costs will eat the model. Push the definition until a query against a database can return true or false.

  2. Where does it live? Name the single authoritative system of record. If the answer is "the vendor's dashboard," the buyer should push back. If it's "we'd have to build that," nobody's ready.

  3. What's the verification lag? Measure the gap between the agent acting and the outcome being confirmable-and-final. Short lags (minutes to days) are workable. Long lags (weeks to a quarter) create cash-flow risk for the vendor and attribution fog for everyone.

  4. What's the counterfactual? If you can't describe what would have happened without the agent, you can't claim the agent caused the outcome. A holdout group is the honest answer; a documented baseline is the minimum.

  5. How does the proxy diverge from the goal? Write down the dumbest way an agent could hit the metric without delivering the value. If that failure mode is cheap and likely, add a quality gate before you add a price.

The vendors winning at outcome pricing in 2026 aren't the ones with the boldest "pay only when it works" headline. They're the ones who quietly built the measurement stack, the event logging, the attribution model, the clawback machinery, and can hand a skeptical buyer an auditable trail on day one. The pricing model is downstream of the measurement system. Get the measurement right and the pricing follows. Get it wrong and no clever rate card will save you.

Insights Most People Overlook

The measurement system is the actual product, and almost nobody prices it in. When a vendor offers per-outcome pricing, the impressive part isn't the agent, it's the instrumentation that can define, attribute, verify, and settle outcomes well enough to bill on them. That instrumentation costs real money to build and operate, and it shows up nowhere on the rate card. Vendors who treat measurement as a billing afterthought rather than a core product investment are the ones who quietly revert to usage pricing six months in.

Outcome pricing can be worse for the buyer even when it's "fairer." Buyers assume paying only for results protects them. But a well-defined outcome that's cheap for the agent to produce at scale can cost the buyer far more in aggregate than a flat usage fee would have, especially for high-volume, low-effort outcomes. "Pay per resolved ticket" sounds buyer-friendly until you're paying premium per-unit rates on a million password resets that cost the vendor fractions of a cent each. Per-outcome aligns incentives; it does not guarantee the buyer the better deal.

The clawback window is the most important number in the contract, and it's usually an afterthought. A "resolved" ticket that reopens, a "qualified" lead that bounces, a "closed-won" deal that cancels, these reversals are where outcome pricing either stays honest or becomes fiction. A contract with no clawback window bills phantom outcomes; one with a window that's too short lets reversals slip through; one that's too long destroys the vendor's cash flow. This single parameter quietly determines whether the model is sustainable, yet it's routinely buried in a schedule nobody reads.

Per-outcome pricing punishes the buyer's best customers and rewards their worst. Attribution fog cuts in a direction people miss: the agent gets full credit for outcomes that would have happened anyway (your motivated, high-intent customers) and gets blamed for non-outcomes among hard cases it never had a chance with. Without a holdout, you systematically overpay for easy wins and underinvest in exactly the segments where agent lift would matter most. The pricing model's blind spot becomes a strategic blind spot.

"Outcome" is a moving target the agent will chase. Unlike a SaaS feature, an agent optimizes against whatever you measure, continuously and creatively. The outcome definition you write today is the spec the agent will reverse-engineer and exploit tomorrow. This makes outcome definitions less like contract terms and more like adversarial reward functions, they need the same paranoid red-teaming you'd give a security boundary, because in effect that's what they are.

References

#per-outcome pricing#outcome-based ai pricing#agent pricing models#gaas unit economics#per-task pricing

More in Economics