Pricing Experiments: What Top Agent Startups Are Actually A/B Testing Right Now
Agent startups are running pricing experiments at a speed that would give a traditional SaaS finance team a stroke. The big ones aren't testing $49 vs. $59 a seat, they're testing whether to charge per seat at all, whether to surface token counts, and whether "we only charge when it works" converts better than a flat retainer. This piece breaks down the specific live experiments running across support, sales, coding, and ops agents, what the early signal says, and where most teams are getting the experiment design wrong. The short version: the winning tests aren't about price points, they're about which *unit* the buyer trusts.
Table of Contents
- Why Agent Pricing Is a Live Experiment, Not a Decision
- The Experiments Top Startups Are Actually Running
- Per-Resolution vs. Per-Seat in Support
- Outcome Fee vs. Flat Retainer in Sales and SDR Agents
- Showing vs. Hiding Token Counts
- Credits and Prepaid Pools vs. Pure Metering
- Autonomy Tiers as a Pricing Axis
- How to Design a Pricing Experiment That Doesn't Lie to You
- The Metrics That Actually Tell You Which Model Won
- Insights Most People Overlook
- References
Why Agent Pricing Is a Live Experiment, Not a Decision
There's a reason the pricing page of a fast-moving agent company looks different every quarter. Nobody knows the right answer yet, and the cost structure underneath the product keeps moving. Inference prices drop, model routing gets smarter, a customer's usage triples overnight because the agent started doing more of the job. A pricing model that was margin-safe in January can be underwater by June.
So the serious teams have stopped treating pricing as a one-time decision and started treating it as a product surface they iterate on. They run experiments. The interesting part is what they test, because it's almost never the thing a SaaS veteran would expect.
Traditional SaaS A/B pricing tests optimize within a fixed model: $20 vs. $25, three tiers vs. four, monthly vs. annual default. Agent startups are running tests one level up, they're testing the pricing model itself. Per-task against per-outcome. Per-seat against per-resolution. The unit of value is still genuinely contested, and changing the unit changes everything downstream: the sales motion, the buyer's mental model, the contract, the margin. This is the core tension running through the whole GaaS pricing taxonomy, and it's why a single pricing experiment can move conversion by double digits where a SaaS price tweak moves it by a point or two.
A useful frame, borrowed from the way a16z writes about AI business models, is that agents collapse the old line between "software you buy" and "labor you hire." When you're selling something that behaves like labor, buyers expect to pay like they pay for labor, for results, or at least for work done, and that expectation is what the experiments keep bumping into.
The Experiments Top Startups Are Actually Running
Let me get specific, because vague talk about "value-based pricing" helps nobody. Here are the live tests, by category, with what the early read seems to be.
Per-Resolution vs. Per-Seat in Support
Customer-support agents are the most mature corner of the market, so the experiments here are the furthest along. The headline test, popularized by Intercom's Fin and now copied widely, is per-resolution pricing, you pay roughly $0.99 (Fin's anchor number) only when the agent actually resolves a ticket, versus the old per-seat model where you pay for "agent seats" whether they work or not.
The per-resolution framing wins on the sales call almost every time, because it sounds like it shares the risk. But the experiment underneath is subtler than "which converts better." The real test is what counts as a resolution. Vendors are A/B testing definitions: does a deflection count, or only a confirmed-solved ticket? Does a conversation the customer abandoned count as resolved? Teams that tested a loose definition got higher revenue per account but a spike in disputes and churn; teams that tested a strict, auditable definition got lower revenue per ticket but far better renewal rates. That tradeoff, revenue capture versus trust, is the actual variable, and it connects directly to the harder question of who defines and audits the outcome.
The case against per-seat here is strong enough that it's becoming its own talking point, but a few vendors are quietly testing a return to it for enterprise buyers who want a predictable line item. Predictability sells to procurement even when it's economically worse for the buyer.
Outcome Fee vs. Flat Retainer in Sales and SDR Agents
Sales-development agents, the ones that book meetings or qualify leads, are running the most aggressive outcome experiments. The classic test is a flat monthly retainer (say $2,000/month for the agent) against a per-meeting-booked or per-qualified-lead fee (say $40 per booked meeting).
Here's what makes this one fascinating: the outcome model usually wins the deal and loses the account. Buyers love "pay per meeting" at signup. Then the agent gets good, books 200 meetings a month, and the buyer is suddenly paying $8,000 against a $2,000 retainer they'd have happily signed. The bill scales with the agent's success, which feels like a punishment for the agent working. This is the land-and-expand dynamic turned into a liability, expansion is automatic, and automatic expansion that the buyer didn't choose breeds resentment.
The smart experiment here isn't outcome-vs-flat. It's floor-and-ceiling: a base fee plus per-outcome pricing that caps out at a ceiling so the buyer's budget can't run away. Early data from teams running this hybrid suggests it converts nearly as well as pure outcome pricing at signup and renews dramatically better, because the buyer never gets a surprise invoice. McKinsey's work on how generative AI is reshaping go-to-market economics keeps landing on the same point: the models that survive contact with a renewal cycle are the ones that protect the buyer from their own success.
Showing vs. Hiding Token Counts
This is the quiet experiment almost nobody talks about publicly, and it's one of the most consequential. Should a usage-based agent product show the customer the underlying token counts and model costs, or abstract them away behind a clean unit like "tasks" or "credits"?
The instinct is that transparency builds trust. The experiments mostly say otherwise. When vendors A/B tested an invoice that exposed raw token usage against one that showed a clean "1 task = 1 credit" abstraction, the abstracted version reduced support tickets, reduced billing anxiety, and improved renewal, even though it gave the customer less information. Showing token counts invited the customer to start optimizing against your meter, second-guessing every agent action, and treating the bill as adversarial. The psychology of metered billing is genuinely strange, and exposing the raw machinery tends to make buyers more anxious, not less.
The counter-experiment, run mostly by developer-facing tools, found the opposite: technical buyers distrust abstraction and assume they're being marked up. For that audience, showing token counts (with the markup stated honestly) won. So the real finding isn't "show" or "hide", it's that the right answer is segment-dependent, and the experiment worth running is who you're showing it to.
Credits and Prepaid Pools vs. Pure Metering
A pattern that's hardening into a default across the GaaS space is the prepaid credit pool: the customer buys a chunk of credits up front, the agent draws them down per task, and they refill. This is being A/B tested against pure post-paid metering (you use it, we bill you monthly for exactly what you used).
Prepaid pools win on three axes the experiments keep confirming: cash flow (you get paid before you deliver), psychological commitment (a prepaid balance gets used, like a gift card), and budget anxiety (a finite pool feels safer than an open meter). The losing edge is that breakage, unused credits, can feel like you're charging for nothing, which trips the same "we only charge when it works" nerve that outcome pricing pokes at. The vendors getting this right are testing rollover policies and expiration windows as the real variable, not the credit price.
Autonomy Tiers as a Pricing Axis
The newest and least-settled experiment: pricing by how much autonomy the agent has. A "suggests actions, human approves" tier priced lower than a "takes actions independently" tier, priced lower than a "operates across systems unsupervised" tier.
This is clever because it maps price to the thing the buyer is actually nervous about, letting the agent loose, rather than to raw usage. The early experiments suggest autonomy tiers convert well as an upsell ladder: buyers start supervised, get comfortable, and climb. The risk being tested is whether the lower tiers anchor the agent as "just a suggestion tool" and cap the perceived value. A few teams found their supervised tier cannibalized the autonomous one because buyers never felt safe enough to climb. Tiering by autonomy is promising, but it's the experiment with the least data behind it so far.
How to Design a Pricing Experiment That Doesn't Lie to You
Most agent-pricing A/B tests are quietly broken, and the breakage is almost always the same: the sample is too small and the feedback loop is too slow.
A SaaS checkout test gets thousands of sessions a week. An agent pricing test in a B2B motion might get forty deals a quarter, half of which are still mid-negotiation when you want to read results. You cannot run a clean statistical A/B test on forty enterprise deals. Anyone telling you their per-outcome model "converts 23% better" off that sample is reading noise.
So the teams doing this well do three things differently. First, they test pricing models in sequence, not in parallel, run flat for a quarter, run outcome for a quarter, compare cohorts, because true randomized parallel tests on small B2B samples don't reach significance. Second, they instrument the renewal, not just the signup, because in agent pricing the signup conversion and the renewal conversion often point in opposite directions (outcome pricing being the canonical example). A test that only measures close rate is measuring the least important half of the funnel. Third, they separate the price-point question from the price-model question and never test both at once, because if you change the unit and the number simultaneously you can't attribute the result.
There's a deeper trap worth naming. The thing you're really testing is the buyer's mental model, not their willingness to pay a number. When per-resolution pricing wins, it's not because $0.99 is the magic price, it's because the buyer's brain files it under "I pay for outcomes" instead of "I pay for software." That reframing is the product of the experiment. Which means qualitative signal, what buyers say in the call, where they hesitate, what they ask the legal team to redline, is often more diagnostic than the conversion number, and the best teams treat sales-call transcripts as primary experimental data.
The Metrics That Actually Tell You Which Model Won
If you only watch close rate, you'll pick the wrong model roughly half the time. The metric stack that actually separates winners from losers in agent pricing experiments looks like this:
- Net revenue retention by pricing cohort. The single most predictive number. Outcome and usage models often signup-convert better but retain worse; flat and credit models often the reverse. NRR catches the gap that close rate hides.
- Gross margin per account, tracked monthly. Because inference costs move, a model that's margin-safe at signup can erode. You're testing margin stability as much as margin level. This is where model routing, quietly serving a cheaper model when the task allows, interacts with pricing: the same price can be 40% or 70% margin depending on routing, and your experiment has to hold routing constant or it's measuring the wrong thing.
- Billing-related support volume. A pricing model that generates a flood of "why is my bill like this" tickets is failing even if it converts, because that volume predicts churn and eats your support margin.
- Invoice variance. How much does the bill swing month to month? Low variance correlates strongly with renewal across nearly every experiment. Buyers will pay more for a predictable number than less for a volatile one, a finding that should make every "pure metering" advocate nervous.
- Time-to-value relative to first invoice. If the customer's first real bill arrives before they've felt the agent earn it, you lose them regardless of model. Some teams are testing deliberately delayed billing, let the value land first.
The honest read across all of these is that the pricing model that wins is rarely the one that extracts the most revenue per transaction. It's the one that keeps the buyer calm. Calm buyers renew, expand, and refer. Anxious buyers churn even when the agent is doing brilliant work. Most pricing experiments in this space are, when you strip them down, anxiety experiments wearing a revenue costume.
Insights Most People Overlook
The "we only charge when it works" message can hurt you. It's the most copied positioning in GaaS, and it backfires in a specific way: it teaches the buyer to scrutinize every outcome and dispute the borderline ones. Vendors who tested a softer "you pay for results" framing against the aggressive "only when it works" guarantee found the guarantee increased disputes and lowered trust, because it invited adversarial accounting. The strongest outcome claim is sometimes the quieter one.
Outcome pricing structurally favors incumbents, and startups keep adopting it anyway. To price on outcomes you need to define, measure, and audit them, which requires historical data on what "good" looks like. The incumbent with five years of resolution data can price outcomes confidently; the startup is guessing. Every time a startup A/B tests outcome pricing and "wins," it may just be inheriting a definitional advantage it doesn't actually have, and the bill for that guess arrives at the first dispute.
The best experiments change the unit, not the price, and the unit is a one-way door. You can A/B test $0.99 vs. $1.29 per resolution forever with no lasting damage. But if you switch from per-seat to per-resolution and back, you've taught your market that your pricing is unstable, and enterprise procurement remembers. Unit changes aren't reversible the way price changes are. Treat the unit experiment with the gravity of a product pivot, because that's what it is.
Repricing as model costs drop is the experiment nobody schedules. Inference gets cheaper every quarter, which means a margin-safe price slowly becomes a price-gouging price in the customer's eyes, or a competitive vulnerability when a rival passes the savings through. Smart teams are running deliberate experiments on when and how to lower prices and how to handle grandfathered customers, which is the inverse of every pricing test in the SaaS playbook and almost completely undiscussed.
Qualitative beats quantitative in this market, and admitting it is the unlock. Because sample sizes are tiny and feedback loops are long, the teams winning pricing experiments are the ones who treat the sales-call redline, the hesitation, and the FinOps team's questions as the real data. The conversion number is a lagging, noisy confirmation. What the buyer's finance team flags in the contract is the leading signal, and it's available weeks earlier.
References
More in Pricing
- The "Minimum Viable Margin" Every GaaS Startup Needs Before It Scales
- Why Metered Pricing Quietly Rewards Your Power Users and Taxes Your Growth Customers
- Currency of Value: Should Your AI Agent Charge for Time Saved, Revenue Gained, or Cost Cut?
- Refunds and SLAs When an Agent Fails the Task
- Pricing Agents in Regulated Industries: Why Audit Overhead Changes the Math