How to Stress-Test Revenue Durability at an Agentic AI Company
Revenue durability is the single hardest thing to underwrite in an Agentic AI-as-a-Service (GaaS) company, because the headline ARR can be real, growing, and structurally fragile all at once. A proper stress test pulls apart usage-based revenue into its load-bearing and decorative parts, models what happens when a foundation model gets 10x cheaper or a customer's pilot budget dries up, and asks whether the agent is doing work the buyer would pay for indefinitely or work that evaporates the moment a procurement committee blinks. This guide gives you the specific scenarios, the cohort math, and the diligence questions that separate durable agent revenue from a number that looks great until renewal season. It is written for investors, operators, and acquirers who need to know whether the run-rate survives contact with reality.
Table of Contents
- Why Agent Revenue Breaks Differently
- The Four Failure Modes a Stress Test Must Catch
- Building the Base Case Before You Shock It
- The Scenarios That Actually Matter
- Scenario 1: The Model-Cost Collapse
- Scenario 2: The Pilot Cliff
- Scenario 3: The Insourcing Threat
- Scenario 4: The Reliability Recession
- Cohort Math: Where the Truth Lives
- A Practical Diligence Checklist
- Insights Most People Overlook
- References
Why Agent Revenue Breaks Differently
Most investors learned revenue-quality analysis on SaaS, and most of those instincts mislead them on agent companies. A SaaS seat is a fairly stable thing: a person logs in, the company pays per head, and the revenue moves slowly with headcount. You can stress that, but the failure modes are gentle and well understood.
Agent revenue is different in kind, not just degree. When a company sells an AI agent on a per-task or per-outcome basis, the revenue is tied to work volume, and work volume is far more elastic than seat count. It can double in a quarter because a customer rolled the agent out to a new department. It can also halve because someone in finance noticed the bill and capped usage, or because the underlying business slowed and there was simply less work to do. The same mechanism that makes usage-based GaaS revenue expand so beautifully in a good quarter makes it contract just as fast in a bad one. This is exactly the tension behind the broader debate over whether usage revenue is durable, and a stress test is how you stop debating and start measuring.
There is a second structural issue. A large and variable share of an agent company's cost of revenue is inference, which means gross margin is hostage to model pricing the company does not control. When token prices fall, that can be pure margin expansion, or it can be competed away in a price war, or it can quietly remove the company's reason to exist if a foundation lab ships the same capability natively. A revenue durability stress test for agent companies has to hold both the top line and the margin line in tension at the same time. You cannot stress one without the other.
The third difference is the most uncomfortable. A meaningful fraction of GaaS revenue in this market is experimental spend, not operational spend. Enterprises are running pilots out of innovation budgets, and innovation budgets are the first thing cut when the macro turns. That money looks identical to durable revenue on a P&L. The whole job of a stress test is to tell the two apart before you wire the funds.
The Four Failure Modes a Stress Test Must Catch
Before you build any model, get specific about what you are trying to detect. A durable-revenue stress test for an agent company is hunting for four distinct ways the number can be a mirage:
- Volume fragility, revenue that depends on a few customers running the agent at high volume, where a usage cap or a single churned logo erases a quarter of run-rate.
- Margin fragility, revenue that only converts to a healthy gross margin under today's specific inference prices, and inverts if a price war forces the company to pass savings through.
- Budget fragility, revenue funded by experimental or innovation line items rather than operating budgets, which disappears in a downturn regardless of how well the product works.
- Moat fragility, revenue earned doing a task that a model provider or an incumbent SaaS vendor can absorb into their platform, turning a standalone product into a feature.
A good stress test produces a number for each. A great one tells you which failure mode dominates, because the remedy is different for each, and so is the right valuation.
Building the Base Case Before You Shock It
You cannot stress a number you have not decomposed. Start by rebuilding reported revenue from the transaction level up, not from the summary the company hands you.
Pull the per-customer, per-month revenue series for at least the trailing 24 months, or the full life of the company if it is younger. Then tag every dollar with attributes that matter for durability: contract type (committed vs. pure consumption), budget source (operating line vs. innovation/pilot), pricing basis (per-seat, per-task, per-outcome), and the specific workflow the agent performs. This tagging is tedious and it is where the real findings come from. It is common to discover that 60% of "ARR" is actually uncommitted consumption revenue with no contractual floor, which is a completely different asset than a committed subscription even if the trailing numbers look the same.
Next, separate gross retention from net retention and compute both at the cohort level, not blended. Blended net revenue retention is the single most over-trusted metric in GaaS diligence, because a handful of explosively expanding accounts can mask a leaky base. The framing that usage-based companies need different benchmarks than seat-based SaaS is now well established among growth investors, and it applies with extra force to agents, where expansion and contraction both run faster.
Finally, establish your inference-cost base case honestly. Get the actual token consumption per task, the current per-token cost, and the blended gross margin after inference, not before. A surprising number of agent companies report gross margins that quietly exclude the model bill. If you do not nail down the real base case here, every downstream scenario is built on sand.
The Scenarios That Actually Matter
With a clean base case, you run shocks. The point of each scenario is not to predict the future but to find the break point: how far does this variable have to move before the company's economics stop working? Run each one as a distinct sensitivity, then combine the two or three most plausible into a single adverse case.
Scenario 1: The Model-Cost Collapse
Foundation model prices have fallen by roughly an order of magnitude per unit of capability on a recurring basis, and there is no reason to assume the trend stops. For most software businesses, cheaper inputs are unambiguously good. For an agent company, it is ambiguous, and the stress test has to model both branches.
Branch A: the cost saving drops to the bottom line and gross margin expands. Branch B: a competitor passes the saving through as a price cut and the company must follow to keep volume, so revenue per task falls while volume holds flat. Model both. The dangerous companies are the ones whose current margins depend on customers not noticing that the underlying model got cheaper. When you stress this, you are partly stress-testing pricing power: a company that prices on outcomes rather than tokens is far more insulated, because the value of the outcome did not change just because the input got cheaper. This is where pricing architecture and revenue durability meet, and it connects directly to the question of why model-cost compression triggers valuation haircuts.
Scenario 2: The Pilot Cliff
This is the scenario most decks pretend does not exist. Tag the revenue that originated as a pilot, a proof-of-concept, or an innovation-budget purchase, then model a downturn in which that spend is cut by 40% to 70% at renewal. The number that survives is your operational revenue base, and it is frequently a fraction of headline ARR.
The tell to look for is whether pilots have converted into production deployments with operating-budget ownership, or whether the company is simply stacking new pilots on top of old ones to keep the growth curve smooth. A business growing on serial pilots can post wonderful numbers right up until the innovation-budget cycle turns, at which point it discovers it never had durable revenue at all. McKinsey's ongoing work on enterprise AI adoption keeps surfacing the same pattern: a wide gap between organizations experimenting with AI and those capturing durable, scaled value from it. The pilot cliff is that gap showing up in someone's revenue.
Scenario 3: The Insourcing Threat
Every agent company is one platform release away from irrelevance in some slice of its market. Stress this by asking, customer cohort by customer cohort: what happens to this revenue if the foundation lab ships this capability natively, or if the incumbent SaaS vendor the customer already pays bundles an equivalent agent for free?
The companies that survive this shock are doing work that is deeply embedded in a customer's specific data, workflows, and accountability structures, the kind of vertical agent that would be expensive and risky to rip out. The companies that fail are thin wrappers whose entire value-add is a prompt and a nice UI. This is the same dynamic that drives the "agent attach" acquisition thesis: incumbents buy or build agents precisely to absorb this revenue. If your stress test shows a cohort's revenue evaporating under a plausible platform release, that revenue should be valued as an option, not as an annuity.
Scenario 4: The Reliability Recession
Agents fail in ways software does not. When an autonomous workflow produces a wrong outcome at scale, customers do not just file a support ticket, they cap usage, pull the agent from production, or churn. The stress test should model a reliability incident: a measurable drop in task success rate, and the revenue impact of customers throttling back to manual review.
This is where agent reliability stops being an engineering topic and becomes a revenue-quality topic. Ask for the actual task-success and intervention-rate metrics over time. A company whose reliability is improving can take a reliability shock and recover; a company papering over a flat or declining success rate with sales growth is carrying hidden revenue risk that a single high-profile failure can crystallize.
Cohort Math: Where the Truth Lives
Scenarios tell you how bad things can get. Cohort analysis tells you whether the business is actually getting healthier or just getting bigger. For an agent company, three cohort curves matter more than any blended metric.
First, gross revenue retention by cohort. Strip out all expansion and ask: of the revenue a cohort represented at month zero, how much remains at month 12 and month 24? This is the closest thing to a structural durability measure you have. Expansion can be bought; gross retention has to be earned. If gross retention is below the high-80s for an enterprise agent product, you are looking at a leaky bucket no matter how good the new-logo story is.
Second, time-to-production by cohort. Track how long it takes a new customer to move from pilot to operating-budget production, and whether that time is shrinking across cohorts. Falling time-to-production is one of the strongest leading indicators that a GaaS company's revenue is becoming more durable, because it means the budget-fragility problem is resolving itself.
Third, margin by cohort. Because inference cost per task tends to fall over time, older cohorts on a well-run agent product should show expanding gross margins as the same workflow gets cheaper to serve. If older cohorts are not showing margin expansion, either the company is competing the savings away or its inference efficiency is not improving, both worth knowing before you sign.
The discipline here is the same one that separates serious growth investors from tourists: trust the cohorts, distrust the blend. Andreessen Horowitz's analysis of how AI-native businesses scale differently from classic SaaS keeps returning to the same point, the unit economics only become legible at the cohort level, and the aggregate can hide both the upside and the rot.
A Practical Diligence Checklist
If you are running this in a real deal process, here is the short list that catches the most problems in the least time:
- Reconstruct ARR from transactions, tagged by contract type, budget source, and workflow. Refuse the company's summary version.
- Compute gross and net retention by cohort, never blended. Flag any gap between an impressive net number and a mediocre gross number.
- Quantify pilot-sourced revenue as a percentage of run-rate, and model a 40-70% cut to it.
- Get true post-inference gross margin per cohort, and model a model-cost shock in both the margin-expansion and price-war branches.
- Map insourcing exposure customer by customer: what does a foundation-lab or incumbent release do to each cohort?
- Pull reliability metrics over time, task success rate, human-intervention rate, and treat declining reliability as a revenue risk, not just a product one.
- Check concentration: how much run-rate sits in the top five accounts, and how many of those are on uncommitted consumption with no floor?
Run those seven and you will know more about the durability of an agent company's revenue than most term sheets are written on. The output is not a single number but a structured view: how much of this revenue is annuity, how much is option, and how much is mirage. That distinction is what you are actually buying.
Insights Most People Overlook
Usage-based pricing is a durability liability dressed as a growth asset. Everyone celebrates consumption revenue because it expands without a sales motion. The flip side is rarely priced in: the same lack of friction that lets revenue grow lets it shrink. A customer can cut an agent's usage in half with a config change and no awkward cancellation call. Committed contracts trade some upside for a floor, and in a downturn that floor is the only thing standing between you and a 30% revenue air pocket. The companies that look least exciting on a usage-growth chart are sometimes the most durable.
Falling model costs can be bearish for an agent company, not bullish. The market reflexively treats cheaper inference as margin upside. But if a company's pricing is anchored to token cost rather than to the value of the outcome, cheaper models mean the company has to lower its prices, and its revenue per task falls in lockstep with its costs. Worse, cheap capable models lower the barrier for the customer to just build the thing themselves. The agent companies that benefit from cheaper inference are the ones priced on outcomes; the ones priced on consumption may watch their revenue compress as the inputs they resell get commoditized.
Net revenue retention above 130% can be a red flag, not a trophy. A spectacular NRR usually means a few accounts are expanding explosively. That is concentration risk wearing a party hat. If three customers drive most of your net expansion, your "best-in-class retention" is one reorg or one budget review away from inverting. Durable revenue tends to show boring NRR built on broad, shallow expansion across many accounts rather than fireworks from a handful.
The most durable agent revenue often hides in the least glamorous workflows. Agents doing exciting, frontier-capability work attract the most pilot dollars and the most insourcing risk, because everyone wants to build there. Agents doing tedious, deeply embedded, compliance-laden work nobody wants to touch, reconciliations, claims, document drudgery, generate revenue that is sticky precisely because ripping it out is risky and unrewarding. When you stress-test, the boring cohorts frequently survive every scenario while the flashy ones evaporate.
Reliability is a revenue metric with a delay. A company can grow through a quietly declining task-success rate for several quarters, because trust erodes slowly and churn lags. By the time reliability shows up in the revenue numbers, the damage is structural. The stress test that pulls reliability trends forward, treating a flat or declining success rate as a leading indicator of future revenue contraction, catches what the trailing financials will not show for a year.
References
More in Market
- The Late-Stage Investor's Agent Due-Diligence Checklist: What to Verify Before You Write the Check
- The Exit Landscape: Who Actually Buys Agent Startups
- Is the GaaS Valuation Bubble Real? Inside the Agentic AI Funding Debate
- SPVs and the Retail Rush Into Agent Investing
- Bootstrap-to-Scale: How to Build a Profitable Agent Company Without Taking VC Money