The Economics of "Unlimited Agent" Plans: Why "All You Can Eat" Is the Riskiest Bet in GaaS
"Unlimited" in agent-as-a-service is a marketing word, not an economic one. Behind every flat-rate "unlimited agent" plan sits a vendor quietly betting that the average customer won't use much, because the marginal cost of an autonomous agent isn't zero like a SaaS seat, it's a live inference bill that grows with every task. This piece breaks down who actually wins on unlimited plans, the three failure modes that bankrupt naive vendors, and the hidden fair-use machinery that makes "unlimited" survivable. If you're pricing or buying agents, the question isn't whether unlimited is generous, it's who's absorbing the variance.
Table of Contents
- Why "Unlimited" Means Something Different for Agents
- The Cost Curve That Breaks the SaaS Playbook
- Who Actually Profits From an Unlimited Plan
- The Three Ways Unlimited Plans Go Bankrupt
- The Fair-Use Machinery Behind the Word "Unlimited"
- When Unlimited Is the Right Call
- A Back-of-Envelope Model for Vendors
- What Buyers Should Actually Ask
- Insights Most People Overlook
- References
Why "Unlimited" Means Something Different for Agents
In classic SaaS, "unlimited" is nearly free to offer. Once you've built the software and provisioned the servers, an extra user clicking around your dashboard costs you fractions of a cent. That's why "unlimited projects," "unlimited storage" (within reason), and "unlimited users" became standard pricing flourishes, the marginal cost of usage rounds to zero, so the generosity is mostly theatrical.
Agentic AI breaks that assumption at the foundation. When a customer runs an agent, they aren't loading a cached page. They're triggering chains of model calls, often dozens of reasoning steps, tool invocations, retries, and verification passes, each one metered in tokens that the vendor pays a model provider for in real money. An "unlimited agent" plan is therefore a promise to absorb an open-ended, usage-correlated cost. That's not generosity; it's risk transfer. The vendor is taking the variance off the customer's books and putting it onto their own.
This is the central tension of the entire GaaS pricing taxonomy. Per-task and per-outcome models pass cost variance through to the buyer. Unlimited absorbs it entirely. Everything in between, credits, floors-and-ceilings, hybrid base-plus-usage, is an attempt to split that variance in a way both sides can live with. Understanding unlimited well is the fastest way to understand why the rest of the pricing landscape looks the way it does.
The Cost Curve That Breaks the SaaS Playbook
Here's the uncomfortable shape of agent economics. A SaaS company has high fixed costs and near-zero marginal costs, which is why gross margins of 75 to 85 percent are normal and why the industry's whole financing model assumes them. Agent companies inherit a different curve: meaningful, persistent marginal cost per unit of work delivered.
Andreessen Horowitz flagged this early, noting that AI-native companies often run structurally lower gross margins than traditional software because inference is a real cost of goods sold, not an amortizable fixed cost. Their analysis of the new business of AI and how it's different from traditional software is still the clearest articulation of why you can't just bolt SaaS pricing onto an inference-heavy product and expect the margins to follow.
What this means for unlimited plans specifically: every customer you sign on a flat fee is a customer whose costs can climb above their price. In SaaS, your worst-case customer costs you a bit of support time. In GaaS, your worst-case unlimited customer can cost you more in inference than they pay you, and they can do it quietly, automatically, at 3 a.m., because the agent never sleeps and never gets bored. The downside isn't capped by human attention spans the way classic "unlimited" usage was.
The saving grace, and it's a real one, is that inference costs have fallen fast and keep falling. Frontier-model token prices have dropped by roughly an order of magnitude across several generations, and cheaper distilled or routed models can do much of the work the flagship model used to. A vendor pricing unlimited today is implicitly betting that their cost-to-serve will keep deflating faster than their heaviest users can escalate. That's a defensible bet in some markets and a reckless one in others. The whole calculus connects directly to margin expansion via model routing, sending the cheap model whenever it's good enough is often the only thing that makes unlimited survivable.
Who Actually Profits From an Unlimited Plan
Every flat-rate plan is a cross-subsidy. The light users fund the heavy ones, and the vendor pockets the spread on the average. This works beautifully in one specific condition: when usage is concentrated in a long, thin tail and the median customer barely touches the product.
Gym memberships are the canonical example, and the analogy is exact. The gym prices for the person who comes twice a week, profits enormously from the person who pays and never shows, and tolerates the handful of fanatics who basically live there. The fanatics would bankrupt the gym if everyone were a fanatic, but they aren't, so the model holds.
Unlimited agent plans work the same way, and they break the same way. They're profitable precisely when most buyers under-use, when there's a wide gap between the median and the mean, and when the heavy tail is small enough to subsidize. They become catastrophic when usage is uniform and high, when every customer is a gym fanatic, because then there's no light-user surplus to cover the heavy load.
This is why unlimited tends to appear in two very different places. First, in prosumer and SMB tiers where buyers are erratic, easily distracted, and rarely sophisticated enough to maximize an autonomous workflow. Second, as a deliberate land-grab in competitive markets, where a vendor eats negative unit economics on purpose to win logos and starve rivals, a strategy adjacent to the discounting death spiral in early GaaS deals, and just as dangerous if you can't exit it. The difference between a smart unlimited plan and a fatal one is almost entirely about which of those two situations you're actually in, and whether you're honest with yourself about it.
The Three Ways Unlimited Plans Go Bankrupt
I've watched enough usage-based businesses to recognize the recurring failure patterns. For unlimited agent plans, three stand out.
Failure one: the power-user concentration
A small cohort of customers discovers they can point your agent at enormous, repetitive workloads and let it run continuously. They're not abusing anything, they're using exactly what you sold them. But their cost-to-serve dwarfs their flat fee, and because agents scale work without human bottlenecks, a single determined power user can consume more inference than a hundred median accounts combined. Metered models naturally bill these users more; unlimited eats the cost. This is the same dynamic explored in why metered pricing rewards power users, seen from the vendor's side of the ledger.
Failure two: the automation flywheel
Agents beget agent usage. A customer wires your agent into a trigger, every inbound email, every new record, every webhook, and now usage grows on its own with the customer's business, not with any conscious decision to use more. This is wonderful for retention and lethal for an unlimited plan. The whole land-and-expand motion in GaaS is built on the fact that expansion is automatic, but automatic expansion under a flat fee is automatic margin erosion. What's a feature under usage pricing is a bug under unlimited.
Failure three: the model-cost shock
You priced unlimited assuming a certain cost-per-task. Then a capability you depend on requires a more expensive model, or a provider raises prices, or your routing logic regresses and starts sending easy tasks to the flagship model. Your cost-to-serve jumps and your revenue doesn't move, because you've contractually frozen it. Pass-through is impossible by definition, that's the whole point of unlimited, so the margin compression lands entirely on you. This is the inverse of the problem in margin-safe pricing for volatile inference costs: unlimited is the one model that structurally forbids passing volatility through.
The Fair-Use Machinery Behind the Word "Unlimited"
Here's the part the marketing page won't tell you: almost no "unlimited" agent plan is actually unlimited. It's unlimited the way an "unlimited" mobile data plan is unlimited, generous until you hit an invisible line, after which something quietly changes.
The mechanisms vendors use to make unlimited survivable are worth naming, because as a buyer you're agreeing to them whether you read them or not:
- Soft concurrency caps. You can run unlimited tasks, but only so many at once. This flattens the cost spike from a power user trying to run ten thousand parallel jobs without technically limiting their total volume.
- Throttling and queueing. Past a threshold, your tasks slow down. Throughput, not total quantity, becomes the real constraint, a subtle form of the protection described in how usage caps protect customers and vendors alike.
- Fair-use clauses. The terms of service reserve the right to limit "excessive" or "abusive" usage, with the vendor defining both words. This is the escape hatch that makes the word "unlimited" legally defensible.
- Model downgrade under load. Heavy accounts get silently routed to cheaper models. The work still gets done; it just gets done with less expensive inference, protecting margin without touching the headline price.
- Autonomy ceilings. Some "unlimited" plans cap the depth or breadth of what an agent will do per task, fewer reasoning steps, fewer tool calls, which connects to pricing tiers based on autonomy level even when it's not advertised as such.
None of this is necessarily dishonest. The honest version of unlimited is "unlimited for any reasonable use, and here's roughly what reasonable means." The dishonest version hides the caps until a customer hits them and feels cheated. The difference is disclosure, and the vendors who get this right tend to be the same ones who think hard about the psychology of metered billing and buyer anxiety, they understand that the entire appeal of unlimited is the absence of anxiety, and that a surprise cap detonates exactly that.
When Unlimited Is the Right Call
Despite all the risk, unlimited is sometimes genuinely the correct choice, and not just as a loss-leader.
It works when your cost-to-serve per task is low and stable enough that even your heaviest plausible user stays profitable. If model routing keeps your marginal cost trivial, the variance you're absorbing is small, and "unlimited" becomes a cheap, powerful marketing asset rather than a liability.
It works when the buyer's number-one objection is bill anxiety. Some segments, especially teams burned by surprise cloud bills, will pay a premium for cost certainty and refuse anything metered on principle. For them, unlimited isn't about getting more; it's about never having to think about it again. This is the same buyer psychology pushing some vendors back toward flat pricing, a trend examined in why some GaaS vendors are returning to flat pricing.
And it works as a deliberate, time-boxed strategy: unlimited during a land grab to win share, with a planned migration to usage-based or hybrid pricing once you've locked in the relationship and the switching costs. The danger there is the grandfather problem, customers you signed on unlimited will fight to keep it, and repricing them later is one of the hardest conversations in the business. If you go unlimited as a strategy, write the exit plan before you launch it, not after.
A Back-of-Envelope Model for Vendors
You don't need a finance team to sanity-check an unlimited plan. You need four numbers.
Take your median customer's monthly task volume and your mean (average) volume. The wider the gap between them, the more the long tail is skewing your costs, and the more dangerous unlimited becomes. A median of 200 tasks against a mean of 900 tells you a small group is doing enormous damage, and that group will only grow as customers automate.
Multiply your mean task volume by your cost-per-task (all-in inference, not just the headline token price, include retries, verification, and tool calls). That's your average cost-to-serve. Your unlimited price needs to clear that with enough room to survive the cost-per-task climbing and the mean drifting upward over time, because both will.
Then ask the question that actually decides it: what does your 95th-percentile customer cost you, and can you survive a plan full of them? If the answer is no, you're not running an unlimited plan, you're running a bet that your worst customers never become your average ones. Sometimes that's a fine bet. Just know you're making it, and revisit it every quarter, because the mean in a healthy GaaS business only moves one direction.
What Buyers Should Actually Ask
If you're evaluating an unlimited agent plan, flip the analysis around. The vendor's risk is your leverage.
Ask where the fair-use line actually sits, in concrete numbers, and get it in writing. Ask what happens at that line, throttling, downgrade, or a hard stop, because the answer determines whether "unlimited" is real for your workload. Ask whether the model you're getting changes under load, since a silent downgrade can quietly gut the quality you're paying for. And ask what happens at renewal, because a vendor losing money on your account will come back to the table, and you want to know now whether you're the kind of customer they'll want to keep or quietly reprice.
If you're a light user, unlimited is often a genuine bargain, you're the one being subsidized, so take it. If you're a heavy, automated, mission-critical user, unlimited is the plan most likely to get yanked, throttled, or repriced out from under you, and a transparent usage-based or floor-and-ceiling deal may actually be safer despite looking scarier on the pricing page. The comfortable-sounding plan and the durable plan are not always the same plan.
Insights Most People Overlook
-
Unlimited transfers variance, not value. Buyers think unlimited means "more for less." Economically it means "the vendor now owns my cost variance." That ownership has to be priced in somewhere, usually as a higher base fee, a quality cap, or a fair-use trapdoor. There's no free lunch; there's only a relocated one. The buyer who understands this negotiates better than the one who just sees a flat number.
-
Falling inference costs are the only thing making unlimited viable, and they make it temporarily viable for everyone, which is the trap. When token prices drop, unlimited looks safe and competitors pile in. But the same deflation that lets you offer unlimited lets your rivals offer it too, compressing the whole market toward flat pricing right as usage automates upward. The cost curve that opens the door also crowds the room.
-
The most dangerous unlimited customer is the happiest one. In SaaS, your power users are your champions and your lowest churn risk. In an unlimited GaaS plan, your most engaged customer is also your biggest loss. You end up in the perverse position of quietly hoping your best users get a little bored, which is a terrible foundation for a business and a sign you've mispriced.
-
"Unlimited" is often a data-collection play disguised as a pricing play. Some vendors offer unlimited specifically to remove the meter so customers use the agent freely and generate maximum usage data. The flat fee isn't where the value is, the behavioral data and the workflow lock-in are. This connects to why outcome pricing favors incumbents with data: whoever sees the most agent runs builds the best agents, and unlimited is one of the cheapest ways to buy that visibility.
-
A fair-use cap that's never disclosed is a reputational time bomb with a usage-based fuse. The customers who hit the hidden cap are, by definition, your most committed users, the ones who'll write the angriest posts. Undisclosed caps don't just cost you those accounts; they convert your most engaged customers into your loudest critics at exactly the moment they were ready to expand.
References
More in Pricing
- Margin Expansion via Model Routing: Use the Cheap Model When You Can
- Pricing Agents by Seniority: How Junior vs. Senior Agent Tiers Actually Work
- Why Outcome Pricing Quietly Hands the Advantage to Incumbents With Data
- Cost-Plus vs. Value-Based: The GaaS Pricing Philosophy Debate
- Pricing Pilots vs. Production Deployments: The GaaS Trap Hiding in Your Pilot Invoice