THE INDEPENDENT RECORD · AGENTIC AI AS A SERVICE AboutStandardsContact
GAASAGENTIC AI · AS A SERVICE
INDEPENDENT · SINCE 2026
UPDATED DAILY
NO HYPE · NO PAY-TO-PLAY
PER-TASK PRICING NOW STANDARD ● NEW BENCHMARK: 71% TASK COMPLETION ● ENTERPRISE PILOTS UP 4X ● RUNTIME FUNDING ACCELERATES ● "AGENTS ARE THE NEW SEATS" ● MARGINS UNDER PRESSURE ● THE INDEPENDENT RECORD ON GAAS
Market

The "Picks and Shovels" Funding Thesis for Agent Infra: Why Smart Money Is Betting Below the Application Layer

The "picks and shovels" thesis says the surest way to profit from a gold rush is to sell the tools, not pan for gold yourself. In Agentic AI-as-a-Service (GaaS), that means funding the infrastructure agents run on, orchestration, memory, evals, observability, identity, payments, rather than betting on which vertical agent wins. The pitch is seductive: infra captures value no matter who wins at the app layer. The reality is more contested. Plenty of "picks" get absorbed into model providers or commoditized into open source. This piece separates the durable shovel businesses from the ones that are about to get bulldozed.

By J. Okafor · May 11, 2026 · 13 min read

Table of Contents

Where the Metaphor Comes From, and Where It Breaks

Levi Strauss didn't dig for gold. He sold denim to the people who did, and his name outlasted nearly every claim staked in California. The line "in a gold rush, sell shovels" has become Silicon Valley's favorite shorthand for backing the supplier instead of the prospector. NVIDIA is the canonical modern example: it didn't have to win the model race because every contender bought its chips.

So far so clean. The trouble is that the metaphor smuggles in an assumption people rarely state out loud, that the shovel is a durable, defensible product the prospectors can't make themselves. That held for picks and shovels in 1849 because forging steel tools required a foundry, distribution, and capital the average forty-niner didn't have. It holds for GPUs today because fabbing leading-edge silicon costs tens of billions and a decade of process knowledge.

It does not automatically hold for software. A lot of what gets pitched as "agent infrastructure" is a few thousand lines of orchestration code that a well-funded model lab can ship as a free SDK on a Tuesday. When that happens, your shovel company isn't Levi Strauss, it's a guy selling slightly nicer pans next to a hardware store that's giving pans away. The picks-and-shovels thesis is only as good as the moat under the specific pick. Investors who chant the slogan without interrogating the moat are doing pattern-matching, not analysis.

What Actually Counts as "Agent Infra"

Before you can underwrite the thesis you have to define the territory, because "infra" gets stretched to cover almost anything. In the GaaS stack, the infrastructure layer sits between the foundation models and the vertical agents that customers actually buy. Roughly, it breaks into these buckets:

Andreessen Horowitz has argued in its writing on the emerging LLM app stack that the orchestration and tooling layers are where a lot of net-new company formation is happening. That's the bull case for infra. The bear case is that several of those buckets are land the model labs consider their own backyard.

Why VCs Like the Thesis

Set aside whether it's right for a moment. The reasons the picks-and-shovels framing is catnip to investors are worth naming, because they explain a lot of where the money is flowing.

First, it sidesteps the hardest question in GaaS investing: which vertical agent wins? Picking the winning legal agent, or coding agent, or support agent requires a strong view on a specific market, a specific founder, and a specific wedge. Backing the infra "everyone needs" feels like a hedge, you're long the category without being short any particular bet. This is the same instinct behind index investing, and it's why infra rounds often draw the more risk-averse end of the capital spectrum. (How investors are actually underwriting these bets differently from classic SaaS is its own deep topic; see the related beat work on VC underwriting models.)

Second, infra promises horizontal scale. A vertical agent's market is capped by its vertical. An orchestration platform, in theory, serves every vertical at once, bigger TAM, better multiples. The valuation premium that agent companies command is partly a bet on exactly this kind of horizontal leverage.

Third, infra businesses can look like classic developer-tools plays, and dev tools have a beloved playbook: bottoms-up adoption, open-source distribution, land-and-expand into enterprise. VCs have made enormous returns on that motion (Datadog, MongoDB, HashiCorp, GitLab), so they recognize the shape and reach for their checkbooks faster.

The problem is that all three reasons are about why infra is attractive to fund, not about whether a given infra company is durable. Attractiveness and durability are different variables, and conflating them is how you end up with an over-funded category full of companies selling commoditized shovels.

The Layers Most Likely to Survive

Here's where I'll plant a flag. Not all agent infra is created equal, and the durable picks share a common trait: they get harder to replace as the system scales, not easier. Three categories clear that bar.

Reliability and evals. As agents move from demos to production, the question stops being "can it do the task?" and becomes "can it do the task 99.5% of the time, and can you prove it?" Evaluation infrastructure compounds in value because every deployed agent generates more edge cases that need testing, and that test corpus becomes proprietary. A model lab can give away a framework; it can't give away your accumulated golden datasets and your specific failure taxonomy. This is the layer I'd weight most heavily, and it ties directly to the revenue-durability questions investors are stress-testing across the GaaS market.

Observability and cost attribution. Agents are non-deterministic and expensive. The burn-rate problem in GaaS, agents that cost real money every time they run, means buyers desperately need to see where tokens and dollars are going. Observability gets stickier with usage because the historical traces are the asset. Migrating off your observability vendor means losing your history, which is exactly the lock-in that made Datadog hard to dislodge.

Identity, permissioning, and agent security. When agents act autonomously on behalf of users, moving money, sending email, editing production data, the security layer becomes load-bearing and regulated. Regulated, audited infrastructure is among the hardest to rip out, because replacing it triggers a compliance review. This is structurally defensible in a way a prompt-orchestration wrapper never will be.

The common thread: these layers accumulate proprietary state (test cases, traces, audit logs, permission graphs) that lives on the customer's side and grows with usage. State is the moat. Code is not.

The Layers Most Likely to Get Bulldozed

Now the uncomfortable half. Several popular "infra" categories are, in my read, sitting in the model labs' expansion path or on a collision course with open source.

Thin orchestration wrappers. If your entire product is "we make it easier to chain model calls and tools," you're building in the exact lane OpenAI, Anthropic, and Google are paving with their own agent SDKs and assistant APIs. Every time a lab ships native tool-use, function calling, or a built-in agent loop, it lops off a slice of the third-party orchestration market for free. McKinsey's research on the economic potential of generative AI underscores how much value concentrates around the model providers themselves, and labs have every incentive to climb the stack toward the applications, absorbing the orchestration layer on the way up.

Generic vector databases as standalone businesses. Retrieval matters, but "store embeddings and do nearest-neighbor search" turned out to be a feature, not a company, for many entrants. Postgres got pgvector. Every major cloud database bolted on vector search. The pure-play vector DB became a commoditized middle.

Single-model inference brokers. Routing to one provider's API with a markup is a margin business with no floor. The moment the lab's own pricing moves, your spread evaporates.

The pattern here is the inverse of the survivors: these layers hold little proprietary customer state, so switching costs are low and the incumbent above you (the lab) or below you (open source, the cloud) can absorb the function at near-zero marginal cost. When you evaluate an agent-infra pitch, the first question isn't "is this useful?", almost everything is useful. It's "what stops a model lab from shipping this next quarter, and what stops the open-source community from cloning it next month?"

How the Economics Actually Pencil Out

The picks-and-shovels thesis has a quiet financial assumption baked in: that infra margins are healthy. In classic dev tools, they are, software gross margins run 75-85%. But agent infra carries a cost the slogan ignores: a meaningful chunk of "infra" companies pass through model inference, which means their cost of goods rides on someone else's pricing.

If you're an eval platform that runs the customer's agent thousands of times against a test suite, every run burns tokens. If you mark that up modestly, you have a real margin. If you eat it as a loss leader to drive adoption, you have a SaaS valuation stapled to a cloud-reseller cost structure, and that's the kind of company that looks great until model costs and revenue quality get scrutinized in a down market. Valuation haircuts hit hardest where the reported "software" multiple was actually sitting on top of a pass-through cost base.

The durable infra businesses are the ones whose value isn't proportional to inference volume, the ones selling the system of record (the traces, the test corpus, the audit log) rather than reselling the system of compute. That distinction is the single best predictor I know of for whether an agent-infra company earns a software multiple or a services multiple at exit.

How to Pressure-Test an Agent-Infra Pitch

If you're an investor, an operator deciding what to build, or a founder writing the deck, here's the checklist I'd run before invoking the magic words "picks and shovels":

  1. State test. Does the product accumulate proprietary customer state that grows with usage and lives stickily on the customer's side? If yes, real moat. If it's stateless code, suspect.
  2. Lab-encroachment test. Draw the line between this product and the nearest model lab's roadmap. If they overlap, assume the lab wins the overlap for free. Build outside it.
  3. Open-source test. Is there an OSS project one GitHub-trending-week away from doing 80% of this? Charging for the managed/enterprise wrapper of an OSS core is viable (HashiCorp, GitLab), but only if the enterprise layer, security, compliance, support, is the real product.
  4. Margin-honesty test. Strip out inference pass-through. What's the true software gross margin underneath? That's the number that survives a down round.
  5. Regulatory-stickiness test. Does ripping this out trigger a compliance or security review? Audited infrastructure is the hardest to displace.

Pass three of five and you might have a Levi Strauss. Pass one and you're selling slightly-nicer pans.

Insights Most People Overlook

References

More in Market