THE INDEPENDENT RECORD · AGENTIC AI AS A SERVICE AboutStandardsContact
GAASAGENTIC AI · AS A SERVICE
INDEPENDENT · SINCE 2026
UPDATED DAILY
NO HYPE · NO PAY-TO-PLAY
PER-TASK PRICING NOW STANDARD ● NEW BENCHMARK: 71% TASK COMPLETION ● ENTERPRISE PILOTS UP 4X ● RUNTIME FUNDING ACCELERATES ● "AGENTS ARE THE NEW SEATS" ● MARGINS UNDER PRESSURE ● THE INDEPENDENT RECORD ON GAAS
Verticals

The Vertical-Agent Moat Is Workflow Data, Not Models

The defensibility of a vertical AI agent doesn't come from the model it calls. Foundation models are rented, swappable, and increasingly commoditized. The durable moat is proprietary workflow data: the messy, accumulated record of how a specific job actually gets done in a specific industry, including the edge cases, the corrections, and the human overrides that no public dataset contains. This piece explains why that data compounds, how it differs from generic training corpora, and where the moat is thinner than founders claim.

By J. Okafor · Jun 23, 2026 · 13 min read

Table of Contents

Why the Model Was Never the Moat

Spend ten minutes at any AI conference in 2026 and you'll hear a founder describe their "proprietary model." Press a little, and it usually turns out to be GPT-class or Claude-class weights with a system prompt, some retrieval, and a fine-tune on a few thousand examples. That is not a moat. It is a configuration, and configurations get copied in a weekend.

The uncomfortable truth for vertical-agent companies is that the intelligence layer is the part competitors can most easily match. When the underlying foundation model improves, and it improves roughly every quarter, everyone's agent gets smarter at once, including the one a rival shipped last month. a16z's analysis of the emerging agent stack makes the point bluntly: model capability is a rising tide, not a private island. If your differentiation rides on the model, you're renting your advantage from someone who will rent the same thing to the company trying to kill you.

So where does durable advantage live? In the one asset a foundation model can't supply and a competitor can't scrape: the proprietary record of how a particular kind of work gets done, correctly, thousands of times, inside one industry. This is the connective tissue of the entire GaaS cluster, it's why depth of integration becomes the new defensibility and why the system-of-record advantage matters so much. The agent is the interface. The workflow data is the asset.

What "Proprietary Workflow Data" Actually Means

People hear "proprietary data" and picture a big table of customer records. That's not it, and the distinction matters enormously.

Generic training data tells a model what language looks like. Workflow data tells an agent what good work looks like in a context that never appears on the public internet. Consider a medical-coding agent. The public web contains the ICD-10 code book, that's freely available reference material, and it gives you no edge. What it does not contain is the sequence of decisions a senior coder at a specific hospital system makes when a chart is ambiguous: which modifier they reach for when a payer historically denies a claim, how they interpret a physician's shorthand, the corrections they apply after a denial comes back. That accumulated judgment is the moat, and it lives nowhere except inside the workflow.

The same pattern holds across verticals. A contract-review agent becomes defensible not because it can read a lease, but because it has seen ten thousand redlines from one firm's partners and learned which clauses this firm always fights and which it concedes. A prior-authorization agent earns its keep by knowing the undocumented quirks of how a regional insurer actually adjudicates, knowledge that exists in approval-and-denial patterns, not in any published policy.

The key property: this data is generated by the work itself. You can't buy it, and you usually can't even collect it until you're already inside the workflow doing the job. That chicken-and-egg quality is exactly what makes it defensible once you have it.

The Three Layers of Workflow Data

Not all workflow data is equally valuable. It helps to think in three layers, each harder to acquire and each more defensible than the last.

Process Traces

The first layer is the raw record of steps taken to complete a task: which systems were touched, in what order, what inputs produced what outputs. A logistics dispatch agent that has logged a hundred thousand routing decisions holds a map of how dispatch actually flows in practice, including the workarounds humans invented that no process document describes. Process traces are valuable but the most copyable layer, because a competitor who lands the same customer starts generating their own.

Edge Cases and Exceptions

The second layer is where real defensibility starts: the long tail of weird situations that break naive automation. Any agent can handle the 80% of cases that look like the textbook. The money, and the moat, is in the 20% that don't. A claims-adjudication agent that has encountered every flavor of fraudulent, malformed, or genuinely unusual claim has an exception library a new entrant simply cannot reconstruct from scratch. Edge cases accumulate slowly and unevenly, which is precisely why they're hard to copy: a rival can't fast-forward through three years of strange Tuesdays.

Outcome Labels

The third and most valuable layer is the record of what worked. Did the contract clause the agent flagged actually cause a dispute later? Did the prior-auth strategy get approved? Did the dispatched route arrive on time? Outcome data closes the loop between action and result, and it is brutally hard to obtain because outcomes arrive late, attribute messily, and require staying embedded in the customer's operation long enough to observe them. An agent with outcome labels isn't just imitating expert behavior, it's learning which expert behaviors were actually right. This is the layer that turns a per-task tool into something closer to outcome-priced infrastructure.

How the Moat Compounds: The Outcome Feedback Loop

A static dataset is a head start, not a moat. What makes workflow data defensible is that it compounds, and it compounds through a specific mechanism: the outcome feedback loop.

Here's how it runs. The agent does the work. A human reviews, corrects, or overrides some fraction of it. Those corrections become new training signal. The agent gets better, which means it handles more volume, which means it observes more cases, including rarer edge cases it never would have reached at lower volume. Each turn of the loop widens the gap between the incumbent and anyone starting fresh.

This is a genuine data network effect, but it's worth being precise, because the term gets abused. The classic network effect (more users make the product more valuable to other users) often doesn't apply to vertical agents, one hospital's coding data doesn't directly help another hospital, and may be legally walled off. What does apply is a learning effect: more usage makes the product better at its job, which compounds within an account and, when data can be pooled, across them. Researchers have noted that these data-driven learning effects behave differently from traditional network effects, they can plateau, and they're vulnerable to a competitor who finds a steeper learning curve. The moat is real, but it's a flywheel you have to keep spinning, not a wall you build once.

The compounding is strongest where the loop is tightest. Coding agents are a clean example: a coding agent that ships autonomous PRs gets a near-immediate outcome signal, the tests pass or they don't, the PR merges or gets reverted. That fast, unambiguous feedback lets the data moat thicken quickly. Compare that to a wealth-management agent, where the "outcome" of advice may not be legible for years. Tighter loops build moats faster.

Why Horizontal Platforms Struggle to Copy It

The recurring fear for every vertical-agent founder is that a horizontal giant, OpenAI, a hyperscaler, a Salesforce, will simply absorb their category. Sometimes that fear is justified (sometimes a horizontal platform does eat your vertical agent). But the workflow-data moat is one of the few things that genuinely slows the giants down, for structural reasons.

A horizontal platform optimizes for breadth. Its data advantage is wide and shallow, it knows a little about everything. The vertical agent's advantage is the inverse: narrow and deep. To replicate it, the horizontal player would have to embed itself in the same specific workflow, earn the same trust, survive the same regulatory review, and wait out the same multi-year accumulation of edge cases and outcomes. That's not a capability gap they can close with a bigger model; it's a time-in-the-workflow gap.

Regulation deepens the trench. In regulated industries, healthcare, legal, insurance, banking, the workflow data is often legally bound to the customer and the use case. A general-purpose platform can't hoover it up into a shared corpus without tripping HIPAA, attorney-client privilege, or financial-data rules. The vertical specialist, who has built the compliance posture and signed the business-associate agreements, can use that data where the generalist cannot. The regulatory wall that looks like a burden is also a moat.

This is the deeper argument behind why vertical agents beat horizontal platforms when they do: it's rarely about being smarter. It's about being somewhere the platform can't easily go, accumulating something the platform can't easily take.

Where the Moat Is Thinner Than Founders Claim

Now the contrarian half, because the workflow-data moat is real but routinely overstated, and pretending otherwise gets companies killed.

First: a moat made of a few thousand examples is not a moat. There's a threshold problem nobody likes to discuss. Below some volume of data, your "proprietary" advantage is within fine-tuning distance of any well-resourced competitor, and as base models get better at few-shot learning, that threshold keeps rising. The capability that used to require ten thousand of your examples now needs five hundred, or a clever prompt. Every model upgrade erodes the bottom of your moat.

Second: customer-owned data isn't yours. Many founders quietly assume the data their agent generates inside a customer's account is theirs to learn from across the customer base. Often it isn't, contractually, the customer owns it, and a savvy enterprise buyer will forbid you from training shared models on their data. If your moat thesis depends on pooling data you're not actually allowed to pool, you don't have a moat; you have a lawsuit waiting to happen and a flywheel that never turns.

Third: data decays. Workflows change. A new regulation, a payer policy update, a software migration on the customer's side, and a chunk of your hard-won edge-case library is suddenly describing a world that no longer exists. Workflow-data moats require constant refresh; they are not annuities.

Fourth, and most overlooked: a competitor doesn't need your data, they need enough data. If a rival lands three big customers in your vertical and those customers generate sufficient volume, the rival's learning curve may simply be steeper than yours, especially if they architected for the feedback loop from day one and you bolted it on later. The moat protects the company that compounds fastest, not the one that started first.

Building a Data Moat on Purpose

If the moat is workflow data, then the strategic implication is that you should engineer for data capture from the very first line of product, not discover it later. A few principles that separate companies who build a moat from those who merely accumulate logs.

Instrument the corrections. The single most valuable signal is the human override, every time a user fixes the agent's output, that's a labeled training example landing in your lap for free. Capture it structured, capture the why where you can, and treat your correction-capture UX as core product, not an afterthought.

Close the outcome loop deliberately. Most agents never find out whether their work was right. Architect to observe outcomes, integrate deep enough into the system of record that you can see the downstream result, even if it arrives weeks later. The companies that win regulated verticals are usually the ones that stayed embedded long enough to learn from outcomes, not just actions.

Pick verticals with tight feedback loops first. If you have a choice of where to start, favor domains where the outcome is fast and legible. The moat compounds faster, and you'll out-learn slower-feedback competitors. This is part of why the services-to-software flip works so well, agencies that became agent companies already sat inside the workflow, watching outcomes, for years before they ever wrote code.

The honest summary: the model is a commodity you rent, the integration is a barrier you build, and the workflow data is the asset that compounds. Foundation models will keep getting better and that helps everyone equally. The only thing that helps you specifically is the accumulating, corrected, outcome-labeled record of a job well done, captured somewhere your competitors can't reach.

Insights Most People Overlook

The moat is in the corrections, not the completions. Most teams obsess over making the agent's first-pass output better. The defensible asset is actually the stream of human overrides, the moments the agent was wrong and a domain expert fixed it. Those labeled corrections are worth more than a hundred successful completions, and the companies that win are the ones whose product is quietly engineered to capture them.

"Data network effects" is usually the wrong term, and using it wrong is dangerous. Vertical agents rarely have classic network effects (one customer's data legally can't help another's). What they have is a learning effect that compounds within an account. Founders who pitch cross-customer network effects often can't legally deliver them, and investors who underwrite that thesis are underwriting something the contracts forbid.

Better base models erode the bottom of your moat while raising the top. Each foundation-model upgrade does two opposite things: it makes your agent better (good) and it lowers the amount of proprietary data a competitor needs to match you (bad). Sub-threshold data moats quietly evaporate with every model release. Your moat has to grow faster than few-shot learning improves.

Speed of the feedback loop matters more than the size of the dataset. A company with less data but a same-day outcome signal will out-compound a company with more data and a six-month signal. When choosing a vertical, loop latency is a more important variable than market size, and almost nobody picks their market that way.

The regulatory burden and the moat are the same object. Founders treat compliance as a cost center. In regulated verticals it's the wall that keeps horizontal platforms out of your data. The BAAs, the privilege handling, the data-residency work, that's not overhead slowing you down, that's the trench. Companies that resent the compliance work are resenting their own moat.

References

#vertical agent moat#vertical ai agents

More in Verticals