Agent-native article available: Enterprise AI Pipelines Don't Lose Money on Tokens: They Lose It BeforeAgent-native article JSON available: Enterprise AI Pipelines Don't Lose Money on Tokens: They Lose It Before
Enterprise AI Pipelines Don't Lose Money on Tokens: They Lose It Before

Enterprise AI Pipelines Don't Lose Money on Tokens: They Lose It Before

There comes a moment when the accumulation of artificial intelligence pilots stops looking like ambition and starts looking like disorder. That moment arrived for many large enterprises in 2026, and the clearest signal wasn't a technological collapse or a model failure. It was something more mundane and harder to defend in a board meeting: token consumption ran ahead of budget without generating proportional value.

Ignacio SilvaIgnacio SilvaJuly 21, 20269 min
Share

The enterprise AI pipeline doesn't lose money on tokens: it loses it before that

There is a moment when the accumulation of artificial intelligence pilots stops looking like ambition and starts looking like chaos. That moment arrived for many large companies in 2026, and the clearest signal was not a technological collapse or a model failure. It was something more mundane and harder to defend in a board meeting: token consumption ran ahead of budget without generating proportional value.

Uber was one of the cases that came to light. The company adjusted its internal AI spending after consumption exceeded what had been planned. It was not a strange exception; it was the visible symptom of a pattern that repeats itself across organizations that adopted AI with a maximalist logic during the previous two years: more use cases, more deployed agents, more employees incorporated into the system, more tokens consumed. The logic was defensible at first. When a technology is new and its potential for transformation is not yet clear, broad exploration makes sense. The problem is that such exploration did not stop when it should have given way to a deliberate architecture.

Sumeet Agrawal, Vice President of Product Management for Data, AI Governance, and Context Engineering at Salesforce, published a diagnosis in Fortune that deserves more attention than a corporate opinion column typically receives. His central argument is precise: lowering the price of tokens does not solve the problem because the problem is not in the price. It lies in how companies are architected to use them.

A pipeline that leaks at every stage

The metaphor Agrawal uses is useful because it is exact: the modern AI pipeline inside a large enterprise behaves like a sieve. It filters tokens, and with them money, at every phase of execution. And it does not do so by accident; it does so by design. Or more precisely, by the absence of design.

When an agent receives a sales or customer service query, the first thing it does — if it does not have a well-built data infrastructure underneath — is flood the prompt with raw context. Uncurated data, duplicate records, unprioritized history. The model then processes that mass of information, most of which is noise. According to Agrawal, that excess can mean five to ten times more tokens than necessary per interaction. At prices of between ten and fifteen dollars per million tokens, and with thousands of daily interactions at a medium-to-large-scale company, the arithmetic becomes hard to ignore very quickly.

The second point of leakage is ungoverned access to data. Without a clear catalog, without lineage, without quality signals, agents navigate data stores searching for reliable information. The process is slow, expensive, and produces inconsistent results. Governance, when it exists at all, tends to function as an after-the-fact control rather than a routing signal that directs the agent toward certified data from the very first attempt.

The third point of leakage is perhaps the most costly in terms of direct budget: sending every task to the largest available model, regardless of the complexity of the task. A routine classification or a simple search does not require the same model as complex reasoning or a sensitive decision. Treating every case with the same frontier model is the organizational equivalent of deploying a team of senior executives to handle tasks that a junior analyst could resolve: technically possible, functionally absurd.

The two remaining points of leakage are less visible but equally costly. Agents without persistent memory start every interaction from scratch: they reload context, reprocess history, and rediscover exceptions that had already been resolved. And agents without reusable semantics regenerate responses that could have been cached or precomputed. Every recurring interaction is paid for as if it were the first.

What vendors cannot resolve for you

Anthropic, OpenAI, and Google have lowered input token prices and launched prompt caching mechanisms. Cursor, in its Composer 2.5 version, incorporates cost as a variable in model selection, not only performance. These are rational responses to customer pressure, but they attack the wrong variable if the company has not yet resolved its internal architecture problems.

Reducing the price per token in a system that consumes ten times more tokens than necessary produces a proportional saving, but it does not close the structural gap. It is a linear improvement applied to a problem that has an order-of-magnitude solution. The company that resolves its architecture first gains an advantage that price reductions cannot replicate, because that advantage does not lie in the token market: it lies in the quality of its own data, in the governance of its workflows, and in the capacity to route work to the right model according to the nature of each task.

Agrawal formulates this with clarity: any company can buy more tokens. Very few know how to extract more value from fewer tokens. The difference between the two is not technological in the narrow sense of the term. It is architectural and organizational.

The example he offers is concrete: the integration between Informatica's master data management and Data 360, Salesforce's customer data platform, ensures that each agent operates on verified and semantically enriched customer context. The result is not merely token efficiency: it is the conversion of ungoverned AI consumption into measurable, auditable business value.

The true cost of absent design

From an organizational design perspective, what Agrawal describes is neither a technological problem nor a pricing problem. It is the deferred cost of having skipped the phase of disciplined exploration to settle prematurely into an exploitation phase of a technology that did not yet have the foundations to be exploited efficiently.

The companies that adopted AI with maximalist logic between 2024 and 2025 did so under legitimate pressure: the uncertainty about which models, which workflows, and which teams would generate value justified a broad deployment strategy. What did not justify itself — and what many organizations failed to do — was building in parallel the data and governance infrastructure that would determine whether that deployment would scale sustainably or simply accumulate technical debt.

The problem is not that they explored. It is that they explored without any underlying design. And now that absent design presents itself in the form of token bills that exceed plans and results that cannot be attributed to specific investments.

There is a pattern in cases of enterprise technology adoption that is worth naming: organizations tend to measure too soon with the wrong criteria, condemning initiatives that should not yet be judged by the same metrics as the core business. But they also tend to leave initiatives without any metric at all for too long — initiatives that should already be producing value. With enterprise AI, many companies did the second: they deployed without measuring or architecting, and now they face the adjustment from a position of greater disorder and greater accumulated cost.

The correction is not expensive in absolute terms. A well-built data catalog, quality signals that function as routers, persistent memory for agents, clear rules for assigning models according to task complexity: none of these decisions requires an exceptional budget. They require something harder to obtain in organizations that are already in scale mode: deliberate priority and architectural discipline sustained over time.

The advantage that cannot be bought in the model market

Agrawal frames token efficiency as the next competitive advantage in enterprise AI. The reading is correct but can be refined. The real advantage does not lie in token efficiency as an isolated metric. It lies in the organizational capacity to convert proprietary data into reliable context for agents that operate at scale, with sufficient governance so that results are auditable and attributable.

That is not a capability that can be purchased from a model provider or obtained by reducing the price per million tokens. It is a capability that is built internally, with data architecture decisions that precede agent deployment — not the other way around. Companies that already have that infrastructure gain an advantage that amplifies over time: each interaction is cheaper, faster, and more reliable than the one before. Those that do not have it face costs that do not fall simply because token prices drop.

The language model market will continue to get cheaper over time. That is almost certain. But the gap between companies that know how to use AI efficiently and those that do not will remain a problem of organizational design, data quality, and governance. And that gap has no solution in any external vendor's catalog.

The organizations that in 2026 are still operating with stateless agents, without functional data catalogs, and sending every workload to the most expensive available model, are not paying for tokens. They are paying the deferred price of not having designed their AI infrastructure when it was still cheap to do so.

Share

You might also like