{"version":"1.0","type":"agent_native_article","locale":"en","slug":"enterprise-ai-pipeline-loses-money-before-tokens-mrusqsoo","title":"Enterprise AI Pipelines Don't Lose Money on Tokens: They Lose It Before","primary_category":"innovation","author":{"name":"Ignacio Silva","slug":"ignacio-silva"},"published_at":"2026-07-21T14:02:33.545Z","total_votes":86,"comment_count":0,"has_map":true,"urls":{"human":"https://sustainabl.net/en/articulo/enterprise-ai-pipeline-loses-money-before-tokens-mrusqsoo","agent":"https://sustainabl.net/agent-native/en/articulo/enterprise-ai-pipeline-loses-money-before-tokens-mrusqsoo"},"summary":{"one_line":"Enterprise AI cost overruns are not a token pricing problem but an architectural and organizational design failure that precedes token consumption.","core_question":"Why do large enterprises overspend on AI despite falling token prices, and what structural changes would actually close the value gap?","main_thesis":"The root cause of enterprise AI budget overruns is not token price but the absence of data architecture, governance, and routing discipline before agents are deployed. Companies that fix their internal infrastructure gain a compounding advantage that no external vendor price reduction can replicate."},"content_markdown":"## The enterprise AI pipeline doesn't lose money on tokens: it loses it before that\n\nThere is a moment when the accumulation of artificial intelligence pilots stops looking like ambition and starts looking like chaos. That moment arrived for many large companies in 2026, and the clearest signal was not a technological collapse or a model failure. It was something more mundane and harder to defend in a board meeting: token consumption ran ahead of budget without generating proportional value.\n\nUber was one of the cases that came to light. The company adjusted its internal AI spending after consumption exceeded what had been planned. It was not a strange exception; it was the visible symptom of a pattern that repeats itself across organizations that adopted AI with a maximalist logic during the previous two years: more use cases, more deployed agents, more employees incorporated into the system, more tokens consumed. The logic was defensible at first. When a technology is new and its potential for transformation is not yet clear, broad exploration makes sense. The problem is that such exploration did not stop when it should have given way to a deliberate architecture.\n\nSumeet Agrawal, Vice President of Product Management for Data, AI Governance, and Context Engineering at Salesforce, published a diagnosis in Fortune that deserves more attention than a corporate opinion column typically receives. His central argument is precise: lowering the price of tokens does not solve the problem because the problem is not in the price. It lies in how companies are architected to use them.\n\n## A pipeline that leaks at every stage\n\nThe metaphor Agrawal uses is useful because it is exact: the modern AI pipeline inside a large enterprise behaves like a sieve. It filters tokens, and with them money, at every phase of execution. And it does not do so by accident; it does so by design. Or more precisely, by the absence of design.\n\nWhen an agent receives a sales or customer service query, the first thing it does — if it does not have a well-built data infrastructure underneath — is flood the prompt with raw context. Uncurated data, duplicate records, unprioritized history. The model then processes that mass of information, most of which is noise. According to Agrawal, that excess can mean **five to ten times more tokens than necessary per interaction**. At prices of between ten and fifteen dollars per million tokens, and with thousands of daily interactions at a medium-to-large-scale company, the arithmetic becomes hard to ignore very quickly.\n\nThe second point of leakage is ungoverned access to data. Without a clear catalog, without lineage, without quality signals, agents navigate data stores searching for reliable information. The process is slow, expensive, and produces inconsistent results. Governance, when it exists at all, tends to function as an after-the-fact control rather than a routing signal that directs the agent toward certified data from the very first attempt.\n\nThe third point of leakage is perhaps the most costly in terms of direct budget: sending every task to the largest available model, regardless of the complexity of the task. A routine classification or a simple search does not require the same model as complex reasoning or a sensitive decision. Treating every case with the same frontier model is the organizational equivalent of deploying a team of senior executives to handle tasks that a junior analyst could resolve: technically possible, functionally absurd.\n\nThe two remaining points of leakage are less visible but equally costly. Agents without persistent memory start every interaction from scratch: they reload context, reprocess history, and rediscover exceptions that had already been resolved. And agents without reusable semantics regenerate responses that could have been cached or precomputed. Every recurring interaction is paid for as if it were the first.\n\n## What vendors cannot resolve for you\n\nAnthropic, OpenAI, and Google have lowered input token prices and launched prompt caching mechanisms. Cursor, in its Composer 2.5 version, incorporates cost as a variable in model selection, not only performance. These are rational responses to customer pressure, but they attack the wrong variable if the company has not yet resolved its internal architecture problems.\n\nReducing the price per token in a system that consumes ten times more tokens than necessary produces a proportional saving, but it does not close the structural gap. It is a linear improvement applied to a problem that has an order-of-magnitude solution. The company that resolves its architecture first gains an advantage that price reductions cannot replicate, because that advantage does not lie in the token market: it lies in the quality of its own data, in the governance of its workflows, and in the capacity to route work to the right model according to the nature of each task.\n\nAgrawal formulates this with clarity: any company can buy more tokens. Very few know how to extract more value from fewer tokens. The difference between the two is not technological in the narrow sense of the term. It is architectural and organizational.\n\nThe example he offers is concrete: the integration between Informatica's master data management and Data 360, Salesforce's customer data platform, ensures that each agent operates on verified and semantically enriched customer context. The result is not merely token efficiency: it is the conversion of ungoverned AI consumption into measurable, auditable business value.\n\n## The true cost of absent design\n\nFrom an organizational design perspective, what Agrawal describes is neither a technological problem nor a pricing problem. It is the deferred cost of having skipped the phase of disciplined exploration to settle prematurely into an exploitation phase of a technology that did not yet have the foundations to be exploited efficiently.\n\nThe companies that adopted AI with maximalist logic between 2024 and 2025 did so under legitimate pressure: the uncertainty about which models, which workflows, and which teams would generate value justified a broad deployment strategy. What did not justify itself — and what many organizations failed to do — was building in parallel the data and governance infrastructure that would determine whether that deployment would scale sustainably or simply accumulate technical debt.\n\nThe problem is not that they explored. It is that they explored without any underlying design. And now that absent design presents itself in the form of token bills that exceed plans and results that cannot be attributed to specific investments.\n\nThere is a pattern in cases of enterprise technology adoption that is worth naming: organizations tend to measure too soon with the wrong criteria, condemning initiatives that should not yet be judged by the same metrics as the core business. But they also tend to leave initiatives without any metric at all for too long — initiatives that should already be producing value. With enterprise AI, many companies did the second: they deployed without measuring or architecting, and now they face the adjustment from a position of greater disorder and greater accumulated cost.\n\nThe correction is not expensive in absolute terms. A well-built data catalog, quality signals that function as routers, persistent memory for agents, clear rules for assigning models according to task complexity: none of these decisions requires an exceptional budget. They require something harder to obtain in organizations that are already in scale mode: deliberate priority and architectural discipline sustained over time.\n\n## The advantage that cannot be bought in the model market\n\nAgrawal frames token efficiency as the next competitive advantage in enterprise AI. The reading is correct but can be refined. The real advantage does not lie in token efficiency as an isolated metric. It lies in the organizational capacity to convert proprietary data into reliable context for agents that operate at scale, with sufficient governance so that results are auditable and attributable.\n\nThat is not a capability that can be purchased from a model provider or obtained by reducing the price per million tokens. It is a capability that is built internally, with data architecture decisions that precede agent deployment — not the other way around. Companies that already have that infrastructure gain an advantage that amplifies over time: each interaction is cheaper, faster, and more reliable than the one before. Those that do not have it face costs that do not fall simply because token prices drop.\n\nThe language model market will continue to get cheaper over time. That is almost certain. But the gap between companies that know how to use AI efficiently and those that do not will remain a problem of organizational design, data quality, and governance. And that gap has no solution in any external vendor's catalog.\n\nThe organizations that in 2026 are still operating with stateless agents, without functional data catalogs, and sending every workload to the most expensive available model, are not paying for tokens. They are paying the deferred price of not having designed their AI infrastructure when it was still cheap to do so.","article_map":{"title":"Enterprise AI Pipelines Don't Lose Money on Tokens: They Lose It Before","entities":[{"name":"Uber","type":"company","role_in_article":"Case example of a large enterprise that adjusted AI spending after token consumption exceeded planned budgets."},{"name":"Sumeet Agrawal","type":"person","role_in_article":"VP of Product Management at Salesforce whose Fortune column provides the central diagnostic framework for the article."},{"name":"Salesforce","type":"company","role_in_article":"Employer of the article's primary source; its Data 360 platform is cited as an architectural solution example."},{"name":"Informatica","type":"company","role_in_article":"Partner whose master data management integrates with Salesforce Data 360 to provide verified agent context."},{"name":"Anthropic","type":"company","role_in_article":"Cited as a vendor that has lowered token prices and added caching, addressing the wrong variable."},{"name":"OpenAI","type":"company","role_in_article":"Cited alongside Anthropic and Google as vendors responding to cost pressure with price reductions."},{"name":"Google","type":"company","role_in_article":"Cited as a vendor that has lowered token prices, addressing symptoms rather than structural causes."},{"name":"Cursor","type":"product","role_in_article":"Cited for incorporating cost as a variable in model selection in Composer 2.5, a partial architectural response."},{"name":"Enterprise AI pipeline","type":"technology","role_in_article":"Central subject of the article; described as a multi-stage system with five distinct leakage points."},{"name":"Data 360","type":"product","role_in_article":"Salesforce customer data platform cited as an example of infrastructure that enables auditable AI value."}],"tradeoffs":["Broad AI pilot exploration generates learning but defers the governance infrastructure needed for sustainable scaling","Lowering token prices reduces unit cost but does not address volume waste from poor architecture","Using frontier models for all tasks maximizes output quality ceiling but creates unnecessary cost for routine workloads","Deploying agents quickly without persistent memory accelerates rollout but compounds per-interaction costs over time","Investing in data governance before agent deployment delays time-to-deployment but reduces long-term token spend and improves auditability"],"key_claims":[{"claim":"Excess raw context in prompts can mean 5 to 10 times more tokens than necessary per interaction.","confidence":"high","support_type":"reported_fact"},{"claim":"At $10-$15 per million tokens with thousands of daily interactions, unoptimized pipelines generate hard-to-ignore cost overruns quickly.","confidence":"high","support_type":"reported_fact"},{"claim":"Uber adjusted internal AI spending after consumption exceeded planned budgets.","confidence":"high","support_type":"reported_fact"},{"claim":"Sending every task to the largest available model regardless of complexity is the primary direct-budget leakage point.","confidence":"high","support_type":"reported_fact"},{"claim":"Agents without persistent memory restart context from scratch on every interaction, paying repeatedly for already-resolved exceptions.","confidence":"high","support_type":"reported_fact"},{"claim":"The integration between Informatica master data management and Salesforce Data 360 converts ungoverned AI consumption into auditable business value.","confidence":"medium","support_type":"reported_fact"},{"claim":"Token efficiency is the next competitive advantage in enterprise AI.","confidence":"medium","support_type":"editorial_judgment"},{"claim":"Companies that fix architecture first gain an advantage that price reductions cannot replicate because it is rooted in proprietary data quality.","confidence":"medium","support_type":"inference"}],"main_thesis":"The root cause of enterprise AI budget overruns is not token price but the absence of data architecture, governance, and routing discipline before agents are deployed. Companies that fix their internal infrastructure gain a compounding advantage that no external vendor price reduction can replicate.","core_question":"Why do large enterprises overspend on AI despite falling token prices, and what structural changes would actually close the value gap?","core_tensions":["Speed of AI deployment vs. discipline of data infrastructure buildout","Vendor-side cost reduction vs. enterprise-side architectural efficiency","Broad exploration logic vs. premature exploitation without foundations","Token price as the perceived lever vs. data quality and governance as the actual lever","Short-term pilot success metrics vs. long-term attributable business value"],"open_questions":["At what organizational scale does the ROI of building a proper data catalog and governance layer become unambiguous?","How should enterprises sequence the five leakage point fixes given limited internal bandwidth?","Will model routing automation (selecting the right model per task) become a commodity feature or remain a differentiating capability?","How do SMEs without large data infrastructure teams replicate the architectural advantages available to large enterprises?","What governance standards will emerge to make AI consumption auditable at the board level?"],"training_value":{"recommended_for":["CTOs and CIOs evaluating enterprise AI infrastructure investment priorities","CFOs reviewing AI budget overruns and seeking structural rather than pricing solutions","Enterprise architects designing agent deployment pipelines","AI strategy consultants advising large organizations on scaling from pilots to production","Product managers responsible for AI cost optimization and ROI attribution"],"when_this_article_is_useful":["When evaluating whether to renegotiate AI vendor contracts vs. invest in internal data infrastructure","When diagnosing why AI token spend is growing faster than business value","When designing the governance layer for an enterprise AI deployment","When advising a company on transitioning from AI pilot phase to scalable production architecture","When building a board-level business case for data catalog and MDM investment"],"what_a_business_agent_can_learn":["How to diagnose AI budget overruns by mapping leakage points rather than negotiating token prices","The five structural causes of token waste in enterprise AI pipelines and their relative cost impact","Why model routing by task complexity is an organizational decision with direct budget consequences","How data governance functions as a cost-reduction mechanism, not just a compliance requirement","The difference between linear vendor-side improvements and order-of-magnitude internal architectural improvements","Why persistent agent memory and semantic caching are infrastructure investments with compounding returns"]},"argument_outline":[{"label":"1. The symptom","point":"Token consumption exceeded budgets at scale enterprises like Uber in 2026 without generating proportional value.","why_it_matters":"This is the visible, board-level signal that the maximalist AI adoption logic of 2024-2025 has hit a structural ceiling."},{"label":"2. The misdiagnosis","point":"Vendors responded by lowering token prices and adding prompt caching, which addresses cost per unit but not volume of waste.","why_it_matters":"A linear price reduction applied to a system consuming 5-10x more tokens than necessary produces proportional but insufficient savings."},{"label":"3. The five leakage points","point":"Raw context flooding prompts, ungoverned data access, uniform use of frontier models, stateless agents, and lack of semantic caching each independently inflate token consumption.","why_it_matters":"Each leakage point is independently fixable with architectural decisions, not vendor contracts."},{"label":"4. The real locus of advantage","point":"The competitive edge lies in converting proprietary data into reliable, auditable agent context, not in buying cheaper tokens.","why_it_matters":"This capability is built internally and compounds over time; it cannot be purchased from a model provider."},{"label":"5. The organizational diagnosis","point":"Companies skipped disciplined exploration infrastructure and moved prematurely to exploitation, accumulating technical debt instead of governance.","why_it_matters":"The correction requires deliberate priority and architectural discipline, not exceptional budget."}],"one_line_summary":"Enterprise AI cost overruns are not a token pricing problem but an architectural and organizational design failure that precedes token consumption.","related_articles":[{"reason":"Directly parallel analysis of hidden cost structures in enterprise AI agent deployment, framing token and operational costs as an unbudgeted tax on corporate AI.","article_id":14501},{"reason":"Quantifies the enterprise AI value gap with Bain survey data showing 40% of large companies measure no ROI, providing empirical grounding for the article's central claim.","article_id":14401},{"reason":"Examines why AI pilots succeed but never scale, which is the organizational failure mode that precedes the token overspend problem described here.","article_id":14521},{"reason":"Analyzes agent gateways as an emerging control layer in enterprise AI, directly relevant to the routing and governance architecture the article prescribes.","article_id":14481},{"reason":"Databricks' $188B valuation is driven by enterprise data infrastructure demand, contextualizing why data architecture investment is becoming a strategic priority.","article_id":14601},{"reason":"Examines how AI contract structures misalign with value creation, complementing the article's argument that the problem is organizational and contractual, not technological.","article_id":14381}],"business_patterns":["Maximalist technology adoption followed by budget correction when consumption outpaces value generation","Vendor price reductions as a response to customer pressure that addresses symptoms rather than structural causes","Technical debt accumulation when exploitation phase begins before exploration infrastructure is complete","Compounding advantage for early architectural movers as each interaction becomes cheaper and more reliable over time","Governance treated as after-the-fact control rather than as a routing signal embedded in the workflow"],"business_decisions":["Whether to prioritize token price negotiation with vendors versus internal data architecture investment","How to assign AI tasks to models of appropriate capability rather than defaulting to frontier models for all workloads","Whether to build persistent agent memory to avoid reprocessing context on every interaction","How to implement a data catalog with quality signals that route agents to certified data on first attempt","When to transition from broad AI pilot exploration to deliberate architectural consolidation","How to measure AI initiative ROI with metrics appropriate to the maturity stage of each deployment"]}}