The most powerful model is not the one that wins in business
There is a number that the president of Alibaba.com chose to open his argument, and it is an uncomfortable number: 61.7%. That is the percentage of real-world commercial tasks successfully completed by the best artificial intelligence agent they tested. Across 107 real e-commerce operations, the most capable system available failed in four out of ten.
Kuo Zhang did not choose it to apologize. He chose it to build a case about what comes after the race for the most intelligent model. And that case, read through the lens of adoption psychology rather than technical architecture, says something more unsettling than any benchmark: companies that continue to bet on raw power as a differentiating advantage are misreading where the real battle is won.
The story published in Fast Company introduces CommerceAgentBench, the open-source test suite that Alibaba.com built from real operations. The methodology differs from conventional language benchmarks: it does not measure whether the model generates a plausible response, but whether the task was completed. The purchase order was dispatched. The listing was published correctly. The shipment was booked according to specifications. That binary success criterion makes the benchmark far closer to how a company measures human work, and far more distant from how the artificial intelligence industry has measured progress over the past few years.
What they found when running different models across those 107 tasks was not a clear hierarchy. It was fragmentation. The model that led on quotations and market research fell behind on claims resolution and listing compliance. A third system proved most efficient at product publishing and returns management. None dominated everything. That dispersion is not a technical accident: it is the central point of the argument.
The illusion of the single model and the cognitive cost nobody quotes
The way most companies adopted artificial intelligence over the past two years followed a recognizable pattern: choose the most capable model available, connect it to existing workflows, and hope that raw power would compensate for the lack of design. It is understandable. It reduces the complexity of the decision. It avoids the discomfort of having to deeply understand what each model does in each context. And it allows the team adopting the technology to feel they made the right decision because they made the most expensive and recognized one.
This pattern has a name in behavioral economics: attribute substitution heuristic. When a question is hard — which is the best system for each of my twenty operational workflows — the mind replaces it with an easier question to answer: which is the best available model. The answer to the second question is applied to the first without anyone noticing the swap.
The cost of that swap is invisible until it is measured. Alibaba.com measured it across 44 tasks and found that its Accio agent, which distributes work among models according to the difficulty and type of reasoning required, cost $1.72 in tokens. Codex cost $3.79. Claude Code cost $3.91. The difference does not come from using cheaper models in every case: it comes from not using the most expensive ones when they are not needed.
That figure is not a return-on-investment promise. It is a clue to the mechanics operating underneath. When a company standardizes on the most powerful model because that simplifies the decision, it is paying for high-complexity reasoning on tasks that do not require it. It is, in operational terms, assigning its best analyst to filing documents because that avoids deciding who should file and who should analyze.
Zhang's proposal is called precision delegation: assigning to each workflow the level of capability and authority that that specific workflow actually requires. It is not a new concept in organizational management. What is new is that it now applies to automated systems, and that the penalty for not doing it has a direct expression in the variable cost of operating artificial intelligence at scale.
Why adoption stalls before the technical problem
There is a dimension that the Fast Company article mentions obliquely but that deserves closer examination: the relationship between the level of autonomy granted to an agent and the operational consequences of its errors.
Zhang frames it with a direct distinction. A weak draft of a social media post carries one type of consequence. An incorrect payment to a supplier carries another. That difference is not merely technical: it is the core of the psychological adoption problem in enterprise environments.
Organizations that have incorporated artificial intelligence tools into low-consequence workflows — text generation, summaries, internal drafts — have done so with relatively low friction. The error is visible, correctable, and produces no immediate financial harm. The person using the tool retains control over the final output. Their professional identity is not at risk because the system still passes through their judgment before producing any effect in the world.
The moment that logic breaks down is when the agent connects directly with systems that execute actions: payments, inventory, logistics, purchase orders. There the equation changes in a fundamental way. The error is no longer a draft that gets discarded: it is an incorrect shipment, a duplicated payment, a misconfigured purchase order that generates contractual obligations. And the person who authorized the automation bears responsibility for that error even though they did not intervene in it.
That shift of responsibility without a corresponding shift of control is one of the most underestimated psychological blockers in the adoption of operational artificial intelligence. Companies do not stall because their teams fail to understand the technology. They stall because their teams understand perfectly well that when an agent fails on a high-impact task, the error has a human name and surname. And that name is usually the one who approved the automation.
The 61.7% completion rate of the best model on real commerce tasks is, from this perspective, a number that operates in two opposite directions. For someone who reads it as an adoption argument, it says there is already enough capability to automate with supervision. For someone who reads it through the fear of responsibility, it says that nearly four out of ten tasks fail, and that those failures happen inside systems connected to money, inventory, and suppliers.
Neither reading is incorrect. They are simply the two sides of the same number, and the company deciding how much authority to delegate to an agent is, at bottom, making a decision about how much error risk it is willing to push downward through its organizational structure before having built solid performance evidence.
The Harrison Nott case and what it reveals about the limits of the argument
The article closes with the example of Harrison Nott, a 16-year-old who used Alibaba.com and Accio to develop a cooling-products business that, according to Fast Company's reporting, has generated $1 million. It is a well-chosen story from a narrative standpoint: it illustrates that capabilities once reserved for teams with budgets and specialists are now within reach of a single individual with no corporate structure.
But the example also reveals, without intending to, an important asymmetry in the precision-delegation argument.
For a solo entrepreneur operating with complete autonomy over their decisions, the question of how much authority to delegate to an agent is almost rhetorical. If the agent makes a mistake, the cost is absorbed by the same person who made the decision. There is no organizational hierarchy that must be held accountable. There are no legacy systems that interfere. There are no compliance teams asking which approval protocol was followed.
In that context, adoption friction is minimal. The utility is immediate. The precision-delegation model works because the person using it has both the incentive to adopt and the flexibility to absorb the error.
The problem is that this context does not describe most of the organizations where the promise of precision delegation would be most valuable. A mid-sized company with ERP systems, approval processes, finance departments, and internal auditors faces a completely different friction architecture. There, adopting a system that distributes work among multiple models according to criteria of difficulty and cost requires not only a technical decision, but aligning governance criteria, defining exception protocols, and building a layer of evidence that justifies expanding the agent's authority incrementally.
That layer is neither expensive nor technically difficult. But it is invisible in the argument as presented, and it is precisely what determines whether precision delegation moves from concept to practice inside a structure with multiple interests and distributed responsibilities.
The advantage lies not in the model but in accumulated trust evidence
Kuo Zhang's argument works best when read as a maturity map rather than an immediate adoption proposal. The companies that will gain an advantage in the next phase of operational artificial intelligence will not necessarily be those with access to the most capable models. They will be those that have built the evaluation infrastructure allowing them to know, workflow by workflow, which model performs best, at what cost, and under what conditions it is safe to expand the system's autonomy.
That infrastructure is not glamorous. It does not appear in product announcements or public benchmarks. But it is the difference between a company that adopted artificial intelligence as an experiment and one that incorporated it as a sustainable operational capability.
What CommerceAgentBench proposes, in its most reduced form, is that continuous performance evaluation at the task level is itself a strategic capability. Not the model you use today, but the ability to measure its real performance, adjust the assignment, and expand the agent's authority as it accumulates evidence of reliability.
That process of gradually expanding authority based on demonstrated performance has a direct parallel in how successful organizations manage people. A new employee does not receive immediate access to all systems or authority to commit resources without approval. That authority expands as their performance generates evidence that they can exercise it without producing costly errors.
Applying that same logic to artificial intelligence agents is not a conservative restriction on the technology's potential. It is the only way to resolve the central psychological blockage that prevents organizations from moving adoption from low-consequence workflows into operations where the financial impact justifies the investment.
The advantage in the next cycle does not belong to whoever has the most intelligent model. It belongs to whoever has built the process for knowing when to trust the agent with a task that matters, backed by performance data that someone can show when the error finally occurs — because it always occurs — and the difference between a company that absorbs that error and one that is paralyzed by it lies in whether it had a protocol or merely had hope.










