Cronkite

Bulletin of September 9, 2026

5 minBusiness

Alibaba President: AI agents can talk, but can they actually do the work?

's president argues that the industry's focus on model intelligence misses the point. Real commercial work requires agents that can execute tasks correctly, not just converse. A new open-source benchmark, CommerceAgentBench, grades outcomes and reveals that even the strongest frontier models fail nearly 40% of real-world e-commerce tasks.

For all the excitement around artificial intelligence, the question most businesses should be asking is not which model is the smartest, but which one can reliably finish the job. That is the argument from Kuo Zhang, president of, who says the industry's obsession with reasoning benchmarks has obscured a more practical problem: AI agents can talk, but can they actually do the work?

Zhang points to the example of Joshua Stancle, who runs Clean Saint, a waterless oral-care film company in Los Angeles. Stancle uses AI for sourcing, marketing, web development, and customer support, describing the experience as having “ten of me.” But Zhang notes that this raises a critical question: if there are ten of you, how do you know the other nine are getting the work right?

The distinction, he argues, is between a model that reasons and an agent that acts. A model can process information, but commerce requires calling a supplier, filing a customs form, chasing a delayed container, and handling a return. None of that happens by itself. Work gets done by an agent, which combines the model with tools, memory, and industry-specific context that defines what a good outcome looks like.

To measure whether agents can actually execute, the Accio team at developed CommerceAgentBench, an open-source benchmark available on GitHub. The test contains 107 end-to-end tasks drawn from real e-commerce operations across procurement, logistics, product listing, fulfillment, and after-sales service. The tasks were assembled from 's own data, which includes 10 million active small-business users, 1.6 million real conversations, and 200,000 execution traces, sorted into seven categories of commercial work.

Grading is based on the end state, not on conversational quality. A product listing either went live with the correct attributes or it did not. Freight either moves on a route that exists or it sits on a dock in Ningbo while someone figures out what went wrong. Zhang argues that commerce has always kept score this way: a customer who receives the wrong product has no interest in how articulate the agent sounded when it placed the order.

The results are sobering. The strongest frontier model tested successfully completed only 61.7% of the tasks. That is high enough to be useful, but low enough to be a warning. Nearly four in ten tasks came back wrong. Failures clustered in recognizable areas: agents struggled to spot payment anomalies hidden inside long supplier email threads, had trouble calculating landed costs with multiple moving variables, broke down on after-sales disputes requiring reconciliation of conflicting documents, and consistently struggled with multi-leg shipping routes.

Zhang warns that the risk changes as adoption scales. Across thousands of businesses using similar agents, individual mistakes could become correlated ones. Inaccurate listings could multiply, fraud signals could be missed, and routing or compliance errors could ripple through supply chains. Measurement, he says, shows where automation is ready to scale and where human oversight still needs to keep pace.

One finding surprised Zhang more than the headline number: no single model won. Leadership rotated by category. The model that ranked first on request-for-quote work and market research slipped behind on claims settlement and listing compliance, where a different model led. A third was strongest at publishing products and handling returns. A ranking built from general reasoning scores, he argues, tells you very little about which system will perform on a particular commercial task.

For an individual business, this translates into a practical question: which workflows can be handed over now, and which still need a person? Zhang calls this “precision delegation.” Where pass rates are high, supplier comparison and routine listing work can come off a desk. Where they are low, on unusual compliance questions, complicated negotiations, and the exceptions that make up more of any commerce operation than anyone expects, a person should stay in the loop and check the work.

Once a business knows where an agent is dependable, it can stop supervising it. Once it knows where the agent breaks, it can catch the failure before a customer does. Both save money, Zhang says, but neither is available without measurement.

Evan Emerson

Author

Political Correspondent

Evan Emerson covers public affairs, politics, business, culture and daily news for Cronkite. The role focuses on verification, context, and clear explanations for readers.

Tail slate

Reporter
Evan Emerson
Filed
Runs
5 min
Source
Fortune | FORTUNE
Block
Business

Next in the Business block

  1. ——:—— Sep 9 TRM Labs valuation doubles to $2 billion as crypto crime fighter scales AI services 4 min
  2. ——:—— Sep 9 Americans’ faith in capitalism hits 15-year low, but the problem is know-how, not ideology 5 min
  3. ——:—— Sep 8 Paramount Skydance Again Demands $1.88 Billion Bond From States and WGA in Merger Lawsuit 3 min
  4. ——:—— Sep 8 OpenAI CFO says enterprise business has been on a tear 4 min