AI product Open source

CommerceAgentBench

CommerceAgentBench is a benchmark developed and maintained by the Accio team at Alibaba International for evaluating whether AI agents can complete long-horizon business workflows in high-fidelity, stateful, reproducible replicas of online services. Its tasks cover browser operations, native-style command-line tools, file and spreadsheet production, API and MCP workflows, public-web research, supplier analysis, product publishing, logistics, storefront configuration, and other commerce operations.

View repository Visit site Mentioned in 1 video ↓

Overview

Each task runs in a fresh container and is graded by a deterministic or LLM-assisted verifier. The benchmark uses local mock services representing commerce, SaaS, messaging, document, and operational systems; runs preserve the resolved configuration, agent trajectory, verifier result, artifacts, logs, and container metadata for auditing and reproducibility. The repository describes 107 tasks spanning CLI, browser, file, and API/MCP interfaces, with text-only, browser-text-capable, and vision-required capability slices. The project was previously called RealReplicaBench.

What CommerceAgentBench is used for

1 use taken from transcripts — each links to the moment in the video.

  • Benchmarks whether AI agents can complete long business workflows inside reproducible replicas of online services, with browser, command-line, file, API, and MCP tasks plus inspectable artifacts and verifier results.

Videos mentioning CommerceAgentBench

1 in the library.