CommerceAgentBench is a benchmark developed and maintained by the Accio team at Alibaba International for evaluating whether AI agents can complete long-horizon business workflows in high-fidelity, stateful, reproducible replicas of online services. Its tasks cover browser operations, native-style command-line tools, file and spreadsheet production, API and MCP workflows, public-web research, supplier analysis, product publishing, logistics, storefront configuration, and other commerce operations.
Each task runs in a fresh container and is graded by a deterministic or LLM-assisted verifier. The benchmark uses local mock services representing commerce, SaaS, messaging, document, and operational systems; runs preserve the resolved configuration, agent trajectory, verifier result, artifacts, logs, and container metadata for auditing and reproducibility. The repository describes 107 tasks spanning CLI, browser, file, and API/MCP interfaces, with text-only, browser-text-capable, and vision-required capability slices. The project was previously called RealReplicaBench.