AI-agent workflow benchmark using reproducible service replicas
Evaluate whether AI agents can complete long business workflows inside reproducible replicas of online services, with isolated runs and inspectable artifacts.
From Github Awesome — GitHub Trending Weekly #43: Cloudflare Computer, waku-agent, querysplat, findphone, Soup, morphicons at 12:31
Problem: Agent evaluations need controlled environments, repeatable initial state, and evidence of what happened during each attempt.
For: AI-agent developers and evaluators who need repeatable tests of browser, command-line, file, API, and MCP workflows.
Examples
- RealReplicaBench contains 107 tasks covering browser work, command-line tools, files, and API or MCP operations, including product publishing and freight booking.
Soon you can unlock the full business plan.
Behind this: 7 build steps · 3 tools and how each is used · how to validate demand · 4 things the video never answers.