AI-agent workflow benchmark using reproducible service replicas

Evaluate whether AI agents can complete long business workflows inside reproducible replicas of online services, with isolated runs and inspectable artifacts.

From Github Awesome — GitHub Trending Weekly #43: Cloudflare Computer, waku-agent, querysplat, findphone, Soup, morphicons at 12:31

Problem: Agent evaluations need controlled environments, repeatable initial state, and evidence of what happened during each attempt.

For: AI-agent developers and evaluators who need repeatable tests of browser, command-line, file, API, and MCP workflows.

Products from this video

anydoc Backchannel Bindwidth Can I Vibecode It? Capptivo cargo-frisk chapter-tgz Claude Code Claude Desktop Cloudflare Computer ComfyUI-Spectrum-MiniMax-H3 CommerceAgentBench creakwork12 da-cli Desko diri doc7 Falco findphone Finger Frame AI GenOffice Gmail Jungle Trail King's Gambit — Medieval 3D Chess MAGI-2 Preview Microsoft Outlook morphicons Node.js OpenAI Codex OpenEdit Outreachr Python QuerySplat Rust Shitty Soup SQLite Stickman Video Director Swiftlet Token Saver Virtual Mac on iPad vPhone Workstation Waku Agent whisper.cpp

Examples

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Related ideas