Sandboxed harness for testing coding-agent learning
Arc Code tests whether a general coding agent can learn unfamiliar ARC games by creating its own parsers, simulators, and search programs during play.
From Github Awesome โ GitHub Trending Weekly #45: comet, blobatar, cumora, herdr, barehands, ip-as-logo-skill, Aura, md2hd at 06:27
Problem: Agent evaluations need to test tool creation and adaptation rather than only predefined task execution.
For: Researchers evaluating general coding-agent capabilities.
Products from this video
ai-data-extractor AI Design Skills APEX Inference Chip arc-code Aura barehands blobatar Core Framework Cumora DesktopFly dgit Endoplexity herdr HQBase icm-architect IP as Logo Skill jit macOS Harness md2hd morphnext NorthCinder Procedural Sounds sloptrim TrueForge unlazy Vibe ASO VibePulse Viscose Zeron
Examples
- ARC games: the tasks used to test whether a general coding agent can learn unfamiliar environments.