Sandboxed harness for testing coding-agent learning

Arc Code tests whether a general coding agent can learn unfamiliar ARC games by creating its own parsers, simulators, and search programs during play.

From Github Awesome โ€” GitHub Trending Weekly #45: comet, blobatar, cumora, herdr, barehands, ip-as-logo-skill, Aura, md2hd at 06:27

Problem: Agent evaluations need to test tool creation and adaptation rather than only predefined task execution.

For: Researchers evaluating general coding-agent capabilities.

Products from this video

ai-data-extractor AI Design Skills APEX Inference Chip arc-code Aura barehands blobatar Core Framework Cumora DesktopFly dgit Endoplexity herdr HQBase icm-architect IP as Logo Skill jit macOS Harness md2hd morphnext NorthCinder Procedural Sounds sloptrim TrueForge unlazy Vibe ASO VibePulse Viscose Zeron

Examples

๐Ÿ”’ Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good โ€” the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis โ€” free

Related ideas