A system for evolving complete AI agents rather than repeatedly adjusting one prompt. It runs populations through epochs, evaluates evidence against a benchmark, and preserves, mutates, or replaces prompts, code, tools, and dependencies while keeping scoring and isolation fixed.