A benchmark for evaluating coding agents on realistic tasks that require the judgment and work quality of senior software engineers. It uses expert-designed instructions, validation specifications, deterministic tests, and a metric called a “tasteful pass.”