AI product Open source
livenerf is an open-source benchmark for detecting post-launch capability changes in frontier models. It calibrates a panel of intermittently solved questions from GPQA Diamond, MMLU-Pro, competition mathematics and AIME, then runs the same panel repeatedly through a hermetic, pinned Claude Code harness and compares paired per-item scores with a launch-period baseline. The benchmark records raw prompts and responses, usage, latency, CLI and harness versions, and token counts; it also runs a control model and applies preregistered statistical thresholds rather than relying on an LLM judge. It is built on the Inspect evaluation framework, supports scheduled daily runs on Linux, macOS and Windows, and is distributed under the MIT license. The repository describes it as an independent project, not affiliated with Anthropic.
1 use taken from transcripts — each links to the moment in the video.
Tests whether an AI model's performance declines after launch by running fixed questions daily, saving responses, and comparing scores with an early baseline.
1 in the library.