Independent, continuously evolving AI model evaluation company

I build a third-party evaluation business that measures frontier-model capabilities independently of the labs developing them. The business creates held-out, private, high-signal benchmarks; evaluates models before release; updates or retires benchmarks as models saturate; and serves both model labs seeking evidence of progress and enterprises choosing models for specific work.

From a16zInside the Race to Measure Frontier Intelligence at 00:55

Problem: Public benchmarks can become saturated, open questions and rubrics can be optimized against, and self-reported scores can diverge from performance on private real-world tasks. Enterprises also lack a clear way to determine which model or agent is the best fit for their own work and spending.

For: Frontier model labs that need independent evidence of capability improvements, and enterprises that need to select models, agents, routers, and usage policies based on productivity, cost, latency, risk, and ROI.

Products from this video

Claude Code Devin vals Vals VI codebench

Examples

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free