Independent, continuously evolving AI model evaluation company
I build a third-party evaluation business that measures frontier-model capabilities independently of the labs developing them. The business creates held-out, private, high-signal benchmarks; evaluates models before release; updates or retires benchmarks as models saturate; and serves both model labs seeking evidence of progress and enterprises choosing models for specific work.
From a16z — Inside the Race to Measure Frontier Intelligence at 00:55
Problem: Public benchmarks can become saturated, open questions and rubrics can be optimized against, and self-reported scores can diverge from performance on private real-world tasks. Enterprises also lack a clear way to determine which model or agent is the best fit for their own work and spending.
For: Frontier model labs that need independent evidence of capability improvements, and enterprises that need to select models, agents, routers, and usage policies based on productivity, cost, latency, risk, and ROI.
Products from this video
Examples
- Meta's Llama 4: it showed incredible capability on major public benchmarks but underperformed on Vals's held-out private benchmarks.