An independent evaluation company for foundation models and agents
Create standardized evaluations that measure models and agent systems consistently, prevent benchmark cheating, and allow competing foundation-model companies to be compared under the same execution conditions.
From Latent Space — The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI at 01:07:36
Problem: Self-reported benchmarks can be difficult to compare because companies may run evaluations differently, use different infrastructure or tools, or inadvertently optimize for the benchmark rather than real-world capability.
For: Foundation-model companies, model users, investors, and developers who need trustworthy comparisons of models and agent systems.
Products from this video
Amazon Web Services (AWS) Apache Spark Artificial Analysis Blender Kilo Code Model Factory OpenCode
Examples
- Artificial Analysis: Poolside identified it as a third party that could run independent evaluations of models, including checking the relevant infrastructure and tools.