
Sep 9, 2026 · 40 min
Independent evaluations race to keep pace with AI models
Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan
As models move into agents, enterprise workflows, and policy decisions, outdated or self-reported benchmarks can obscure both capability and risk.
- 1Independent evaluators must update benchmarks as models gain new capabilities and learn to optimize against existing tests.
- 2Enterprise model selection increasingly depends on task-specific evidence, pricing, tooling, and workflow performance rather than general rankings.
- 3Evaluations could give policymakers concrete evidence on alignment, cybersecurity, reward hacking, and geopolitical differences in model behavior.
Don't miss
Ryan Krishnan describes how VALS spent roughly $1.5 million in token value during one month using coding tools, leading to VALSmith’s cost-aware routing system.
The brief
Ben Horowitz and Ryan Krishnan open with a basic problem: public benchmarks can overstate what models do in practice, while labs’ self-reported results leave too much unexamined.
Krishnan explains VALS’s six-hour pre-release evaluations and its effort to turn fuzzy real-world abilities into explicit tests without slowing model launches or creating conflicts of interest.
The frontier is moving beyond question answering. VALS tests agents across long-running workflows and explores recursive improvement, where models help build improved versions of themselves.
For enterprises, the best general-purpose model may not fit a particular repository or task. VALSmith emerged after roughly $1.5 million in monthly token value, routing work to tools of different capability and cost.
The discussion ends with a policy question: can independent evaluations translate fast-changing model behavior into evidence that governments and international actors can use before standards fall behind?