Back to a16z Podcast

Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan

a16z Podcast

Full Title

Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan

Summary

This episode discusses the critical need for independent and evolving evaluation methods for AI models, as current public benchmarks can be misleading and become obsolete quickly.

The conversation highlights how robust evaluation is becoming essential for both AI labs to justify investment and for enterprises to understand ROI and effectively adopt AI technologies.

Key Points

  • Public benchmarks can provide a distorted view of AI model capabilities, as demonstrated by a model underperforming on private benchmarks while excelling on public ones.
  • Independent evaluation is crucial to provide a more accurate and unbiased assessment of AI models, especially as the industry matures and trillion-dollar markets emerge, mirroring historical needs for testing groups in new industries.
  • AI labs need reliable methods to prove their models' advancements to justify billions in investment, moving beyond self-reported metrics to substantive evidence.
  • The VALS team is developing sophisticated evaluation systems that can operate within tight pre-release windows (hours), leveraging automation and advanced infrastructure.
  • Evaluating AI models faces challenges similar to evaluating human intelligence, where clear, universally agreed-upon frameworks are difficult to establish, and models can be "hacked" to perform well on specific tests.
  • VALS aims to make enterprise workflows and human expertise more explicit and legible to enable better AI model evaluation, addressing the difficulty of testing AI in real-world contexts.
  • The rapid evolution of AI necessitates that benchmarks evolve alongside models; VALS actively deprecates saturated benchmarks and constructs new ones to represent emerging frontiers and the current state of the world.
  • The increasing complexity of AI tasks, including agentic work and multi-day operations, requires more stable and sophisticated evaluation infrastructure that can handle long-running processes and retries.
  • For enterprises, understanding the true ROI of AI is becoming existential, as token costs can eclipse salary expenses, making accurate evaluations critical for strategic adoption and competitive advantage.
  • The development of tools like Valsmith aims to help companies build internal coding benchmarks, identifying cost-effective and high-ROI AI agents for specific tasks.
  • Policy discussions around AI are often abstract, lacking empirical grounding; VALS aims to provide evidence-based data on model capabilities and risks to inform more sophisticated policy conversations.
  • A collaborative approach involving labs, independent evaluators, customers, and government is needed to set AI standards, with government setting and enforcing rules, and private companies providing the detailed, ongoing evaluation.
  • Global AI development, including the rise of "sovereign AI," necessitates a shared language of evaluations to facilitate trust, verification, and alignment on risks, drawing parallels to nuclear arms control.
  • Evaluating AI for cybersecurity, biosecurity, and the potential for recursive self-improvement requires continuous development of new benchmarks that reflect evolving capabilities and risks.

Conclusion

As AI models rapidly advance, the need for independent, dynamic, and sophisticated evaluation methods is paramount for both industry and policy.

Robust and transparent evaluation frameworks are essential for enterprises to understand the value and return on investment of AI technologies, and for policymakers to create effective regulations.

The future of AI safety and responsible development hinges on establishing a shared language of evaluation that can keep pace with innovation and facilitate global cooperation.

Discussion Topics

  • How can we ensure AI evaluation methods evolve as rapidly as AI capabilities themselves?
  • What role should independent third parties play in setting standards and auditing AI models for safety and efficacy?
  • As AI costs potentially rival or exceed human salaries, how can businesses ensure they are investing wisely and maximizing the ROI of AI adoption?

Key Terms

Benchmark
A standard or point of reference against which things may be compared or assessed.
Frontier model
A state-of-the-art AI model that represents the cutting edge of current capabilities.
Recursive self-improvement (RSI)
The ability of an AI system to improve its own intelligence or capabilities, potentially leading to rapid advancements.
Agentic work
Tasks performed by AI systems that can act autonomously, plan, and execute actions to achieve goals.
Token
In the context of AI models, a unit of text or data that the model processes.
ROI
Return on Investment, a measure of profitability or efficiency.
Sovereign AI
The development and deployment of AI technologies within national borders, often with the goal of maintaining technological independence and control.

Timeline

00:02:45

It was clearer to the VALS founder that public benchmarks were insufficient and a new methodology was needed.

00:03:36

A key example of public benchmark inadequacy occurred with Meta's Llama 4 release.

00:04:22

The need for independent testing groups when new trillion-dollar industries emerge is a historical parallel to current AI evaluation needs.

00:04:56

VALS has developed processes to conduct evaluations within a six-hour pre-release window for models.

00:06:05

Evaluating AI models faces challenges similar to human intelligence testing, with models potentially "hacking" benchmarks.

00:06:45

The difficulty of making fuzzy or distributed human evaluation explicit is a bottleneck for testing AI models.

00:07:39

Analogies like the MPAA's rating system highlight the evolving and subjective nature of defining and evaluating content, which is similar to AI.

00:08:46

VALS's decision to not sell training data to labs is a proactive measure against creating conflicts of interest.

00:09:35

VALS offers a catalog of benchmarks, including popular ones for finance and coding, and experimental ones like the recursive self-improvement index.

00:10:45

The recursive self-improvement benchmark attempts to proxy the expensive and slow process of a frontier model training its successor.

00:11:23

VALS deprecates benchmarks that become saturated and constructs new ones to continually challenge AI models.

00:12:26

Benchmarks need to reflect the current state of the world, akin to how professionals like lawyers or doctors must be re-certified.

00:13:20

The rise of agentic work has changed how AI capabilities are evaluated, requiring a focus on more complex, multi-step tasks.

00:13:38

Beyond capability, evaluations now consider parameters like cost, latency, and flexibility for broader domain application.

00:14:04

Evaluating AI models that run for hours or days requires stable infrastructure with retry mechanisms.

00:14:17

Complex AI evaluations now involve fewer samples but a larger set of criteria and expectations for outputs.

00:15:13

While platforms like OpenRouter act as gateways, the core challenge remains building evaluations to determine optimal model routing for specific applications.

00:15:57

Enterprise adoption of AI is existential, as AI costs can become a significant line item, necessitating clear ROI justification through evaluations.

00:18:47

Valsmith is a new product designed to help companies build internal coding benchmarks from their GitHub repositories.

00:20:39

The "messy middle" of AI models presents non-intuitive choices for enterprises, where cost and capability are not always directly proportional.

00:21:39

The existing repository of human work can be leveraged to build dynamic AI evaluations across industries.

00:22:38

VALS's internal use of Valsmith revealed surprising insights, such as the token efficiency of certain coding tools, informing their strategy.

00:24:43

Policy discussions around AI need empirical grounding; VALS aims to provide data on model capabilities and risks to inform this.

00:25:27

The government's role in AI policy includes identifying fears and then using evaluators to determine if models can perform those actions.

00:26:53

Balancing rapid innovation with ensuring AI is in the best interest of people requires careful reconciliation of competing forces.

00:28:14

VALS broadly categorizes AI alignment issues, including reward hacking and ensuring models align with user intent.

00:29:17

Government agencies can set rules based on their fears, but private companies are better suited to test if models can indeed perform those actions.

00:30:32

The trend of large labs keeping advanced models internal highlights the challenges of controlling potentially harmful capabilities.

00:30:55

Bridging the gap between the narrative of AI capabilities and real-world deployment in bespoke enterprise environments is crucial.

00:32:03

Insights and data from AI evaluators should be passed to government branches to inform policy development.

00:33:08

The rise of "sovereign AI" and geopolitical competition necessitates a shared language of evaluations for verification and risk alignment.

00:35:52

Harmonizing global AI policy is a complex challenge, but starting with mutual interests like cybersecurity and biosecurity is a step.

00:36:21

The potential for unchecked recursive self-improvement in AI poses a significant long-term risk that requires joint international conversations.

00:37:24

VALS is focused on building benchmarks that capture the frontier of AI capabilities and risks, including infrastructure-level concerns.

Episode Details

Podcast
a16z Podcast
Episode
Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan
Published
September 9, 2026