
a16z general partner Anjney Midha sits down with LMArena cofounders Anastasios N. Angelopoulos, Wei-Lin Chiang, and Ion Stoica to talk about the future of AI evaluation. As benchmarks struggle to keep up with the pace of real-world deployment, LMArena is reframing the problem: what if the best way to test AI models is to put them in front of millions of users and let them vote? The team discusses how Arena evolved from a research side project into a key part of the AI stack, why fresh and subjective data is crucial for reliability, and what it means to build a CI/CD pipeline for large models. They also explore: - Why expert-only benchmarks are no longer enough - How user preferences reveal model capabilities — and their limits - What it takes to build personalized leaderboards and evaluation SDKs - And why real-time testing is foundational for mission-critical AI Chapters: 00:00:04 - LLM evaluation: From consumer chatbots to mission-critical systems 00:06:04 - Style and substance: Crowdsourcing expertise 00:18:51 - Building immunity to overfitting and gaming the system 00:29:49 - The roots of LMArena 00:41:29 - Proving the value of academic AI research 00:48:28 - Scaling LMArena and starting a company 00:59:59 - Benchmarks, evaluations, and the value of ranking LLMs 01:12:13 - The challenges of measuring AI reliability 01:17:57 - Expanding beyond binary rankings as models evolve 01:28:07 - A leaderboard for each prompt 01:31:28 - The LMArena roadmap 01:34:29 - The importance of open source and openness 01:43:10 - Adapting to agents (and other AI evolutions)

Building an AI Physicist: ChatGPT Co-Creator’s Next Venture

From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki

Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building

The Current Reality of American AI Policy: From ‘Pause AI’ to ‘Build’

Rick Rubin: Vibe Coding is the Punk Rock of Software