
Alexander Novikov
100% confidenceNgân Vũ
100% confidenceMarvin Eisenberger
100% confidenceEmilien Dupont
100% confidencePo-Sen Huang
100% confidenceAdam Zsolt Wagner
100% confidenceSergey Shirobokov
100% confidenceBorislav Kozlovskii
100% confidenceFrancisco J. R. Ruiz
100% confidenceAbbas Mehrabian
100% confidenceM. Pawan Kumar
100% confidenceAbigail See
100% confidenceSwarat Chaudhuri
100% confidenceGeorge Holland
100% confidenceAlex Davies
100% confidenceSebastian Nowozin
100% confidencePushmeet Kohli
100% confidenceMatej Balog
100% confidenceGoogle DeepMind just dropped AlphaEvolve, a Gemini-powered evolutionary coding agent that designs advanced algorithms by pairing LLM creativity with automated evaluation. The headline result: it beat Volker Strassen's 56-year-old record for 4x4 matrix multiplication, finding a method that uses 48 scalar multiplications instead of 49. No human or AI had managed that in over half a century. Tim sits down with two of the researchers behind the work -- Matej Balog and Alexander Novikov -- to walk through the system, the results, and what it means. In this episode: - How AlphaEvolve works: an evolutionary pipeline that pairs LLM-generated code proposals with rigorous automated evaluators, iteratively improving solutions rather than relying on one-shot generation. The gap between single-shot LLM sampling and scaled evolutionary search turns out to be enormous. - The matrix multiplication breakthrough: AlphaTensor tried for years with reinforcement learning and only cracked the Boolean case. AlphaEvolve found a 48-multiplication algorithm for general 4x4 matrices almost by accident, running "for completeness." The result generalises from complex to real matrices, which is counterintuitive -- solving the harder problem actually made the search easier. - Three ways to represent the search target: direct solution, constructor function, or search algorithm. For matrix multiplication, AlphaEvolve designed a gradient-based search algorithm that finds matrix multiplication algorithms -- a meta-level approach that produced loss functions and update rules no human would have tried. - Real-world impact at Google scale: AlphaEvolve recovered 0.7% of fleet-wide compute resources in the Borg data center scheduling system and sped up Gemini training by 1%. These are already-heavily-optimized production systems. - The human-AI collaboration loop: AlphaEvolve is not autonomous research. Humans choose the problems, design evaluators, seed initial solutions, and interpret results. Alexander Novikov argues this back-and-forth is the whole point -- the system improves your questions as much as your answers. - Keith Duggar probes the halting problem and evaluation cascade limitations. The researchers acknowledge the constraint but note that practical framing (time-bounded evaluation, evaluation cascades from cheap to expensive) sidesteps the theoretical issue for now. - The recursive self-improvement question: AlphaEvolve improved the infrastructure that trains the models that power AlphaEvolve. The feedback loop exists but currently operates on a timescale of months. SPONSOR MESSAGES: *** Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on o-series style reasoning and AGI. They are hiring a Chief Engineer and ML engineers. Events in Zurich. GTC is coming, the premier AI conference, great opportunity to learn about AI. NVIDIA and partners will showcase breakthroughs in physical AI, AI factories, agentic AI, and inference. Register for virtual GTC for free, using Tim's link (https://nvda.ws/4qQ0LMg) Tufa AI Labs is a new research lab in Zurich started by Benjamin Crouzier focused on o-series style reasoning and AGI. They are hiring a Chief Engineer and ML engineers. Go to https://tufalabs.ai/ --- TIMESTAMPS: 00:00:00 Introduction: AlphaEvolve's Breakthroughs and DeepMind's Lineage 00:11:24 Introducing AlphaEvolve: Evolutionary Architecture and LLM Pairing 00:16:56 The Halting Problem and Evaluation Constraints 00:23:20 Knowledge Augmentation: Meta-Prompting, Library Learning, and Self-Generated Data 00:29:08 Matrix Multiplication Breakthrough: From Strassen to 48 Multiplications 00:39:11 Problem Representation: Direct Solutions, Constructors, and Search Algorithms 00:46:06 Surprising Outcomes: What Researchers Did Not Expect 00:51:42 Hill Climbing, Program Synthesis, and Intelligibility 01:00:24 Real-World Applications: Complex Evaluations and Robotics 01:05:39 The Role of LLMs, Recursive Self-Improvement, and Future Directions --- REFERENCES: Blog Post: [00:00:00] AlphaEvolve Blog Post https://deepmind.google/discover/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/ Paper: [00:03:00] FunSearch https://www.nature.com/articles/s41586-023-06924-6 [00:12:00] MAP-Elites https://arxiv.org/abs/1504.04909 Person: [00:11:24] Matej Balog https://x.com/matejbalog [00:14:10] Alexander Novikov https://x.com/SashaVNovikov Company: [00:11:24] Tufa AI Labs https://tufalabs.ai/ --- LINKS: Full Transcript: https://app.rescript.info/share/de65b6ec8da83b261ce6018039fec289 Download PDF transcript: https://app.rescript.info/api/public/sessions/855d5b15e733d182/pdf Guests: Matej Balog: https://x.com/matejbalog Alexander Novikov: https://x.com/SashaVNovikov

Are Human Drivers Finally Obsolete? | Freakonomics Radio

Are Human Drivers Finally Obsolete? | Freakonomics Radio

What is “reasoning” in modern AI?

LIFE IS SPIRITUAL PRESENTS- ABIGAIL'S STORY ``HOW POVERTY LED ME TO SELL MY BODY ONLINE''

Google's 'DeepMind' does Mathematics - Numberphile Podcast