
Eiso Kant is the co-founder and CTO of Poolside AI, one of roughly seven companies worldwide with the technical muscle to build frontier foundation models from scratch. He sat down with Tim to explain why Poolside deliberately rejected the prevailing wisdom of "just scale up the next GPT" and instead bet the company on a thesis most labs were ignoring: that reinforcement learning from code execution feedback is the missing scaling axis. The argument is straightforward and, once you hear it, hard to unsee. Next-token prediction is imitation learning. Reinforcement learning is trial-and-error learning. Poolside maintains close to a million fully containerized code repositories -- each with its own test suite -- as a massive, diverse RL environment. The model writes code, executes it, gets deterministic feedback, and learns. This is why software turns out to be the ideal domain for RL scaling: it is deterministic enough to provide clear reward signals, but diverse enough to avoid model collapse. Along the way, Eiso and Tim get into the weeds on frontier lab operations (over 4000 experimental runs per month), why Chinchilla optimality breaks down once you account for inference cost, what DeepSeek V3 tells us about the second-generation AI company playbook, and why the R1 zero-shot reasoning result should have been the real headline rather than the dollar figure. There is a candid exchange on whether the software development lifecycle itself will progressively collapse into the model, what Chris Olah interpretability work means for alignment, and whether Karpathy was right about Software 2.0 -- with some important caveats Eiso has developed after a decade of building AI for code. SPONSOR MESSAGES: *** Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on o-series style reasoning and AGI. They are hiring a Chief Engineer and ML engineers. Events in Zurich. --- TIMESTAMPS: 00:00:00 Introduction and Guest Background 00:02:50 Poolside AI's Vision and Three-Step Plan 00:06:50 Foundation Models vs. Enterprise Customization 00:10:25 The Missing Scaling Axis: Reinforcement Learning 00:15:40 Reinforcement Learning from Code Execution Feedback 00:22:20 Model Economics and Experimental Optimization 00:26:00 Enterprise Deployment Strategy and Market Focus 00:30:30 DeepSeek, Distributed Training, and Hardware Architecture 00:36:40 Emergent Reasoning and Chain-of-Thought Scaling 00:45:00 AI-Assisted Software Development Today 00:58:20 Architecture Innovation and Model Interpretability 01:15:00 Karpathy's Software 2.0 and the Future of Code 01:25:00 AWS Partnership, Enterprise Security, and Closing --- REFERENCES: website: [00:01:40] Tufa AI Labs https://tufalabs.ai/ [00:02:50] Poolside AI https://poolside.ai/ social: [00:02:50] Eiso Kant on X https://x.com/eisokant paper: [00:15:40] The Curse of Recursion: Training on Generated Data Makes Models Forget https://arxiv.org/abs/2305.17493 [00:20:00] Efficient Estimation of Word Representations in Vector Space (Word2Vec) https://arxiv.org/abs/1301.3781 [00:22:40] On the Measure of Intelligence https://arxiv.org/abs/1911.01547 [00:30:30] DeepSeek-V3 Technical Report https://arxiv.org/abs/2412.19437 [00:34:30] Training Compute-Optimal Large Language Models (Chinchilla) https://arxiv.org/abs/2203.15556 [00:45:45] The Bitter Lesson http://www.incompleteideas.net/IncIdeas/BitterLesson.html [00:49:55] Mastering the Game of Go with Deep Neural Networks and Tree Search https://www.nature.com/articles/nature16961 [01:10:15] Mamba: Linear-Time Sequence Modeling with Selective State Spaces https://arxiv.org/abs/2312.00752 [01:14:25] Zoom In: An Introduction to Circuits https://distill.pub/2020/circuits/zoom-in/ benchmark: [00:46:10] ARC Prize Challenge https://arcprize.org/ blog: [01:19:10] Software 2.0 https://karpathy.medium.com/software-2-0-a64152b37c35 --- LINKS: Full Transcript: https://app.rescript.info/share/a8144052b0a5fc7210ab37f9651a3557 Download PDF transcript: https://app.rescript.info/api/public/sessions/2f5beb6d70f79973/pdf Eiso Kant: https://x.com/eisokant https://poolside.ai/