
Iman Mirzadeh is a machine learning research engineer at Apple and the lead author of the GSM-Symbolic paper, which exposed deep fragility in how large language models handle mathematical reasoning. In this conversation, he draws a sharp line between intelligence and achievement -- between what a system can score on a benchmark and what it actually understands. The discussion starts with chess. Mirzadeh explains how grandmasters don't use engines to memorize moves; they use them to develop theory. AlphaZero discovered unprecedented strategies, but that knowledge stays trapped in the game. Humans, by contrast, extract abstract principles like 'control the center' and transfer them to entirely different domains. That capacity for abstraction is what he thinks current AI architecturally lacks. His critique of LLMs is structural. These systems are trained to minimize cross-entropy loss over a distribution, and by construction they cannot reason beyond what that distribution contains. Change the surface form of a problem -- swap names, add irrelevant clauses -- and performance varies wildly, even on grade-school math. That's the core finding of GSM-Symbolic. By generating templated variants of math word problems, Mirzadeh's team showed that even frontier models exhibit large performance variance from changes that should be semantically irrelevant. The implication: what looks like reasoning is closer to sophisticated pattern matching across memorized distributions. Mirzadeh proposes that intelligence should be measured by the slope of a system's scaling -- how fast it can learn novel things -- rather than its current benchmark position. The conversation also covers the connectionism-symbolism divide, active engagement and agency as necessary conditions for learning, and why we might need a fundamentally different vessel to reach genuine reasoning. SPONSOR MESSAGES: *** Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on o-series style reasoning and AGI. They are hiring a Chief Engineer and ML engineers. Events in Zurich. --- TIMESTAMPS: 00:00:00 Intelligence vs Achievement in AI 00:03:27 AlphaZero and Abstract Understanding in Chess 00:10:10 Language Models as Distribution Learners 00:14:47 The State of AI Research Methodology 00:24:24 Interpolation vs True Reasoning in LLMs 00:29:00 Measuring Intelligence: From Chollet to the Iman Moon Test 00:35:35 Agency, Active Learning, and World Models 00:47:15 Scaling Laws and the Connectionism-Symbolism Debate 00:58:09 GSM-Symbolic: Exposing LLM Reasoning Fragility --- REFERENCES: paper: [00:00:55] Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm https://arxiv.org/abs/1712.01815 [00:17:15] GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models https://arxiv.org/abs/2410.05229 [00:21:20] Connectionism and Cognitive Architecture: A Critical Analysis https://www.sciencedirect.com/science/article/pii/001002779090014B [00:29:35] On the Measure of Intelligence https://arxiv.org/abs/1911.01547 [00:33:25] On definition of intelligence https://www.sciencedirect.com/science/article/pii/S0160289624000266 [00:35:25] Defining Intelligence https://cis.temple.edu/~wangp/papers.html [00:43:10] Chain-of-Thought Prompting Elicits Reasoning in Large Language Models https://arxiv.org/abs/2201.11903 [00:47:45] Scaling Laws for Neural Language Models https://arxiv.org/abs/2001.08361 [00:55:10] Tensor Product Variable Binding and the Representation of Symbolic Structures in Connectionist Systems https://www.sciencedirect.com/science/article/abs/pii/000437029090007M book: [00:07:05] Game Changer: AlphaZero's Groundbreaking Chess Strategies https://www.amazon.com/Game-Changer-AlphaZeros-Groundbreaking-Strategies/dp/9056918184 [00:37:35] How We Learn: Why Brains Learn Better Than Any Machine... for Now https://www.amazon.com/How-We-Learn-Brains-Machine/dp/0525559884 [00:39:30] Surfaces and Essences: Analogy as the Fuel and Fire of Thinking https://www.amazon.com/Surfaces-Essences-Analogy-Fuel-Thinking/dp/0465018475 reference: [00:11:30] NLP Course: Language Modeling http://lena-voita.github.io/nlp_course/language_modeling.html [01:08:40] GSM8K: Training Verifiers to Solve Math Word Problems https://huggingface.co/datasets/openai/gsm8k --- LINKS: Full Transcript: https://app.rescript.info/share/72689965572c7fd460954f70e26b2eaa Download PDF transcript: https://app.rescript.info/api/public/sessions/abf4f57b59fcbbd5/pdf