
Panel discussion with Francois Chollet, Kevin Ellis, and Zenna Tavares on why program synthesis matters and where deep learning falls short. Chollet recounts how his early work on theorem proving with Christian Szegedy at Google made him realise gradient descent cannot learn discrete algorithms, even when the correct solution is representable by the network. Ellis, whose PhD with Armando Solar-Lezama helped shape the modern program synthesis field, asks how much of the bottleneck is the learning mechanism versus the representation. Tavares considers a deeper integration of neural networks into programming language semantics, where neural operators implement the interpreter rather than sitting outside it. The group discusses the limits of transformers at function composition, the failure of Cyc-style hand-built ontologies, and what ARC has revealed about strong generalisation. Chollet explains how test-time training and O1-style iterative program writing let static models adapt to novelty, then previews ARC 2, which will include human difficulty data and push harder on compositional complexity. Ellis and Tavares describe MARA, their new project that extends ARC-style tasks toward active experimentation where the agent must choose what questions to ask. Recorded as part of a broader discussion at the intersection of program synthesis, neural-symbolic integration, and abstract reasoning. Published March 2025. SPONSOR MESSAGES: *** Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on o-series style reasoning and AGI. They are hiring a Chief Engineer and ML engineers. Events in Zurich. --- REFERENCES: website: [00:00:01] Basis Research Institute https://www.basis.ai/ [00:14:30] Keras https://keras.io/ [00:14:50] Armando Solar-Lezama https://www.csail.mit.edu/news/solar-lezama-wins-robin-milner-young-researcher-award [00:15:05] Kevin Ellis https://www.cs.cornell.edu/~ellisk/ [00:18:50] Cyc Project https://en.wikipedia.org/wiki/Cyc [00:28:10] ARC Prize https://arcprize.org/ paper: [00:01:00] HolStep Dataset https://openreview.net/pdf?id=ryuxYmvel [00:05:20] On the Measure of Intelligence https://arxiv.org/abs/1911.01547 [00:07:00] Neural Turing Machines https://arxiv.org/pdf/1410.5401 [00:07:20] Manifold Hypothesis https://arxiv.org/abs/2208.05314 [00:21:50] Test-Time Training on Nearest Neighbors https://ekinakyurek.github.io/papers/ttt.pdf [00:26:20] AlphaZero-style Program Synthesis https://arxiv.org/abs/2205.14229 --- LINKS: Full Transcript: https://app.rescript.info/share/b8e9612724c01a88ef103804be1e79d5 Download PDF transcript: https://app.rescript.info/api/public/sessions/3380bd2c998bf22d/pdf Francois Chollet: https://x.com/fchollet https://ndea.com/ https://arcprize.org/ [00:21:55] Test-Time Training, Akyurek et al. https://ekinakyurek.github.io/papers/ttt.pdf

Why Program Synthesis Is Next (Kevin Ellis and Zenna Tavares)

ARC Prize Version 2 Launch Video! [Francois Chollet, Mike Knoop]

François Chollet on OpenAI o-models and ARC

Pattern Recognition vs True Intelligence - Francois Chollet

It's Not About Scale, It's About Abstraction