
Evan Hubinger is Anthropic’s alignment stress test lead. Monte MacDiarmid is a researcher in misalignment science at Anthropic.The two join Big Technology to discuss their new research on reward hacking and emergent misalignment in large language models. Tune in to hear how cheating on coding tests can spiral into models faking alignment, blackmailing fictional CEOs, sabotaging safety tools, and even developing apparent “self-preservation” drives. We also cover Anthropic’s mitigation strategies like inoculation prompting, whether today’s failures are a preview of something far worse, how much to trust labs to police themselves, and what it really means to talk about an AI’s “psychology.” Hit play for a clear-eyed, concrete, and unnervingly fun tour through the frontier of AI safety. 00:00:00 - Introduction 00:03:32 - Why Do Models Cheat? 00:12:14 - Understanding Alignment and Alignment Faking 00:24:05 - The AI's Psychotic Behavior 00:26:38 - AI's Misalignment 00:30:49 - Using AI for Research 00:32:58 - The AI's Internalization of Cheating 00:35:20 - Context-Dependent Misalignment 00:36:06 - The AI's Understanding of Cheating 00:39:24 - Preventing Misalignment 00:43:08 - The AI's Interpretation of Cheating 00:44:32 - Concerns Over AI Behavior 00:45:52 - Inoculation Prompting and Mitigation 00:51:21 - Addressing Criticisms of Anthropic 00:57:09 - Anthropomorphizing AI Models

Could LLMs Be The Route To Superintelligence? — With Mustafa Suleyman

Is Something Big Happening?, AI Safety Apocalypse, Anthropic Raises $30 Billion

Ex OpenAI Researcher: Total Job Loss IMMINENT

Rogue Coordinated AI SWARM Commits Secret CRIME SPREE

Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question