Episode 494

A Coin Toss for the Future

A Conversation with Ryan Greenblatt

September 22, 2026

Sam Harris speaks with Ryan Greenblatt about AI misalignment and the risk of losing control of increasingly capable systems. They discuss the Hugging Face incident, the spectrum of concern about AI risk, why companies are racing ahead despite high odds of catastrophe, reward hacking, the distinction between alignment and control, the danger of AIs reasoning in "neuralese," alignment faking, how an AI takeover might unfold, and other topics.

Ryan Greenblatt is Chief Scientist at Redwood Research, where he works on technical AI security and safety research focused on risks from misaligned AI. He was the primary empirical researcher on the METR/Redwood independent investigation of the OpenAI/Hugging Face agent hacking incident. He was the lead author of the paper that introduced the AI control agenda (ICML 2024) and of "Alignment Faking in Large Language Models" (2024). He holds a BS in Applied Mathematics and Computer Science from Brown University.

 

Website: redwoodresearch.org

X: @RyanGreenblatt