How Agents Decide: Goodfire's Eric Bigelow on Critical Tokens, Phase Shifts, & In-Context LearningGoodfire researcher Eric Bigelow joins the show to investigate how large language models arrive at decisions at the level of mechanistic interpretability. Drawing on his research into forking paths, Eric explains that model reasoning functions as in-context learning from sampled tokens, where outcome distributions can abruptly collapse at single, critical tokens. The conversation explores the performative nature of chain of thought in reasoning models like DeepSeek-R1, the impacts of reinforcement learning, and widespread reward hacking observed in frontier models like Kimi K3. As confidence in chain-of-thought monitoring declines, understanding these underlying decision mechanics becomes essential for evaluating model alignment.
For full show notes, links, and references, read the episode page:https://www.cognitiverevolution.ai/how-agents-decide-goodfire-s-eric-bigelow-on-critical-tokens-phase-shifts-in-context-learning/
Sponsors:
Parallel: Parallel provides enterprise-grade web search APIs for AI agents, offering the optimal balance of quality, speed, and cost. Get started for free at https://parallel.ai/tcr
Claude: Claude is the AI collaborator for problem solvers, helping with writing, coding, financial models, strategy, and more. Get started with Claude and explore Claude Pro at https://claude.ai/tcr
Tasklet: Tasklet empowers your business with AI agents that connect to your tools and automate recurring workflows with no code required. Visit https://tasklet.ai and use code cog rev for $50 in free credits
CHAPTERS:
(00:00) About the Episode
(03:53) Forking paths in reasoning (Part 1)
(15:53) Sponsors: Parallel | Claude
(18:47) Forking paths in reasoning (Part 2)
(19:40) Dynamics of in-context learning (Part 1)
(28:16) Sponsor: Tasklet
(29:53) Dynamics of in-context learning (Part 2)
(29:56) Continual learning and interpretability
(38:55) Shifting distributions through RL
(46:41) Selecting models for research
(57:11) Cognitive science and LLMs
(01:07:35) Modeling beliefs in AI
(01:16:22) Phase shifts and uncertainty
(01:31:49) Uninterpretable chain of thought
(01:45:57) Automating research with Silico
(01:50:57) Safety and alignment frontiers
(01:57:07) Episode Outro
(02:00:02) Outro
PRODUCED BY:
https://aipodcast.ing
SOCIAL LINKS:
Website: https://www.cognitiverevolution.ai
Twitter (Podcast): https://x.com/cogrev_podcast
Twitter (Nathan): https://x.com/labenz
LinkedIn: https://linkedin.com/in/nathanlabenz/
Youtube: https://youtube.com/@CognitiveRevolutionPodcast
Apple: https://podcasts.apple.com/de/podcast/the-cognitive-revolution-ai-builders-researchers-and/id1669813431
Spotify: https://open.spotify.com/show/6yHyok3M3BjqzR0VB5MSyk
389|2T
