
Advertise on podcast: The Information Bottleneck
Rating
5from
This podcast has
72 episodes
Language
EnglishPublisher
Ravid Shwartz-Ziv & Allen RoushExplicit
No
Date created
2025/08/21
Latest episode
2026/09/27
Average duration
70 min.
Release period
4 days
Description
Two AI Researchers - Ravid Shwartz Ziv, and Allen Roush, discuss the latest trends, news, and research within Generative AI, LLMs, GPUs, and Cloud Systems.
Unlock The Information Bottleneck podcast Email contact info,
Listeners & Audience details
Email contact information
Direct podcast contact details

Listeners
Audience numbers & engagement insights

Audience details
Podcast Insights

Podcast episodes
Check latest episodes from The Information Bottleneck podcast
Chris Manning: Language Is the Real Unlock for Intelligence
2026/09/27
Chris Manning is a legend in NLP. He's a professor of linguistics and computer science at Stanford and ran the Stanford AI Lab. His course CS224N, Natural Language Processing with Deep Learning, is where a whole generation of researchers learned the field. If you've worked on anything in NLP over the last 25 years, you've probably used his work.
We start with what linguistics gave machine learning that ML wouldn't have figured out on its own, and what's left for NLP researchers now that LLMs handle most of the classic tasks. Chris thinks the open problems have moved up the stack to pragmatics and dialogue. Models are good at using context but still sound confident when they shouldn't.
Then we get into his recent work on how language models learn verb classes. Small GPT-2 models seem to form abstract categories right away instead of memorizing verbs one at a time, and Chris explains why distributed representations push them in that direction.
The second half is about meaning and representation. Chris makes the case that Yann LeCun underrates the role of language in intelligence, and revisits his debate with Emily Bender over whether text alone can teach meaning. We finish with ReFT, which steers a frozen model by editing its hidden states, and whether concepts really live in linear subspaces.
Timeline
00:00 Intro
01:05 What linguistics gave machine learning
03:27 Is NLP losing its focus on language?
06:29 What's left for NLP researchers in the LLM era
09:37 The next frontier: pragmatics and dialogue
13:01 Constructed languages and AI-to-AI communication
15:35 Do language models learn categories first?
23:43 Why transformers learn abstractions early
26:37 Does it hold at scale?
28:30 LeCun, JEPA and the role of language in intelligence
34:54 Can text alone teach meaning? The octopus debate
40:45 Diffusion language models vs transformers
43:39 ReFT: editing representations instead of weights
50:34 Weights or representations: where knowledge lives
52:37 Do concepts really live in linear subspaces?
Topics
NLP, computational linguistics, large language models, learning dynamics, meaning from form, world models, JEPA, diffusion LMs, ReFT, representation learning, interpretability
Music
"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0
Why AI Still Can't See wtih Andrew Dai
2026/09/24
Andrew Dai spent over a decade at Google Brain and DeepMind, where he co-wrote the 2015 paper that introduced language model pre-training followed by fine-tuning, and later co-led pre-training data for Gemini. He's now co-founder and CEO of Elorian AI, which is building models for visual reasoning.
We talk about how his pre-training result started as a bug, why next-token prediction scales better than other objectives, and what went wrong for Google in the early LLM race. The second half is about vision: why today's frontier models still can't count objects in a photo, why he thinks reasoning is fundamentally visual, and how his view of world models differs from JEPA.
Chapters
00:00 Intro
00:48 Andrew's background and the accidental discovery of pre-training
07:10 Why next-token prediction scales
10:09 How Google fell behind and the early days of Gemini
18:28 What makes training data good
28:16 Where visual understanding breaks down
46:01 World models, JEPA and robotics
55:27 Generation vs understanding, and what's next for Elorian
Topics
Pre-training and fine-tuningScaling and next-token predictionGemini and Google's AI historyData quality and synthetic dataVisual reasoning and countingWorld models and JEPAMusic
"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0
Can AI Replace Doctors? Zachary Lipton on the Future of Healthcare and AI
2026/09/21
Zachary Lipton is an associate professor at Carnegie Mellon University and a co-founder of Abridge, a healthcare AI company (and a jazz saxophonist!). His research spans machine learning, healthcare, and the broader impact of AI.
In this episode, we discuss what AI can (and cannot yet) do in healthcare. We talk about why medicine is harder to automate than coding, AI scribes and clinical decision support, drug discovery, whether AI could eventually replace doctors, and the role of open models and intelligent routing.
In the second half, we turn to the future of AI research: whether academia can still compete, what AI PhD students should work on, and how Zach thinks about automation and the future of research.
Timeline00:00 — Introduction
00:28 — Why healthcare is hard for AI
05:00 — Zach’s path into healthcare AI
22:13 — Where AI can have the biggest impact
32:01 — Can AI replace doctors?
41:41 — Open models and AI infrastructure
56:20 — The future of AI research
1:03:03 — What should AI PhD students work on?
1:27:00 — Advice for researchers and founders
TopicsAI in healthcare • AI doctors • drug discovery • clinical decision support • open-source AI • AI research • academia vs. industry • future of work
Music
"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0
Sara Hooker on the End of Static AI
2026/09/16
What comes after scaling?
We talk with Sara Hooker, co-founder and CEO of Adaptation Lab, about why the next generation of AI may look very different from today's static models. Sara argues that models should continuously adapt to new tasks, data, users, and environments—and that doing this efficiently will require rethinking much more than fine-tuning.
We discuss continual learning, AutoScientist and automated research, why non-verifiable tasks may become the next major bottleneck, and why interfaces could be as important as the models themselves. We also get into open vs. closed models, distillation and Chinese AI labs, AI regulation and safety, cybersecurity and biorisk, AI companionship, and what may eventually come after Transformers and tokenization.
TopicsContinuous learning and adaptive AIFine-tuning, memory, and AutoScientistAI agents and automated researchNon-verifiable tasks and human feedbackAdaptive interfacesOpen vs. closed models and distillationAI safety, regulation, cyber risk, and bioriskAI companionship and persuasionThe limits of TransformersMultilingual models and tokenizationChapters00:00 — Introduction
02:15 — Why start another AI lab? The return of research
05:46 — What continuous learning actually means
12:04 — Should every company have its own adapting model?
13:59 — Fine-tuning and platforms like Tinker
18:04 — AutoScientist and automated optimization
22:52 — Can AI really improve its own research?
28:38 — The problem of non-verifiable tasks
31:30 — Human feedback and the limits of exponential progress
34:43 — Why the AI interface matters
40:36 — Distillation, China, and open models
49:05 — Open-model licensing
52:19 — Will open models catch closed models?
58:43 — AI regulation and compute thresholds
1:03:07 — AI safety and agent failures
1:10:19 — Biorisk vs. cybersecurity
1:14:03 — Persuasion, AI companionship, and overlooked risks
1:20:41 — Where will AI have the biggest real-world impact?
1:25:41 — What is missing from current AI architectures?
1:29:03 — Neurosymbolic AI
1:31:30 — Multilingual models and tokenization
1:34:02 — Byte-level models and alternatives to tokenization
1:35:03 — Closing
Music
"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0
Tiny Recursive Models Beat the Giants - Alexia Jolicoeur-Martineau (Microsoft)
2026/09/15
Alexia Jolicoeur-Martineau is a Principal Researcher at Microsoft and the author of "Less is More: Recursive Reasoning with Tiny Networks," the paper behind the Tiny Recursive Model that hit about 45% on ARC-AGI-1 with a fraction of the parameters of frontier systems. It won the 2025 ARC Prize paper award.
She read the hierarchical reasoning paper, thought the potential was real and the explanation was not, and rebuilt it without the mouse brains: a small network that carries a hidden state and a current answer, thinks for a few steps, updates, and repeats, with the gradient truncated at each loop. We get into why puzzles suit this and autoregression doesn't, why she thinks LLMs are bad at molecules and more data won't fix it, and what she'd do with a trillion dollars.
Timeline
00:01 Intro01:06 Leaving biostatistics, and why the field stagnated06:47 GANs, diffusion, and research on four GPUs12:58 What was wrong with the hierarchical reasoning paper16:31 Tiny recursive models explained without the biology22:35 Why puzzles favor recursion over left to right generation24:15 Is the bitter lesson really bitter?27:28 With infinite compute, would you still want small models?32:00 Self improvement, memory, and a trillion dollars37:01 Test time compute beyond chain of thought40:41 Why chain of thought fails on molecules45:17 Is there a universal representation?48:06 What people are already building with TRM55:22 Fixed point models and DEQMusic
"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0Topics
Tiny Recursive Models and the ARC-AGI resultsWhat the hierarchical reasoning model was really doingDeep supervision and truncated backpropLooping transformers and parameter efficiencyWhy puzzles favor whole-context iteration over left to right generationTest time compute beyond chain of thoughtLatent reasoning and the Coconut line of workWhy LLMs fail on chemistry and physicsRepresentation learning and whether a universal representation exists
Continual Learning Is the Next Bottleneck | Rohan Anil (Core Automation )
2026/09/10
Rohan Anil spent eleven and a half years at Google, where he went from writing memory allocators to large-scale linear solvers, then optimization at Google Brain, where he co-developed distributed Shampoo and led optimization for PaLM and Gemini pre-training, including the work that produced Gemini Flash. He then joined Anthropic's pre-training team, and left before the IPO to co-found Core Automation with Jerry Tworek (ex-VP of Research at OpenAI). We talk with him about how Brain worked at its peak, why he left two of the world's best labs, and what he thinks is missing from today's models.
Rohan's view is that pre-training and RL were split by organizational convenience rather than by science. Pre-training builds a prior, and RL sharpens it to the tasks we care about, and neither gives a model a way to absorb new data or learn from its own experience once it is deployed. Post-training more every day plateaus, on-policy distillation plateaus, and in-context learning only goes as far as the context does. He argues the next architecture needs better ways to fold in new knowledge at inference time, and that this is a fundamental optimization question rather than a harness-engineering one.
We also get into why coding agents still fail on low-level systems work, his take on Muon, why second-order methods matter once you leave the noise-dominated regime, and why nobody can yet use a few million GPUs for a single training run.
Timeline
00:00 Intro
01:09 From computer vision to Google systems engineering
02:37 Large-scale linear solvers and sparse features
06:01 Getting into optimization: SDCA and Yonghui Wu's team
07:29 Joining the Shampoo crew
09:35 The Google Brain ethos, and why 2017 to 2019 was special
14:23 Is open research going to keep winning?
15:53 Frontier models are only as good as the prior you give them
17:30 Missing the language model wave, then Common Crawl and online distillation
18:31 Paternity leave, DALL-E Mini, and the 14 days that became two years
20:32 PaLM, Gemini pre-training, and Gemini Flash
23:59 The Shampoo origin story: Tomer Koren's two-week proof
26:54 Why leave Google for Anthropic
30:20 Why leave Anthropic for a startup
31:33 Meeting Jerry Tworek at Dolores Park
34:00 What Core Automation is building
36:17 Continual learning and the pre-training vs RL split
39:11 Why coding agents fail at kernels and low-level pipelines
42:00 The QR factorization kernel competition and reward hacking
44:57 Numerics, verification, and hardware that keeps changing
46:07 Are LLMs creative, or just good at search?
49:50 Getting models to extrapolate instead of interpolate
52:03 Why did we ever call it pre-training?
55:02 What RL is really learning
56:27 Competing with the big labs with fewer people
58:31 Will kernel generation keep old GPUs alive? Amdahl's law
1:01:58 Open source plans
1:02:55 Audience question: agentic optimizers
1:04:53 Audience question: Muon, Shampoo, and the future of second-order methods
1:08:58 Hiring at Core Automation
key topics
Journey from Google Brain to startupEvolution of AI research and optimizationPre-training and reinforcement learningKernel optimization and system efficiencyOpen source AI and collaborative researchChallenges in AI creativity and explorationFuture directions in continual learning and model scaling
Music"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0
World Models | John Langford (Microsoft AI Labs)
2026/09/05
John Langford, one of the heads of Microsoft's AI Labs, the creator of Vowpal Wabbit, and a co-inventor of CAPTCHA, joins us to talk about world models. Transformers need orders of magnitude more data than humans to learn the same thing, and John argues a compact, implicit world model is how you close that gap. He explains why he's skeptical of JEPA-style objectives, why a transformer's KV cache is the Ptolemaic epicycle model of belief states, and what his Next Latent work does differently.
We also get into whether research still matters in the age of scale; open versus closed models; agent-driven research after running 2,000 pre-training experiments in 90 days; the origin story of CAPTCHA; and why Muon and orthonormal optimizers actually work.
Topics:
Implicit vs. explicit world models, and the case against JEPA-style objectivesCompact belief states: why compression beats a growing KV cacheDoes research still matter in the age of scale? The Kimi K3 argumentAgent-driven research: 2,000 pre-training experiments in 90 daysThe invention of CAPTCHAOptimizers from SGD and Vowpal Wabbit to Muon
Chapters00:00 Why world models: the sample-complexity gap09:48 The case against JEPA; a transformer-style implicit world model15:52 Compact belief states: epicycles vs. heliocentrism23:41 Does research still matter? The Kimi K3 argument27:35 Open vs. closed models35:57 Recursive self-improvement and agent-driven research42:30 2,000 pre-training experiments in 90 days; weak baselines and reproducibility54:54 The invention of CAPTCHA1:00:53 Optimizers: from Vowpal Wabbit to MuonMusic
"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0
Which Tabular Model Should You Actually Use? | David Holzmüller (INRIA)
2026/09/03
Description
Tabular data is still where most of machine learning actually happens in industry, and the field has changed a lot in the last few years. In this episode we talk with David Holzmüller, a researcher at INRIA and one of the people behind TabArena, TabICL and RealMLP, about what the state of the art looks like right now and how to pick a model for your own data.
We cover the shift to TabPFN-style foundation models that learn to learn from whole tables, why TabArena was built and what earlier benchmarks got wrong, what Google's new TabFM means for the leaderboard, and when gradient boosted trees are still the right tool. David explains why LLMs struggle with tables, shares an early result comparing Claude Opus against TabICL on tiny datasets, and walks through how to embed text columns for tabular models. We also get into time series vs tabular data, the open research problems he thinks matter most, and why classical ML libraries are so bad out of the box.
Links:
TabArena: https://tabarena.ai
Topics
Tabular foundation models and in-context learning on tablesTabArena and Beyond Arena: building a benchmark that stays honestTabFM, TabPFN, TabICL and the tradeoffs between themWhen boosted trees and MLPs still win (large data, CPU, fast inference)Why LLMs are inefficient on tabular data and where they might helpEmbedding text columns with language modelsExplainability, calibration and class imbalanceTime series vs tabular dataOpen problems: invariances, synthetic data, uncertainty, scaling downWhere the field is heading in the next five years
Chapters
0:00 Intro
0:31 What changed in tabular ML: TabPFN-style foundation models
2:22 Which model to try first? TabArena and how it was built
5:14 What older benchmarks got wrong, and Beyond Arena
8:45 GPU AutoML vs foundation models
10:40 Reading the leaderboard: TabFM, TabPFN, TabICL and the tradeoffs
12:47 Calibration, class imbalance and small vs large data
19:45 Explainability for black-box tabular models
21:34 Why LLMs are bad at tabular data
25:39 Claude Opus 4.6 vs TabICL on tiny datasets
27:51 New classifiers, five-year outlook, real vs synthetic pretraining
33:13 Embedding text columns for tabular foundation models
36:13 Time series vs tabular data
39:59 When gradient boosted trees still win, and feature engineering
45:31 Open research problems and where the field is heading
52:54 Better MLPs and why classical defaults are bad out of the box
Music"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0
Why You Can't Just Rent 1,000 GPUs | Charles Frye (Modal)
2026/09/01
Charles Frye (Modal, ex-Weights & Biases, Berkeley PhD) joins Ravid and Allen to explain why modern AI research is bottlenecked by compute, and why simply buying more GPUs doesn't solve it. We cover the three problems every lab hits (underutilization, saturation, resource sharing), when companies should actually train their own models, why inference is a "bad algorithm" for today's hardware, NVIDIA's monopoly, the OpenAI/Hugging Face hack and what it says about open models, and whether we're in a compute bubble.
Key topics
AI infrastructure challenges and when to train your own modelsGPU resource management and virtualizationInference optimization and speculative decodingThe economics and future of AI hardwareAgents, sandboxing, and open-model security
Chapters
00:00 Intro
01:03 Why AI needs special-purpose compute
03:22 Buying vs renting GPUs: the three problems
07:15 Modal's approach, and doing more with less compute
09:46 Do we actually need to spend more? The conflict-of-interest question
13:08 Should companies train their own models?
14:47 Efficient fine-tuning and prompts as fast weights
17:37 Are we in a compute bubble?
20:21 Why inference will dominate compute (the SQLite analogy)
22:42 Speculative decoding
26:44 Why scaling inference is hard, and neuromorphic hardware
28:36 Why NVIDIA's monopoly persists
33:09 Inference chip startups and the hardware lottery
35:24 How Modal stays hardware-agnostic (GPU snapshot restore)
38:45 Will agentic coding erode CUDA's moat?
41:18 Running one agent vs thousands: sandboxing at scale
46:27 The OpenAI/Hugging Face hack and open models as defenders
52:28 Rogue AI, self-replication, and fast takeoff
56:09 What's next: evals, embodiment, edge inference
1:00:27 Modal is hiring (modal.jobs)
Music"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0
Stella Biderman (EleutherAI) - Open Source, AI Safety, and Who We Can Trust
2026/08/28
Stella Biderman, Executive Director of EleutherAI, joins us the week an OpenAI model autonomously broke out of its sandbox and hacked Hugging Face. Stella calls it what she thinks it is, an offensive cyber operation, and argues it's part of a pattern: this is not the first containment failure at a frontier lab, and sandboxes have failed basically every time they've been tested for real.
So we spend a good chunk of the episode on what actual containment would look like. Stella's argument is that the tools already exist, the labs just don't use them: run dangerous capability evals on air-gapped networks with no route to the public internet, put the most sensitive testing in SCIF-style secure facilities, and treat model evaluation the way the security world treats classified systems rather than the way startups treat staging environments.
And yet Stella remains one of the world's most prominent open-source advocates. From her perspective, the biggest risk isn't the technology; it's unchecked corporate power, and the only durable check on it is an independent scientific research establishment that doesn't depend on the AI industry for its funding or its facts.
From there the conversation spans the geopolitics of Chinese open models and whether governments can restrict them, sovereign AI and what it would actually take for other countries to train their own models, why harnesses and UX drive more of AI's perceived progress than raw intelligence, the AI-found counterexample to the Jacobian conjecture, and EleutherAI's "Deep Ignorance" approach to making open-weight models safe by filtering hazardous knowledge out of pretraining.
key topics
AI governance and regulationCybersecurity incidents involving AI modelsOpen source AI safety and securityThe role of independent research in AI safetyLegal and ethical considerations in AI developmentTimeline
00:13 — Intro: Stella Biderman and EleutherAI, a real non-profit in AI02:05 — News of the week: Kimi K3, and OpenAI's model autonomously hacking Hugging Face05:49 — "Frontier labs can't be trusted": repeated containment failures, air-gapped networks and SCIFs vs. sandboxes22:45 — Can governments ban open or Chinese models? Import restrictions and the six-month open/closed gap27:05 — Why Stella is still pro-open-source: unchecked corporate power as the real danger31:11 — The opioid epidemic analogy: avoiding both regulatory failure and overcorrection34:57 — Offense vs. defense: why open access to AI has empirically favored defenders37:28 — Chinese labs, the CCP, and why safety and fine-tuning are low-prestige work in China42:19 — Sovereign AI: does every country need its own foundation model?49:29 — Sampling, harnesses, and why ChatGPT was really a UX breakthrough54:09 — AI solves the Jacobian conjecture: domain data beats raw intelligence58:02 — Safety is contextual, not a model property — and what HAL 9000 got right1:01:42 — Is Stella optimistic about the future?1:02:50 — Deep Ignorance, the science of AI training dynamics, and how to get involved with EleutherAIMusic"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.
Why Deep Learning Finally Works on Tables | Frank Hutter (Prior Labs)
2026/08/24
In this episode, Frank Hutter joins us to talk about TabPFN and why tabular data is suddenly the hottest problem in deep learning. Frank is a professor at the University of Freiburg and spent 15 years building the AutoML field before founding Prior Labs, which SAP just acquired for over a billion dollars.
We get into why deep learning failed on tables for a decade and what in-context learning changed, how TabPFN is trained entirely on synthetic data, and why a model that never saw a real time series ended up beating specialized forecasting models. Frank also explains the architecture tricks behind scaling from 10,000 to a million rows, where LLMs fit into data science (and where they embarrassingly don't), and what happens to XGBoost from here.
Beyond the research, Frank talks about the jump from professor to co-CEO, why he refused to merge his 45-person team into SAP's 110,000 employees, the open-weights licensing debate, and the case for building a frontier lab in Freiburg rather than San Francisco.
key topics
The role of foundation models in tabular dataImpact of SAP acquisition on Pro LabsThe evolution of AutoML and hyperparameter optimizationChallenges and solutions for large context in modelsOpen source models and licensing strategiesThe importance of independence for startup agilityFuture directions in AI for science and medicine00:00 Intro
00:34 The SAP acquisition and staying independent
07:39 Why tabular data is the next big thing in deep learning
14:19 What makes tabular data hard
19:14 AutoML, AutoGluon, and fifteen years of hyperparameter tuning
28:27 Scaling TabPFN: context limits and architectures
34:35 Agentic data science and LLMs
39:30 Online learning, time series, and Bayesian inference in a forward pass
47:05 Open weights and the license debate
54:51 Will LLMs and tabular models merge?
1:00:01 From academia to startup
1:09:42 Why build in Europe
1:12:53 Audience questions and hiring
Music"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.
Surya Ganguli: The Physics of Intelligence
2026/08/17
Surya Ganguli is a professor at Stanford and VP at General Catalyst, working at the intersection of physics, neuroscience, and AI. He started in string theory, moved to theoretical neuroscience, and now uses tools from statistical physics to understand both brains and neural networks.
We talk about why deep learning theory is finally catching up to practice, including his group's recent work explaining neural scaling laws, and why smarter data selection could beat them entirely. He also tells the origin story of diffusion models, which were invented in his lab as an attempt to violate the second law of thermodynamics.
The second half turns to the brain: what happens to a mouse's sense of self on ketamine, how stimulating a handful of neurons can induce hallucinations, and a method his lab developed to get a neuron deep in a monkey's brain to describe, in English, what makes it fire.
We close on where he thinks AI is going wrong: models train on ten trillion tokens while humans hear a hundred million words, because we don't teach children with gradients; we tell them the algorithm.
key topics
Connections between physics, neuroscience, and AIEmergent properties in complex systemsScaling laws in language modelsData efficiency and pruning in AINeuroscience insights into consciousness and selfThe future of AI and brain modelingChapters
00:00 Introduction to Surya Ganguli
00:57 Surya's Background: From String Theory to Neuroscience
02:22 Emergent Properties in Physics, Neuroscience, and AI
03:16 Energy Landscapes and Loss Landscapes in High Dimensions
04:07 Why Local Minima Don't Exist in High-Dimensional AI
05:22 Gradient-Based vs. Gradient-Free Learning Methods
08:21 AI in Mathematics and Drug Discovery: Opportunities and Challenges
13:48 Scaling Laws and Data Efficiency in Language Models
18:10 Properties of Data that Affect Scaling Laws
22:04 Constructing Non-Redundant Data Sets for Better Learning
24:32 Theory vs. Empirical Results in AI Research
32:19 Fundamental Components of Deep Learning: Are They Changing?
34:31 Future Paradigms in AI Beyond Current Models
37:22 Teaching AI and Humans: Paradigm Shifts in Learning
41:37 Consciousness, Self, and the Brain: Surya's Perspectives
49:49 Neuroscience and AI: Understanding the Brain and Consciousness
01:02:03 Understanding the Brain: Challenges and Opportunities
01:09:21 Brain-Computer Interfaces and AI in Neuroscience
Music
"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.
Text Diffusion Models with Brendan O'Donoghue (Google DeepMind)
2026/08/14
Brendan O'Donoghue, research director at Google DeepMind, makes the case for text diffusion as a real alternative to autoregressive generation. He walks through how discrete diffusion works, why diffusion samples are far more diverse and what that unlocks for RL, where the Gemma diffusion model actually stands against frontier models, and why the whole training and serving stack being hyper-optimized for autoregression is the main thing holding the approach back. The conversation also covers hardware trends favoring flops over bandwidth, AGI timelines and real-world bottlenecks, and why he thinks RL is still underhyped.
Key topics
- Discrete diffusion for text vs autoregressive generation
- Why diffusion samples are more diverse, and what that unlocks for RL
- Where diffusion already wins: latency, on-device, robotics
- Why serving cost, not quality, is the real blocker
- RL as the most underhyped area in AI
Timeline
00:00 Introduction
00:50 What diffusion models are and how text diffusion works
04:40 Why Brendan bet on text diffusion in 2023
07:15 Diversity, creativity, and why it helps RL
11:00 The best diffusion LLM today and the gap to frontier models
14:25 Latency, serving cost, and why it needs more chips
17:14 Where diffusion already wins: on-device, robotics, battery
20:14 One model, two modes: diffusion for thinking, AR for answering
22:24 Samplers and the stuttering problem
26:27 Theory, BERT, and why now is a good time to work on this
31:48 Pipelines built for autoregression, and continuous diffusion
35:35 Hardware: flops vs bandwidth
39:49 AGI timelines and real-world bottlenecks
50:15 Is AI engineering or science?
54:14 Most overhyped and most underhyped ideas
58:35 RL on diffusion, value functions, and exploration
1:07:30 Go download the model and break it
Music
"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.
Nathan Lambert: Inside Post-Training and the Open Model Fight
2026/08/08
Nathan Lambert spent three years as post-training lead at Ai2, where he built the OLMo models, and he writes Interconnects, one of the most-read technical newsletters in AI. He left Ai2 in June and is now working on a new project. He's also the author of the RLHF book. We talked a lot about open models, their capabilities, and why they are better than he expected. We get into what that means over the next two to five years, why he thinks recursive self-improvement is overblown, what the market for training environments actually looks like now, and why he expects Anthropic's famously open internal culture to break after its IPO.
Key Topics
Open vs closed models and who actually captures the valueAnthropic and OpenAI as opposite cultures, and the talent concentration problemBoom vs bubble, and why token spend hasn't produced 10x better productsContinual learning, RSI skepticism, and what Nathan wants to work on nextWhat the open ecosystem needs economically to survive
Timeline
00:00 Intro
00:27 Open vs closed models, and who actually captures the value
05:12 China, harnesses, and where the real training leverage sits
08:40 Sovereign compute and the national security case for building models
11:18 Uncensored open weights and the bioweapon question
14:29 Anthropic vs OpenAI, ideology and politics
19:35 The Mythos ban and the Fable 5 delays
24:30 The AGI narrative, the talent drain, and antitrust
28:12 Why researchers join Anthropic, and the open Slack culture
34:04 Nathan's next 12 months: character training and big RL runs
37:55 Continual learning, RSI, and why Nathan is skeptical
43:19 Boom or bubble, tokens vs GPUs
45:12 Why all that token spend never produced 10x products
48:38 Job displacement and the small-business future
52:49 Robotics, world models, and why multimodal lags
57:44 What the open ecosystem should actually do
1:03:17 Why NVIDIA isn't building a frontier model
1:07:34 The RLHF book, and whether RLHF still matters
1:11:06 GRPO vs PPO and on-policy distillation
Music
"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.
Daphne Koller - The Future of AI in Biology and Drug Discovery
2026/08/03
Daphne Koller wrote the book that many of us learned probabilistic graphical models from, founded Coursera, and now runs insitro, which is trying to make drug discovery a machine-learning problem.
We start with the bitter lesson. She agrees with most of it and then says where it stops working: biology doesn't have enough data, structure is how people understand anything, and making a drug is a question about an intervention that hasn't happened yet, not a pattern in data you already have.
Most of the episode is about why drug discovery is hard. Ninety percent of drugs that reach the clinic fail, and mostly not because the molecule was bad. The molecule usually does what it was designed to do. It just turns out the thing it was designed to do had nothing to do with the disease. Only 22% of diseases have any approved drug at all, and she calls that an upper bound on what we understand, not a lower bound.
She also gets into what agents are and aren't good for in a wet lab, why cells don't grow faster no matter how many GPUs you point at them, what it would take to have real foundation models for biology, and why almost all of biology is still out of distribution.
Plus GLP-1s and what human data keeps teaching us, whether AI can make the kind of leap that turned a bacterial immune system into CRISPR, and what she'd build if she were starting Coursera today.
Key Topics
The impact of scaling and data in machine learningThe importance of structure and causality in AIChallenges in drug discovery and biological understandingThe role of foundation models in biologyEthical considerations in AI and biomedical research
Chapters
00:00 Introduction to Machine Learning and Drug Discovery
02:00 The Bitter Lesson and Its Implications
06:48 Challenges in Drug Design and Discovery
11:48 Ethical Considerations in Human Research
17:20 The Drug Discovery Pipeline Explained
29:30 Integrating AI in Experimental Design
35:38 The Role of Human Judgment in Drug Design
37:14 Future of Drug Design: Efficiency vs. Automation
39:37 Challenges in AI and Data Availability for Biology
41:08 Foundation Models: Potential and Limitations
43:39 Causality in Biological Data: Importance and Challenges
45:18 Creativity vs. Understanding in Drug Design
48:17 Balancing Investments in Data, Algorithms, and Experiments
50:07 The Value of Simulations in Drug Discovery
52:03 Mathematical Frameworks in Biology: Utility and Limitations
54:14 The Future of Drug Discovery: Optimism and Innovations
56:28 The Impact of Coursera on Education
01:00:33 The Role of Universities in Lifelong Learning
01:04:06 Connecting Dots: The Fun of Variety in Work
01:05:46 Optimism for the Future of Drug Discovery
Music
"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.
Podcast sponsorship advertising
Start advertising on The Information Bottleneck relevant audience podcasts
You may also like to advertise on these Podcasts

4.861381091
The David Pakman Show
David Pakman

4.933781199
Tea Time UNFILTERED With Lovelyti
Lovelyti

4.6619502000
The Dan Bongino Show
Cumulus Podcast Network | Dan Bongino

4.921791077
The Carl Jackson Podcast
Salem Podcast Network

4.928911196
Real Life Ghost Stories
Real Life Ghost Stories

4.890581547
Fearless with Jason Whitlock
Blaze Podcast Network

4.827561497
Talkin' Yanks (Yankees Podcast)
Jomboy Media

4.913551952
Wholesaling Inc with Brent Daniels
Find distressed properties for pennies on the dollar and turn them for huge profits!

4.98331800
The Adam Mockler Show
MeidasTouch Network

4.89021968
Catholic Sprouts: Daily Podcast for Catholic Kids
Nancy Bandzuch