The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0.
We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al.
Full show notes always on https://latent.space
www.latent.space
Advertise on Latent Space: The AI Engineer Podcast
Advertise on Latent Space: The AI Engineer Podcast to promote your brand to thousands of podcast listeners
Unlock Latent Space: The AI Engineer Podcast podcast Email contact info, Listeners & Audience details
Email contact information
Direct podcast contact details
Listeners
Audience numbers & engagement insights
Audience details
Podcast Insights
Podcast episodes
Check latest episodes from Latent Space: The AI Engineer Podcast podcast
Academia is for Ambition — Alex Zhang, MIT
2026/10/02
Last call for regular tickets for AI Engineer NYC! As an exclusive for Latent Space subscribers, the first 30 of you can take a 30% off code if it helps - for new tickets only, no refunds! See you in 2 weeks!
While we tend to cover industry on the pod, every so often we celebrate a clearly emerging superstar PhD. In 2024 we featured Shunyu Yao, who went on to build Operator at OpenAI and is now Chief AI Scientist of Tencent. In 2025 we featured Jack Morris, who went on to cofound Engram at $600m and is now a leading voice on continual learning.
This year we are proud to feature the work of Alex Zhang of MIT.
From GPU kernels and KernelBench to Recursive Language Models, Mismanaged Geniuses, and massive multi-agent swarms, Alex Zhang is exploring how much capability we’re leaving on the table by wrapping increasingly powerful models in primitive systems.
RLMs took over the timeline early this year:
and an RLM based harness was the first to ~solve ARC-AGI-3 before OpenAI’s Astra:
and is even today, influencing new research that has more extreme implications than RLMs:
We go deep on GPU Mode and AI-written kernels, research taste and why academics should take bets industry labs won’t, GEV and alternatives to the standard autoregressive language model, and the idea of harnesses as compositional generalizers. Alex explains RLMs, context offloading, programmatic subagent calling, Prime Agent, persistent subagents, and why the “language model” of the future may actually be an invisible swarm of agents underneath a simple interface. We also discuss OpenAI’s massive agent experiments, Kimi swarms, open-ended research at Sakana AI, speculative programmatic tool calling, capability overhang, Neuralese, and where Alex thinks the next big research opportunities may lie.
We discuss:
* Why AI-generated GPU kernels still leave substantial room for human expertise
* How one expert insight can potentially replace enormous amounts of brute-force token search
* Why PhD students should take research bets that initially look trivial, weird, or pointless
* What SWE-bench, RLMs, ReAct, and Quiet-STaR reveal about research taste
* GEV and why a language model does not have to mean an autoregressive text-to-text decoder
* Why Claude Code, Codex, and Pi are structurally more similar than they look
* How harness design can improve compositional generalization across tasks and domains
* RLMs: context offloading, code execution, recursive subagents, and shared memory
* Prime Agent, continual harnesses, and persistent agent-to-agent communication
* Why the model you query in the future may secretly be an entire swarm or scaffold
* OpenAI’s 10,000-agent experiment, 130B output tokens, and ~$40M-equivalent problem solving
* Why much of an agent swarm may be wasted search — and why convergence is still hard
* Kimi versus OpenAI and different approaches to multi-agent systems
* Open-endedness, Sakana AI, and finding hidden gems in enormous amounts of generated work
* Why current frontier models may already have a large capability overhang
* Speculative programmatic tool calling and overlapping tool execution with generation
* Whether English, code, or an entirely new “Neuralese” constrains how models reason
* AI for science, fast-moving benchmarks, and how Alex chooses what research problems to bet on
Alex Zhang
* Website: alexzhang13.github.io
* X: @a1zhang
Timestamps
00:00:00 Introduction
00:00:49 GPU Mode, KernelBench, and AI-Written Kernels
00:07:38 Human Expertise vs. Brute-Force AI Search
00:13:20 Research Taste and Taking Big Bets
00:19:28 GEV and Rethinking the Language Model
00:29:03 Video Game Agents and the Harness Problem
00:31:01 Why Claude Code, Codex, and Pi Are So Similar
00:36:42 Harnesses as Compositional Generalizers
00:44:24 RLMs Explained
00:52:01 Prime Agent and Persistent Subagents
00:57:41 RLMs in the Wild
01:00:30 OpenAI Swarms and the Future of Language Models
01:07:26 Open-Endedness and Sakana AI
01:15:52 Kimi vs. OpenAI Agent Swarms
01:20:06 Capability Overhang and Speculative Tool Calling
01:28:19 Neuralese, Future Research, and AI for Science
Transcript
Introduction: Alex Zhang, RLMs, and GPU Mode
Swyx [00:00:00]: All right, we’re here in the studio with Alex Zhang, I guess most famously of, RLMs, but you have a few other affiliations. Welcome to the show.
Alex Zhang [00:00:12]: Yeah. Thank you for having me.
Swyx [00:00:13]: Yeah. I guess GPU Mode as well?
Alex Zhang [00:00:15]: Yes, GPU Mode as well.
Swyx [00:00:16]: You were shepherded in by Mark Saroufim. Not everyone gets that kind of welcome.
Alex Zhang [00:00:19]: Yep. Yeah. Yeah. I’m very close to all the people in GPU Mode, so yeah
Swyx [00:00:23]: Yeah
Alex Zhang [00:00:23]: We often end up working together in various capacities, like even beyond just GPU Mode itself, so.
Swyx [00:00:29]: Yeah. Can we explain, so people who are not that close Don’t know about this. It’s just a-- it’s a Discord. It used to be focused on, I guess, CUDA Mode, and then generalized a little bit. it was started by Mark.
Alex Zhang [00:00:41]: Yep.
Swyx [00:00:42]: It was basically like. To me, it’s like the hiring pipeline of the PyTorch team.
Swyx [00:00:45]: And then you left PyTorch.
Alex Zhang [00:00:47]: Yep.
Alex Zhang [00:00:49]: Yeah. So it used to be, I think, it started actually around when I was in college, in like 2023. I think it was started by Mark, Andreas, and Jeremy Howard. The original premise was just like, it was a GPU-- or it was a Discord dedicated to learning how to write GPU kernels, and they had, like, lectures. That was basically the extent of it. and I got interested in it because I was writing GPU kernels. It was actually out of, like, pure chance. I was interning at Snapchat at the time, and
From CUDA Mode to GPU Mode
Swyx [00:01:22]: Rexis.
Alex Zhang [00:01:23]: Yeah, I was very bored with Rexis. So, they had a project where, like, they were interested in writing. It was this paper called Infinite Attention. It was like a Google paper.
Swyx [00:01:35]: Yes, we’ve covered it on Paper Club.
Alex Zhang [00:01:36]: Yes, yeah. So I was interested in whether or not you could write specialized kernels for it at Snapchat. it didn’t. Nothing really came of it, but I joined GPU Mode. At the time, it was called CUDA Mode, I think for, like, legal reasons or something, they changed the name. But I met Mark, I met Matei, I met a bunch of other people that were very involved in the community. And then Mark had pitched this idea called Popcorn, which was now what you see as the leaderboard today. But the general idea was like. I think all of us had this, like, intuition that GPU programming is, like, very similar to if you guys have done, like, competitive programming. It’s a, it’s a not. I don’t mean to say, like, they’re transferable skills.
Popcorn, KernelBench, and Automating GPU Kernels
Swyx [00:02:17]: You have constraints. You code golf a little bit.
Alex Zhang [00:02:19]: Yep.
Swyx [00:02:19]: Yeah.
Alex Zhang [00:02:19]: Yeah. And there’s like. There’s actually a surprisingly small space of optimizations that people do. and there’s actually not that many kernels per se that people are interested in optimizing. And so we kind of had this thought that, like, if you had enough data, like, in the same way that Codeforces, there’s millions of problems. If you could do this with GPU code, like, you could scale and automate kind of GPU kernel development, which for researchers is a huge deal. ‘Cause I think one of the bigger bottlenecks. Like, if you look at, like, Mamba, for example, like, they release the paper with kernels because otherwise, like, you can’t really use it in any meaningful way? And not everyone has, like, a Tri Dao on their team. So we’re very interested in this. KernelBench kind of spawned from that too, of like, can we get LLMs to automate, GPU kernel code? And I think that was like a. It was a very fun time. It was like between college and my PhD, and yeah, I had a really pleasant time doing stuff with GPU Mode. Now I kind of just help with the lectures sometimes. I’m not as involved, and I think in general, like, we don’t have as many competitions as we used to. But, yeah, I still keep in touch a lot with everyone there.
Swyx [00:03:27]: Is there a friendly rivalry? Because I think the previous community that used to do this was like MLSys, MLPerf
Alex Zhang [00:03:32]: Yeah
Swyx [00:03:32]: Kind of thing. Is there a friendly rivalry? Is this like just new generation MLPerf, or what’s going on?
Alex Zhang [00:03:38]: The nice thing about GPU Mode is that it is also a community in the sense that, like, a lot of the lectures are very easy enough for a beginner to follow and ask questions and things like that. And, like, the competitions are, like, somewhat, not secondary, but, like, you can participate in them to learn. I think with a lot of, like, MLSys, MLPerf kind of benchmarks, like, for the most part, like, only serious labs and companies participate. Like, seriously in them, at least. That was my understanding of it. I could be wrong. But I think also beyond GPU Mode now, one thing that has been really exciting is there’s a lot more websites and, like, people that work on hosting competitions. Like, I think there’s this.
Alex Zhang [00:04:21]: I think there’s this website called, like, LeetGPU or something, and it’s like leet code for GPU problems.
Swyx [00:04:26]: Wow.
Alex Zhang [00:04:26]: There’s, like, other ones that I’ve. Like, we’ve, we’ve seen. Like, there’s many that have kind of spawned and, like, talked on GPU Mode, and like, it’s very exciting in the sense that I think GPU programming used to be super niche, like when I was interested in it. And the only reason I got interested in it was Tri Dao gave a talk at Princeton because he was, applying for faculty. he is faculty there now, but I listened to his talk on FlashAttention in li
Why Dwarkesh is Wrong about Computer Use + How OpenAI shipped its Jev competitor in 1 Week
2026/09/30
Three months ago Dwarkesh, who has been posting incredible blogs and episodes about RL, posted a framing question for his video essay on RLVR which upset a lot of Computer Use folks:
We are no strangers to learning in public and are no strangers to the stress of getting things wrong when you have a big platform. However, we were at Anthropic for the Computer Use launch, there for Claude Cowork with the first big podcast on it, organized the first Computer Use track at AIE presenting the state of the art, and were close to the OpenAI-Sky Software acquisition that now powers the complete domination of computer use that Codex enjoys today. This is why we’re excited to bring you today’s first guest, Ari Weinstein, cofounder of Sky and now leading all the amazing CUA progress that casuals might miss:
Ari explains why Computer Use is now “180 degrees different” from where it was months ago, how agents are learning to debug and recover from failures, why combining screenshots with accessibility data, the DOM, Playwright, and generated code changes the speed equation, and why the next frontier is making agents literally superhuman at using software.
OpenAI clones Jev
In the second half, Nikunj Handa from OpenAI’s API team breaks down the new developer stack: async tool calling, mid-turn steering, WebSockets, UltraFast inference, the Decisions API, prompt caching, pre-warming, compaction, and the Agents API. Given that we were the first Jev podcast, we particularly focus on the unusually fast sprint on the Decisions API:
And why it is just a Luna wrapper for now but the team is motivated and egoless enough to clone what they consider to be good patterns.
We discuss:
* Why OpenAI thinks Computer Use has changed dramatically in just the last few months
* Dots and what changes when every agent gets its own Linux computer
* Why Computer Use can now complete some tasks faster than the average human
* The path from human-level to “literally superhuman” computer use
* Why modern agents are much better at debugging and recovering from failure
* How screenshots, accessibility trees, the DOM, Playwright, and generated JavaScript work together
* App Shots and why they give models much richer context than ordinary screenshots
* Why Computer Use can close the loop between writing software and testing it
* Trust, permissions, and safety when agents can make payments and operate websites
* Async function calling and why models no longer need to stop reasoning while tools run
* Mid-turn steering, WebSockets, and the architecture behind more responsive agents
* UltraFast inference and how OpenAI is pushing frontier models toward much lower latency
* The rapid internal story behind the Decisions API
* Why Decisions API is more than structured outputs at low latency
* GPT Live, fast tool calling, and real-time computer control
* How OpenAI is already using Decisions API for support classification and internal workflows
* Longer prompt caching, cache pre-warming, and cache-aware applications
* Server-side compaction vs manual compaction for long-running agent threads
* What should live inside an Agents API versus a developer’s own harness
* OpenAI as an “AI cloud” and the search for higher-level primitives beyond raw model APIs
Ari Weinstein
* Product & Engineering, Computer Use at OpenAI
* X: https://x.com/AriX
* LinkedIn: https://www.linkedin.com/in/weinsteinari/
Nikunj Handa
* Product, API at OpenAI
* X: https://x.com/nikunjhanda
* LinkedIn: https://www.linkedin.com/in/nikunjhanda/
Timestamps
00:00:00 OpenAI DevDay: Dots, GPT-6.1, Agents API, and Decisions API
00:02:52 Dots and Personal Cloud Computers
00:04:59 Why Computer Use Is “180 Degrees Different”
00:06:04 From Sky to Self-Debugging Computer Use Agents
00:09:24 How Computer Use Sees and Operates Software
00:12:09 From Faster Than Humans to Superhuman Computer Use
00:16:03 Agents API: Trust, Permissions, and Safety
00:17:31 Computer Use for Coding, Testing, and QA
00:19:14 GPT-6 APIs, Async Tool Calling, and UltraFast Inference
00:23:21 The Rapid Story Behind Decisions API
00:25:32 What Decisions API Is and How It Works
00:30:24 What OpenAI Is Building With the New APIs
00:32:23 Prompt Caching, Pre-Warming, and API Performance
00:35:20 Context Compaction for Long-Running Agents
00:37:13 Memory, Higher-Level APIs, and the AI Cloud
Transcript
Introduction: OpenAI DevDay and the New Agent Stack
Vibhu [00:00:00]: Okay. We’re very excited to be here. Today is OpenAI DevDay. Special podcast
Swyx [00:00:08]: We’re the first podcast after your livestream.
Vibhu [00:00:10]: First podcast. We have Ari here, who leads the product and engineering team for Computer Use agents. Before we kick in and dive deep on Computer Use, you wanna give a quick recap? What was announced? What’s the quick slew of announcements you guys had today?
Ari Weinstein [00:00:24]: Yeah. yeah, it was a super exciting day. we just got out of the keynote. It was really sick. there were a bunch of Computer Use announcements that I think are worth thinking about. We have, Dots, which is the new, sort of personal assistant product, and, that has some really exciting Computer Use features. There’s GPT-6.1 Sol, which is this amazing new model, that I think is particularly great for Computer Use ‘cause of, sort of the cost and speed, advantages. I think, I think we shared that it’s, a fifth of the cost of Astra and a seventh of the cost if you’re looking at Computer Use specifically, which is really amazing. sorry, there were so many things. I’m trying to sort through it.
Swyx [00:01:02]: And the API.
Ari Weinstein [00:01:03]: Agents API, which now has Computer Use in it, which is really cool, ‘cause now developers can build on the same Computer Use, that is part of Codex, and ChatGPT. and then there were some demos of our existing Computer Use features, like app shots, where you can take the context of something you’re doing on your computer and bring it into Codex and ChatGPT really fast. And then, like, native Computer Use on your Mac, where Roman had it taking screenshots of his app, automatically, and he could do other things on his computer while Computer Use was using his applications. so yeah, really exciting keynote.
Swyx [00:01:35]: And not to mention the Decisions API.
Ari Weinstein [00:01:37]: Decisions API.
Swyx [00:01:38]: Off the bat, are they all the same model? Like, this is. Or the same dataset distilled to different models?
Swyx [00:01:44]: Like, basically, like, is Computer Use using Decisions API, or are they, like, kinda separate?
Ari Weinstein [00:01:49]: So what’s really cool about the Decisions API is it, you know, it has all these new capabilities. It does inference in parallel. it doesn’t have reasoning. It’s a smaller model, than the ones we use for Computer Use. and so those capabilities make it really fast.
Dots and Delegating Work to a Cloud Computer
Swyx [00:02:07]: Yeah.
Ari Weinstein [00:02:07]: They also make it a little bit less good at doing, like, long horizon, sort of sophisticated tasks. And so I think I would say it’s still an open area of research for how we, like, bring those approaches together. But, yeah, I’m really excited to see what people build with the Decisions API.
Vibhu [00:02:24]: One of the interesting things is Dots now have attached personal computers.
Ari Weinstein [00:02:28]: Yeah.
Vibhu [00:02:28]: So it seems like they’re very much more persistent. You’ve been using them for a while. How should people push the bounds? Like, what should people aim for? What should they try? Personally, right now I use it for a lot of customer service. Like
Ari Weinstein [00:02:41]: Cool
Vibhu [00:02:41]: “Oh, this was wrong. I don’t wanna sign in. I don’t wanna authenticate.” Find whatever and just get it fixed.
Ari Weinstein [00:02:45]: Yeah.
Vibhu [00:02:46]: How should we push further? What should people try?
Ari Weinstein [00:02:50]: Dots Are a really cool product because each Dot has access to its own Linux virtual computer in the cloud, which is different from our other products. you know, traditionally, we’ve have access to a browser in the cloud, or it has access to your own computer, but now you get your own entire Linux computer in the cloud. And so it can run full desktop applications, and it can also use a web browser. And so, yeah, you know, I think the powerful thing about Computer Use and the reason why I think it’s so, exciting is because it makes it so that the agent can do anything you as a, as a person can do, because all the software in the world was designed for humans, and now agents can use that same software, and you can delegate to the agent. So, yeah, like, anything that you would do on a computer, you can ask a Dot to do. Yeah, I think what particularly is useful is gonna really depend on who the end user is and what- what’s valuable in their life. but yeah, I would just start by thinking about, like, one of the things that you spend time on and how could you delegate those to an agent.
Swyx [00:03:47]: Yeah, a lot of flight booking and shopping and honestly even, like, playing a game or whatever, right?
Ari Weinstein [00:03:52]: Totally.
Swyx [00:03:52]: Yeah.
Ari Weinstein [00:03:53]: Yeah, I don’t know. For me, something I did recently, I’ve been working on. I’ve, subscribed to a meal prep service ‘cause I was trying to, like, eat healthy, you know? And I really like this meal prep service I found because it lets me customize the meals I order to, like, a high degree of granularity. So I can say, like, “I want this many grams of chicken and this many grams of rice.” but it was so complicated. It took me two hours to do an order, and I found that I could ask Computer Use to do it for me, and it did it in 15 minutes. so I actually saved two hours. it both did it eight times faster than I could, and it saved me two hours on GPT-6.1 Sol.
Swyx [00:04:32]: Yeah.
Ari Wei
Claude Code’s Next Era — Thariq Shihipar, Anthropic
2026/09/29
We are excited to have Anthropic share their latest AI x Finance work at AI Engineer New York, coming up in 2 weeks!
In case you’ve been under a rock, here’s a non-exhaustive list of what Anthropic has been shipping since closing the largest fundraise of all time in May at $47B ARR:
* June: Launched Claude Tag and Sonnet 5 and Fable 5
* July: Opus 5, /checkup. crossed $65B ARR
* Last month: Fable/Mythos 5.1, and EFS (upcoming pod)
* IPO target $2T, end 2026 ARR estimated $100B
* Cowork/chat merged before did
* Claude Mods
* Dario endorses the same Pacing the Frontier message cosigned by all labs
* Last week: Opus 5.5, Plugins portal, Cloud Sessions/Claude Projects
* Today: Sonnet 5.5!
Today’s episode should catch you up, with Thariq Shihipar, the explainer-king of Anthropic, who we last caught up on Fable launch day with The Field Guide to Fable:
The Future of Mutable Software
Pay special attention to Claude Mods (especially the cheatsheet):
In general this is also the inverse of the other viral tweet from Thariq:
Cloud Brain, Local Hands
And give a try to Claude Projects:
The “hands” terminology is not just an analogy for the local/cloud paradigm that is being built up at frontier coding agent companies like Cognition, but is ALSO particularly relevant to the safety systems discussions that we’ll be discussing with Anthropic in an upcoming episode as they prepare to pace to frontier with responsible AI deployment.
For those who want Thariq’s writing tips we teased at the start of the pod, watch the full video here:
From the rapid rise of Claude Code to a future where agents can rewrite their own harnesses, collaborate across teams, and operate across cloud and local environments, the way we build software is changing extraordinarily fast. In this episode, Anthropic’s Thariq Shihipar joins swyx and Vibhu to unpack how power users are actually working with Claude Code today, why prompting remains a high-skill discipline, and where Anthropic thinks the agent harness is headed next.
We go deep on Claude Code’s evolving interface: Ask User Question and elicitation, artifacts as persistent generative interfaces, Claude Tag for multiplayer agent workflows, Projects, model effort, implementation notes, and the new Claude Mods system for customizing the harness itself. Thariq explains why Claude.md may eventually disappear, why the smartest model could also become the cheapest model for many tasks, and why mutable software could become a new paradigm for how applications are built and customized.
The conversation then turns to agent security and Anthropic’s “Pacing the Frontier” argument. Thariq walks through recent incidents where agents discovered unexpected ways to communicate, exploit infrastructure, reverse-engineer benchmark scorers, and chain vulnerabilities together. We discuss sandboxing, prompt injection, autonomous agents, interpretability, constitutional classifiers, probes, fallbacks, Auto Mode, and why securing increasingly capable agents may become one of the defining engineering problems of the next few years.
We discuss:
* Why agentic coding went from controversial to the default in less than a year
* Why prompting is still one of the highest-leverage skills for working with Claude Code
* How expert users build a mental model of Claude and what it can reliably one-shot
* Why discovering your “unknown unknowns” matters more as agents become more capable
* Artifacts as persistent, generative interfaces between humans and agents
* How Claude could split into a cloud-based “brain,” local or remote “hands,” and dynamic interfaces
* Claude Tag, Projects, and multiplayer agents and how collaborative agent workflows could evolve
* Why spending more time on the initial prompt can dramatically reduce wasted agent work
* When to use low, medium, high, or max effort for different engineering tasks
* Why frontier models may eventually outperform smaller models on both intelligence and token efficiency
* Why implementation notes can expose decisions the model considered but chose not to make
* Why Claude.md may eventually disappear — and why starting without one can sometimes be better
* Claude Mods: customizing the execution loop, UI, subagents, routing, and behavior of Claude Code
* Model routers, forked agents, and supervisor agents that automatically improve agent workflows
* Why Claude Mods may be an early preview of “mutable software”
* The bitter lesson of harness engineering and why agent architectures go out of date so quickly
* How Claude Tag is becoming an organizational harness for multiplayer work
* Why giving agents access to company data creates an enormous new security surface
* The Exploit-Bench incident where agents discovered ways to communicate and collaborate
* Why agents hacked Hugging Face for scorer code rather than benchmark answers
* How agents chained sandbox and infrastructure vulnerabilities in unexpected ways
* Why increasingly capable agents make traditional security assumptions harder to maintain
* The argument behind Anthropic’s “Pacing the Frontier” proposal
* Why software engineers are increasingly doing two jobs: engineering and keeping up with AI
* Constitutional classifiers, probes, and fallbacks and what interpretability looks like in production
* How Auto Mode checks whether an agent’s actions actually match the user’s permissions
* Why Thariq can see serious AI risks while still having a relatively low p(doom)
Thariq Shihipar
* X: https://x.com/trq212
* LinkedIn: https://www.linkedin.com/in/thariqshihipar
Timestamps
00:00:00 Introduction
00:04:12 Ask User Question and the Future of Agent Interfaces
00:08:29 Artifacts, Projects, and Multiplayer Agents
00:15:37 Prompting as the Core Claude Code Skill
00:21:52 Context, Effort, and Smarter Model Usage
00:28:10 Is Claude.md Going Away?
00:32:49 Claude Mods: Customizing the Claude Code Harness
00:36:35 Model Routing and the Rise of Mutable Software
00:44:40 The Bitter Lesson of Harness Engineering
00:50:49 Claude Tag as an Organizational Harness
00:55:59 Pacing the Frontier and Autonomous Agent Security
00:58:22 Agents Hack Hugging Face for the Scorer
01:05:34 What Happens When Agents Need More Compute?
01:10:32 AI Coding Is Changing Faster Than Engineers Can Keep Up
01:17:17 Probes, Fallbacks, Interpretability, and Auto Mode
01:28:32 AI Risk, p(doom), and Closing Thoughts
Transcript
Introduction: Life at Anthropic and the Pace of Change
Swyx [00:00:00]: We’re here in the studio with our friend Thariq from Anthropic, and I guess generally the Claude Code, I-- there’s, there’s so much, merging of boundaries and you’ve been so on top of everything since you joined Anthropic. You have been early to Claude Code itself, but then also, and you’ve told that story in other podcasts, and you’ve also been talking about seeing like an agent. Most recently you did the top AIE World Tour talk, Field Guide to Fable, which obviously you guys launched Fable, so that was-- that’s cheating. And mostly you most recently also launching Claude Tag, and we’re also gonna be talking about Pacing the Frontier. There’s a lot going on in Anthropic. I guess top of the question is, what’s it like being at Anthropic when there’s so much going on?
Thariq Shihipar [00:00:48]: I think that It is, like. I think you can get whiplash sometimes. I think, like, going. When I joined Anthropic, I joined because of Claude Code. Like Claude Code had just come out and I was like, “This is so good.” And Opus 4 to me was like just, I could not imagine, like, how good it was? And that was, like, a real moment for me. But I was, like, trying to convince, like, my startup friends to use agentic coding, and they’re like, “Oh, no, like, our engineers don’t think it’s good enough,” or something. And I was like, “That’s insane.” and now you, like, fast-forward, 12 months, less, and, like, it’s just like, yeah, the default way that everyone codes, right? And I think that, like, just having to go from, like, selling it to, like, now, teaching people how to be. make the most use of it and be more efficient and things like that is just like a big, like big change. And, yeah, I think, like, it’s just hard to stay on top of everything as a human? Like, I think things happen so fast and like
Swyx [00:01:51]: You just throw more agents at it.
Thariq Shihipar [00:01:52]: Yeah, like that’s like the agentic stuff scales much better than the, like, human stuff where it’s like, oh, like, there are three things happening right now and, like, they’re all emergencies and, like, how do you, like, respond to it? Yeah.
Teaching People to Use Claude Code
Vibhu [00:02:05]: What do you split your time on? You do a lot of technical writing, engineering work.
Thariq Shihipar [00:02:10]: Yeah, so I think that, like, when I joined the Claude Code team, I wanted to teach people how to use Claude Code and I think that, like, that has been something that, like, I thought, like, maybe I would spend a little bit of time on it or, like, I’d, like, do. I was spending some time on the agent SDK first, and I wasn’t exactly sure, like, how the bitter lesson would go, when it comes to, like, harnesses, right? Like, I think sometimes we were like, “Oh, like, what’s after Claude Code?”? And so initially I was like, I just wanna teach people how to use Claude Code and make it easier to use Claude Code. And I think that has just, like, as the harnesses have gotten better and better, that’s like the dominant problem now is, like, how do you use the agents, right? Like, it’s like such a high skill expression thing. So I do that and then I do engineering work. I give talks, but I think, like, when I’m doing engineering work, my goal is to take that feedback that we get from users and also, like, then be able to talk about, like, hey, how to use Claude Code to do engineering. So there’s like a good l
OpenRouter: from Seed to Stripe — with OpenRouter’s Alex Atallah & AMP’s Anjney Midha
2026/09/25
From the earliest days of open-weight models to becoming the neutral routing layer for more than 10 million developers, OpenRouter is one of the clearest bets that the future of AI will be multi-model. In this episode, OpenRouter co-founder & CEO Alex Atallah, with AMP’s Anjney Midha returning with swyx to unpack how OpenRouter emerged from the first wave of Llama, Alpaca, Mistral, and Midjourney, why model diversity mattered before it was consensus, and how a company dismissed as “just a wrapper” became critical infrastructure for the AI ecosystem.
We go deep on the product and distribution lessons behind OpenRouter: why model labs can spend billions training a checkpoint and still struggle to get it into developers’ hands, how Mistral helped prove the value of a competitive inference marketplace, why OpenRouter chose focus over expanding into fine-tuning, memory, and other adjacent products, and how its rankings became a real-time map of how AI usage was changing. Alex also explains OpenRouter’s early experiments with model fusion, why they deleted the first version and brought it back years later, and how the platform grew to more than 10 trillion tokens per day.
Finally, Anjney explains why Stripe and OpenRouter fit together, why token fraud may become one of the defining security problems of the AI economy, and why the next wave of fraud won’t just come from humans but from autonomous agents attacking increasingly valuable token flows.
We discuss:
* Why OpenRouter bet early that no single AI model would win everything
* Alpaca, Llama, and open models becoming impossible to ignore
* Why Discord’s early AI deployments exposed the limitations of closed models
* Why model labs can spend billions on training and still fail at distribution
* How OpenRouter became a neutral distribution layer for model developers
* Why VCs dismissed OpenRouter as “just a marketplace” or “just a wrapper”
* The Mistral price war and the first real proof of an inference marketplace
* How Midjourney scaled through Discord and what it taught the AI ecosystem
* Why crypto infrastructure became a dress rehearsal for generative AI
* OpenRouter vs. LM Arena and why their missions are fundamentally different
* Why focus became one of OpenRouter’s biggest strategic advantages
* Anthropic’s early focus on AI pair programming and coding
* The OpenRouter products that were prototyped but never launched
* MOM, OpenRouter’s early Mixture of Models experiment
* Why model fusion failed in 2024 — and why it works much better now
* How OpenRouter’s leaderboard became a live map of the AI industry
* OpenClaw, auto-routing, and agents reshaping AI usage
* How OpenRouter reached 10+ trillion tokens per day
* Why inference gateways are increasingly becoming targets for fraud
* Why Stripe’s fraud infrastructure is strategically important to OpenRouter
* The coming rise of agentic fraud and attacks on the token economy
* What changes and what stays the same as OpenRouter joins Stripe
Alex Atallah
* LinkedIn: https://www.linkedin.com/in/alexatallah/
* X: https://x.com/alexatallah
* Website: https://alexatallah.com
Anjney Midha
* LinkedIn: https://www.linkedin.com/in/anjney/
* X: https://x.com/AnjneyMidha
* AMP: https://www.amppublic.com/
Timestamps
00:00:00 Introduction
00:02:12 Alpaca, Llama, and the Multi-Model Bet
00:06:04 Discord, Open Models, and OpenRouter’s Origins
00:14:28 Why “One Model Wins” Was the Wrong Bet
00:17:27 Why Model Labs Struggle With Distribution
00:23:04 “Just a Wrapper”: Why VCs Misunderstood OpenRouter
00:27:58 Bootstrapping OpenRouter Through Community
00:36:16 Crypto, Midjourney, and the Early Generative AI Ecosystem
00:43:38 Mistral and the Birth of the Inference Marketplace
00:47:10 OpenRouter vs. LM Arena
00:52:08 Focus, Anthropic, and Roads Not Taken
00:59:34 Mixture of Models and Model Fusion
01:02:44 Sonnet, OpenClaw, and OpenRouter’s Explosive Growth
01:09:03 Why Stripe Acquired OpenRouter
01:12:45 Fraud and the Emerging Token Economy
01:17:47 The Coming Wave of Agentic Fraud
01:19:07 What’s Next for OpenRouter at Stripe
Transcript
Introduction: OpenRouter, Marketplaces, and Pub-Sub as a Product Principle
Swyx [00:00:00]: Okay, we are here in Anja’s house, which is where all big startups in San Francisco start.
Anjney Midha [00:00:08]: Howdy.
Swyx [00:00:08]: And, congrats on Cursor, Mistral. I don’- God knows what else. You got so much stuff going on.
Anjney Midha [00:00:17]: There’s, there’s a lot going on. Well, OpenRouter is probably the - has been the most, I would say, like, one I’m excited about recently.
Swyx [00:00:24]: Yeah. And we have Alex, first time on the pod, but,
Anjney Midha [00:00:27]: Thanks for having me.
Swyx [00:00:27]: You’ve been in the IE a few times. I appreciate every time you’ve shown up, for the community. Congrats. I just, like, what a journey. When I was looking back at your past posts, one of the earliest principles that I saw you write as a product person is sub as a product principle. And I wanted - you to maybe explain how you think about what should exist in the world.
Anjney Midha [00:00:49]: Yeah. The sub piece, which was early 2023, I didn’t think about it until we talked like 10 minutes ago, is about how there is like a way of thinking about products as an intersection between subscribing to data and publishing data. And marketplaces are an easy example of this. You have suppliers that are publishing some product to a SKU. And the SKU is like a sub topic that a consumer is subscribing to and just going to, like, consume whenever they want. And humans consume in a very, like, discreet, ad hoc way. It’s not very scalable. all their attention is on the topic when they’re buying the thing, and their attention is nowhere else when that happens. agents and consumers of inference don’t act like that. They’re consuming continuously, and they’re changing the SKUs that they consume from all the time. So OpenRouter is like a blend between a normal API experience and a marketplace where we create model slug. We have the auto router. We have all kinds of, like, product SKUs that you can subscribe to. And then you can, like, continuously add, like, derive value and make decisions based on those consumers.
Alpaca, Llama, and the Multi-Model Bet
Swyx [00:02:11]: Yeah. This is something that was more consensus now, but not consensus when you guys started, which was that there is such a demand for swapping models and changing things out and, that people would not use the native SDKs. I guess, for each of you, what was your realization moment that this would be it? I, - You’ve, you’ve given a talk at EIE about Alpaca as,
Anjney Midha [00:02:33]: Yeah.
Swyx [00:02:33]: One of your inspiring moments.
Anjney Midha [00:02:35]: Alpaca, I can, like, rehash the Alpaca moment for a sec. Like, the very beginning, at the end of 2022, OpenAI was the only game in town. There was, like, OpenAI, Cohere,
Swyx [00:02:47]: Yes.
Anjney Midha [00:02:48]: And then a smattering of, like, early attempts at open weight models.
Swyx [00:02:54]: Yeah.
Anjney Midha [00:02:54]: When Llama came out in January of 2023, it was like, “Wow, really exciting. This is really big.” It outperforms 3 on, one or two benchmarks. but you can’t chat with it. It wasn’t like - It wasn’t an engaging model, but it seemed like someone just needed to fix a couple things and do some RLHF on it to get it all the way there. And Alpaca was the first model that I saw that did that. It only took $600 to do. A team at Stanford generated a bunch of synthetic data, tuned Llama, and made Alpaca, billion parameter model. Or was - Maybe it was thirteen billion parameters. And it was so good. Like, I was just, like, on an airplane using it. I, - in many cases, I, like, you could not discern a ChatGPT versus an Alpaca result. And I figured if it was this easy to make a model, one, we have a whole new way of monetizing data for the first time. you can just, like, take really valuable data and turn it into a service in $600. and that cost will probably go down over time.
Swyx [00:04:03]: When you - So sorry. when you say monetizing your data as, what eventually will become an MCP endpoint or as a training data for a model?
Anjney Midha [00:04:12]: Yeah, training data for a model.
Swyx [00:04:13]: Awesome.
Anjney Midha [00:04:13]: Like, an abstract way of saying like, “Hey, I have this data.”
Swyx [00:04:15]: Compress it into a model.
Anjney Midha [00:04:16]: Like, it makes sense for me in my product, but, like, I could repackage it in the form of a model and sell it. And so it’s just a whole new business model for the economy. It also, of course, provides, like, a way of following what Frontier Labs are doing, but in a way that, like, a single developer or a small team of developers can roll on their own. And so - Whenever you have an example of that, like a breakout app that’s doing really well, and then some framework for imitating it with - in your own flavor, you have an immediate ecosystem of, like an immediate ecosystem, like, should arise because there’s just a huge gap between the, like, decisions that the single company is making and all of the variations in those decisions that, like, a wider ecosystem can create themselves. And so then, you need a marketplace to, like, discover all of those, services and all of those products. There wasn’t any place on the internet that, like, was like a home base for LLMs in terms of seeing how much they were being used and seeing who was using them and why.
Swyx [00:05:29]: The closest would be Hugging Face.
Anjney Midha [00:05:30]: Hugging Face was the closest at the time, yeah.
Swyx [00:05:31]: They just started Hugging, like, a few years ago before that.
Anjney Midha [00:05:34]: Yeah, and Hugging Face also didn’t have the closed-source models.
Swyx [00:05:37]: Yeah.
Anjney Midha [00:05:38]: And they didn’-
Runway’s WorldPrompt and the Engineering of Real-Time Worlds
2026/09/25
Earlier this month, world model company Runway introduced GWM Worlds 2, a research preview that “turns high-fidelity video and audio generation into real-time interactive simulation.” Runway calls this an “autoregressive diffusion” model; with autoregressive describing how it generates over time.
One new feature in particular caught our eye: WorldPrompt, a proposed input format for specifying a generated world and the actions within it. It allows you to fix some aspects of a simulated environment — including the first frame — and then create a series of timestamped events. The events, or actions, can even be prompted in real-time.
To understand the implications of WorldPrompt, we spoke to Kamil Sindi, Runway’s CTO, and Robin Kahlow, its Principal Research Scientist for generative video and multimodal AI. We also have exclusive comments from Anastasis Germanidis, co-founder & co-CEO of Runway, courtesy of a podcast swyx and Vibhu did with him.
Who’s building real-time interactive world models?
First, some context about world models that can generate interactive video and audio in real-time.
Runway is reportedly valued at $5.3 billion, based on its most recent fund raise of $315 million in February. Its first release, GWM Worlds, was launched last December.
Alongside Runway, there are several other notable projects in this domain: Google DeepMind’s Genie 3 (which also generates at 720p and 24 fps), Odyssey-2 Pro, and World Labs’ RTFM (Real-Time Frame Model). We’ve summarized their differences in the following table:
Given the complexity and massive latency demands of real-time video and audio generation (which we’ll get into below), all of the projects listed above have limitations. For instance, Google notes that Genie 3 “can currently support a few minutes of continuous interaction, rather than extended hours.”
But as our interviews with Runway show, real progress is being made.
The central idea of WorldPrompt
WorldPrompt, a new feature in GWM Worlds 2, helps differentiate Runway from its competition. You can think of it as a control layer for characters, cameras and the environment. As Kahlow put it, it’s a way to “control all the different subjects in the world” — similar to a computer game.
“Like, if there’s an NPC [Non-Player Character] somewhere, the NPC might walk up to you and say something. So you could achieve the same thing with this kind of model, where you can have very detailed control over everything in the scene.”
As the name suggests, WorldPrompt is a prompting mechanism — not a programming language. So, unlike virtual world games like Minecraft or Roblox, GWM Worlds 2 doesn’t offer scripting capabilities or the ability to control state. But there’s a power to that, as Sindi pointed out.
“You can create promptable worlds on-demand with video and audio in sync, across all these different domains and environments. That’s not a distant-future hypothetical thing,” he said.
But there are also limitations to prompting a world model. We asked how reliably the model would follow an instruction to create, for example, a law of gravity or a certain ability in a character?
“Yeah, so it’s a research preview,” Kahlow replied. “So it’s not perfect, of course, and there are still flaws. It really depends on how difficult the action is. I would say movement works quite reliably.”
Sindi added that more training plus scaling the data and models is resulting in “better following.”
How a video model becomes a real-time runtime
Despite the current limitations of GWM Worlds 2 — especially if you compare it to pre-designed and scriptable worlds like Minecraft or Roblox — the true promise of world models like Runway is that they’ll eventually lead to fully self-generated, real-time games and experiences. Which is an extremely hard engineering problem, as Kahlow reminded us.
“There are two challenges. One is making the model not generate a whole clip at once. So instead, you want it to generate frame by frame while you’re looking at it. And the other challenge is actually making the generation fast, so you can play it in real time.”
GWM Worlds 2 offers real-time interactive worlds streamed in continuous 720p video at 24 frames per second (fps) and audio at 48,000 Hz.
Runway achieved this firstly by taking its foundational audio-video generation model and fine-tuning it to the new WorldPrompt format, so the model can follow that. It then post-trains the model to generate autoregressively.
“And after that, we work on making it real-time through distillation methods,” Kahlow added.
Co-CEO Anastasis Germanidis offered more technical details in our podcast with him. He told us that the process starts from “bidirectional diffusion that basically generates an entire video at once and [makes] it autoregressive.” This allows the model to “generate one frame or a few frames at a time.”
Germanidis described two possible forms of distillation in order to make it real-time: distilling a larger model into a smaller one or reducing its diffusion steps. As a general example, he said a model might go from around 50 denoising steps to four, with some quality loss but potentially comparable results.
The challenges of real-time generation
Germanidis admitted that there were issues with how it generates real-time interactive video.
“The biggest challenge with autoregressive models is error accumulation,” he said. “You’re feeding generated frames back into the model to generate the next frames, and if there are any small errors, they accumulate over time.”
Sindi told us there are also challenges dealing with “infinite generations” of content.
“There’s all these challenges around what context to keep, what to discard that’s not important. And so there’s all these optimizations we have to think about, so we’re not blowing up our GPU memory.”
Another current limitation is long-term memory. “The model does not have perfect memory,” Kahlow said. “That’s still an open research problem.”
Causality and correctness
While performance is the primary challenge for Runway at this time, its world model also has to produce plausible consequences when a user takes different actions.
Germanidis used the example of simulating football; he pointed out that online video training data contains more successful goals than failed goal attempts, so a video model might render the first more convincingly.
“If I take this action versus this action, you want it to generate equally realistic outcomes,” he told us. “That’s, I think, the big gap between video models and world models: that idea of counterfactual generation.”
Sindi told us that evaluation gets harder the more complex interactions get.
“If you have this multi-prompt, multi-character, multi-scene [environment], how do you really understand what was causal and what was not?”
To try and solve that, Runway has some automated verifiable tests. But since GWM Worlds 2 is a research preview, Kahlow noted that doing tests yourself is also advisable — “trying out your model to see what doesn’t work is really important.”
More than gaming — there are agent use cases too
Gaming is the obvious use case for what Runway is building, but there are others. Kahlow mentioned robotics — for example using a simulated environment to test how a robot works.
Another, more intriguing, use case is to use it to test agents at scale.
“Having thousands of simulated environments is much less challenging if you have a suitable model like GWM Worlds,” Kahlow said.
But how does an agent know what’s changed in the world — is there a structured state that it can read, or is it just the generated video and audio that it’s consuming and understanding?
“So there’s no structured state here,” Kahlow replied. “It’s just observing the same thing you might observe in real life, just [in this case] from cameras.”
Sindi noted that GWM Worlds can also be used for “synthetic data generation for agents.”
Finally, Germanidis suggested there’s potential to use these world models alongside reasoning models.
“You’re maybe using some reasoning [for] planning of the scene, and then you’re passing it into the diffusion head that’s actually generating the pixels.”
Anastasis Germanidis
* LinkedIn: https://www.linkedin.com/in/agermanidis/
* X: https://x.com/agermanidis
Timestamps
00:00:00 Introduction
00:05:17 Runway’s Origins and the Bet on Generative Video
00:12:23 The Stable Diffusion Story
00:18:44 Gen-2, Controllability, and the Weekend Hack
00:23:02 From Video Generation to World Models
00:28:03 Learning From the World, Not Just Language
00:35:04 Sora, Runway’s Existential Crisis, and Gen-3
00:39:39 Why Real-Time Video Is Inevitable
00:43:06 Interface World Models: Software Without Code
00:50:25 The Fully Neural Operating System
00:55:11 World Models for Robotics
01:02:32 Robot Policies and World Action Models
01:07:47 The Lucid Dream Test
01:11:41 Video Agents and Omni Models
01:23:12 Artists, AI, and Creative Workflows
01:27:14 Physical AI and the Future of World Models
Transcript
Introduction: Runway, Creative AI, and the Early Thesis
Swyx [00:00:00]: Okay, we’re here with, Anastassios from Runway, with, me and Vibhu in the studio. Welcome.
Anastasis [00:00:08]: Good to be here.
Swyx [00:00:09]: Congrats on all your success and progress with Runway. You’re opening offices all over the world. Did you envision this when you first started out?
Anastasis [00:00:16]: Not quite. I think even when we started, we had this idea that, It was more a matter of when, not if, we were seeing the early generative models of 2016, 2017, and just extrapolating, assuming, we resolution, quality increases predictably over time. There’s gonna be a point where most of content will be generated, and that was maybe the initial thesis of Runway was we will need, as a result of t
🔬Bio-security is an AI Arms Race - Eric Nguyen (CEO, Radical Numerics)
2026/09/23
The OpenAI → Hugging Face attack has people asking “what else do we need to worry about?” and Anthropic’s filters flag two things: cyber-security and biology. The natural question is: what about bio-security, then?
Clem Delangue argues that cyber-warfare defensive capabilities need to be open and to keep pace with frontier models’ attack capabilities
Radical Numerics co-founder Eric Nguyen sat down with us and explained why the same models that increase biological capability can also keep defense from falling behind.
Building a virus from scratch
While he was at Stanford, Eric couldn’t get traction on Genomic Language Models (GLMs) for a long time. Biologists didn’t believe it would work, didn’t think they could verify the output, and didn’t see important applications beyond what they could already do. He kept pushing, eventually helping lead the development of Evo and contributing to Evo 2 at Arc Institute. Those models were later used by a separate Arc/Stanford team to generate entire bacteriophage genomes that were synthesized into functional viruses!
Long context unlocks biological intelligence
Early ChatGPT spit out poems and email, and early DNA language models like Evo and Evo-2 could build a genome from scratch. DNA is different, however, from natural language in that it has a very small alphabet (4 characters ACTG) and that its sequences are very long:
* 60K for an average human gene
* long being up to 2.3M
* the whole human genome around 3B.
Innovation in long-context models made this possible about 3 years ago (footnote: striped hyena), long before the frontier labs were building 1M+ context models.
Now Eric and other AI x Bio luminaries have founded Radical Numerics to build and scale GLMs to tack a wide range of biological problems, extending well beyond generating DNA.
Thinking in DNA
Their GLMs already do pretty well with RNA and protein because there are clear markers in the DNA sequence for genes (RNA sequences the perform many functions) and specific genes that encode proteins. This means that the models already generalize to multiple “languages,” before even attempting to train in other modalities, such as 3d protein structure, epigenetics and natural language.
If a model thinks in the DNA language, maybe it understands the imprint that environment left on different genomes as well? Perhaps the model has learned the functional relationship between different sequences, and could extrapolate to new sequences based on that?
And so what we wanted to showcase was that if we show the model progressively better RNAs in a series of steps with its score, right? So you have like low scores first and then you gradually move up the chain. Can the model continue that trajectory on its own? And then in the final step, does it self optimize to a point where it's like the best score it can get? That was the experiment. Can we do that? And so we took a data set, a large data set of aptamers. We held out a portion of the best performing ones and we showed it only the lower ones, but then we ranked it, right? So we showcase lower scores with the RNA aptamers and then progressively got higher, and then ask the model to just like continue with that pattern. And it turns out it was able to recapitulate some of those higher scores that we had not shown it yet.
So, voila: chain-of-thought, thinking in DNA!
The arms race
But much as long-context inference, chain-of-though and multi-modal perception unlocked sophisticated reasoning in natural language LLMs, these capabilities in GLMs are enabling increasingly sophisticated “biological intelligence,” and along with it, greater danger.
According to Eric, defense is currently losing this battle, but Radical Numerics argues to push the frontier harder!
I won’t spoil the details for you. In the episode we talk in detail about:
* Biosecurity as an arms race — and how defense can keep up
* The genome as the imprint of the environment on DNA
* Going truly multi-modal
* How chain-of-though works when you “think” in the language of DNA
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe
🔬 An Oscar, Two Asteroids, and the Algorithm in Your sklearn: John Platt on AI for Science
2026/09/22
How often do you get to talk to a guest who has both an Academy Award and who invented textbook machine learning algorithms? John Platt has an Oscar, two textbook algorithms, two named asteroids, and an Erdos-Bacon number of 6. This was easily the most fun bio of all the guests we’ve read to date. And the result was an epic and fun chat covering Google’s Empirical Research Assistance (ERA), how AI can help battle climate change, and tons of great stories about the co-evolution of science and AI.
John’s colleague Dave Bacon likes to tease John that his career has been defined by being twenty years early to the next big thing. This may be convolutional neural networks (some credit him with coining the term), fusion research, quantum computing. John and Google have been working on solving some of humanity’s hardest problems with AI and computation for well over a decade now. Recently John and his team set their sights on using AI to solve any scientific problem that can be written down as a score.
Google’s Empirical Research Assistance (ERA)
John’s team has taken on many hard scientific problems over the years. In solving these, they noticed a pattern, many scientific problems can be reduced to what John calls a “scoreable task”. Once you have the score function, the goal is to find some code that maximizes the score. The hard part is in formulating the score, but once you have the score finding the maximizer can still be quite a lot of effort.
John’s team set out to automate solutions to this general problem. This came out of the idea of an “auto-Kaggle” AI, which can solve any Kaggle problem you can throw at it. Kaggle is owned by Google, so all the data was ready and easily available to them!
The result is Google’s Empirical Research Assistance or ERA (paper, github, blog). ERA is surprisingly simple conceptually. Gemini (or your LLM of choice) keeps a running tree of past experiments (notebooks) and where they’re going. It’s a close cousin of Monte Carlo Tree Search: at each iteration the Upper Confidence Bound rule picks which notebooks are most promising to mutate. This is optimistic, not greedy, so sometimes even the fifth-best notebook gets chosen. Gemini then proposes mutations for each one, about ten at a time. The history of each branch is shared, so different leaves can learn from each other.
“It’s almost like having a hyper-eager grad student who doesn’t sleep.”
Evolutionary algorithms have been around since the 70s, but this works because Gemini actually knows where to look! What’s even more interesting is that there was a step change between Gemini 2.0 and 2.5, and this went from just not working to working great.
ERA is so powerful that John and his team solved many outstanding problems with it, resulting in at least ten papers. Some of these were climate change related, which we talk about in the next section.
So, we had to ask: if you have an optimization god how do you avoid fooling yourself? John’s answer is that ERA provides predictive models. It’s up to the scientist to make sure they’re truly descriptive. Some of this just involves good old-fashioned careful machine learning science. “It’s a power tool. It can slice your fingers off.” This led to some fun discussion about Kaggle competitions, and the fun ways people can overfit to datasets without meaningfully solving the problem you actually care about: Google’s contrail-detection competition was won by entrants who noticed a half-pixel error in the labels (is the origin at the corner of the pixel or the center?) and this turned out to be a part of the winning special sauce. Great for winning $15,000, not so helpful if you actually want to solve contrails.
“People themselves will act like these LLMs and try to reward hack. It goes back to Goodhart’s law: any metric that becomes a target is no longer good as a metric.”
His advice for where to start instead?
“Always just fit linear regression. Just do it. Just do it. Just do it. Or SVM.”
Tackling Climate Change with AI
John and his team have worked extensively to mitigate the effects of climate change. We talked about several of their initiatives.
Perhaps the most interesting result we talked about was reducing the effects of condensation trails (contrails) from airplanes. Those little streaks you see running behind planes somehow account for 1% of all human-induced global warming?!? Some of these trails of ice crystals can hang out for days. These crystals are black in the infrared, acting like a thermal blanket that traps heat day and night.
It’s easy to understand what’s happening here, a region of atmosphere becomes “ice supersaturated”, and a tiny bit of exhaust seeds water vapor that instantly crystallizes. The scale here is astounding, with a single gram of exhaust resulting in ten kilograms of ice crystals.
The solution to all of this is quite simple, in principle! We know what parts of the atmosphere are most likely for the trails to form. Just have the planes drop a flight level or two. Problem solved, right? Well, the hard part is accounting for how much warming was prevented. This is a counterfactual problem, parts of which stumped John’s team for over two years. They had a working model for the heat-trapping half, but not for the reflected sunlight. ERA was able to find a simple model with some confounders they hadn’t considered. Cracked it!
Modeling climate generally is a hard problem. Climate is best thought of an attractor of many different possible weather outcomes. This makes it much harder to model.
“Weather is where you are on the attractor, and climate is the statistics of the attractor. The problem with climate is that we’re altering it. The attractor itself is changing, it’s moving.”
John and his team have worked on treating both the symptoms and the disease of climate change, with several other works in the area. Another fun example we briefly cover is FireSat, a way of using a constellation of satellites to rapidly identify fires before they grow too big to put out. For anyone living in California, you understand the problem. In dry years a small fire can result in hundreds of thousands of acres. If you could find this fire when it’s the size of a room, it could be put out. By the time it hits an acre we have a much harder problem.
Where is this all going? Looking forward by looking back
By now it should be clear John has an incredible and unique view over the intersection of science, computation, and AI. John talked about a class on physics of computation he took with Richard Feynman back in 1982. This was when quantum computing was an ill-defined concept with no theory or experimental backing. John recalls every Tuesday was a guest lecture, and every Thursday was Feynman explaining why the Tuesday guest was wrong. John also recalls doing science back when there was essentially no compute, a million operations per second was cutting edge.
What is John’s recommendation: the most important skill is developing deep domain expertise. There’s no other way to develop taste than to tackle hard problems. One surprising part of this is that John recommends spending time doing things the old fashioned way. Play with tools, and just implement things yourself.
“You could drive up the mountain, or you could hike up the mountain, and maybe it’s okay, even fun, to occasionally hike.”
Summing it up, John’s message to the audience is that there will still be a place for scientists, and that if anything it will just open up more opportunities for “the creative stuff, the rigorous stuff, the philosophy stuff.” But don’t forget to spend time doing the grunt work.
“There just seems to be this strong impetus in the world to optimize and squeeze everything out. But you do lose something when you hyper-optimize. It’s overfit.”
And whatever tools you end up using, John’s advice is the same one Feynman gave him forty years ago: you must not fool yourself, and you are the easiest person to fool.
We had a great time talking with John. We hope you enjoy!
Also in this episode
* Fusion is three years away, not thirty, if you ask John. And why the Lawson criterion means every fusion approach has an Achilles heel.
* Why superconducting qubits are still finicky.
* The asteroid he named after his mom, which turned out to have a moon.
* The looming helium shortage nobody talks about.
* How NeurIPS started as people crashing a private workshop at Snowbird, and why Hopfield networks are all you need.
* Being Carver Mead’s sysadmin on a VAX with an 80 MB disk the size of a dishwasher.
* Finding asteroids in 1985 with film, a stereoscope, and a letter to Brian Marsden. The Vera Rubin Observatory found 11,000 in six weeks.
* The Feynman effect: total clarity in the room, none once you leave.
* Quantum echoes, the NISQ era, and why he thinks quantum is neither thirty years away nor tomorrow.
* A startup that wants to inject mercury into a fusion reactor and sell the transmuted gold. “It might not work.”
* John’s 20% time rule for his own group: do stuff for learning, and you don’t even have to tell him what.
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe
Jev: System One models for Prod, not God — with Diogo Almeida, CEO, TypeSafe AI
2026/09/21
Tickets for AIE NYC now open, and apply for the invite-only AIE CODE. Join us!
We have an unusual relationship with today’s guest: for years since coauthoring the InstructGPT paper, Diogo Almeida had been saying that API-available frontier models have been going down the wrong path, everything from the alignment to refusals to reliability perspectives, that we have dropped every mode other than autoregressive chat-tuned LLMs because of the overwhelming success of ChatGPT.
In a launch video now viewed ~40M times (by comparison, GPT4o was 22M, Fable 5 was 15M, Navier Stokes was 74M, and 6 Astra was 137M), Diogo introduced Jev and it immediately took over the AI timeline — we’ll skip full Jev explainers because your favorite AI influencer/educator has probably already done one. We also collected:
* the official patterns and cookbooks you should see first, from Allie
* Jev usecases
* speed based - games and computer use
* the voice + computer use example we discuss at 1h34 mins
* voice + browser control
* The must not miss Doom demo
* Driving cars in games
* Excalidraw
* virtual try-ons
* “Smart Games”/smart NPCs
* guided responses in text messages
* Jev for coding agents has an official guide
* jev for linting
* compacting tool calls
* reasonable pushback from Theo - Diogo has published a note on the Tyranny of the KV Cache that you should read as a followup after the pod for Jev + coding agents, because of his belief that Cache Rules Everything
* Programming Languages built atop Jev (Diogo’s fave)
* Jev for analytics replay and user journey review
* “dark data”
* entity resolution
* natural language search
* “smart software”
* a core goal of Jev is to “disappear into the background” - eg as unremarkable as regex
* Jev as a judge
* Jev memes
* Jev vs LLM capabiltiies
* blending transformers and classifiers
* about the confidence api
* Jev vs GLiNER (note difference/pushback, agreed, agreed, agreed)
* Jev on trolley problem
* Jev Bush
Instead we’ll focus on what we can uniquely offer — a broader philosophical and mission-based understanding of how and why Jev was created, and what you should expect next in terms of future models from TypeSafe (ReasoningJev?) and what usecases and ideas you should work on vs the 55th low effort clone of Jev’s API or doing a generic JevBench benchmark - something Diogo has rejected publicly.
Why RLCD: Three kinds of RLHF, and why they are ALL the wrong north star
Diogo knows a good deal about RLHF, given that he was on the team that pioneered post-training at OpenAI — and traces the three branches to Christiano et al 2017 (the robot backflip demo), Stiennon et al 2020 (learning to summarize) and his baby, Ouyang et al 2022 (InstructGPT). From there on, every innovation from Function Calling to Structured Outputs to Reasoning felt like a hack on top of the string based, sequence to sequence prediction paradigm. As he mentions on the pod, from 2023-2024 he struggled unsuccessfully, due to both personal and organization underestimation, to train a model that accurately addressed what he saw as the core problem with making LLMs the heart of software: reliability.
Jev’s core innovation is "Reinforcement Learning for Calibrated Decisions”, a novel, unpublished technique that optimizes for “answers with epistemically honest probabilities on System One tasks” rather than human rated feedback (RLHF) — which causes hallucinations, sycophancy, and permanent reliance on humans — or programmatically verifiable outputs with rubrics (RLVR) — which solves Navier Stokes but exacerbates jagged intelligence and doesn’t integrate well with other software.
We’ve talked about the calibration problem before on the pod, but probably the single best place to understand why RLCD became necessary is Diogo’s AIE talk, which discusses why a generation of training helpful AI assistants for humans has impaired them for training models for composable, programmable AI for automation.
At the end he also teases his contrarian opinion on scaling laws - which teases how to build a modern neolab without the billions of dollars the major labs have…
The Bitterest Lesson: Tasks and Data beats Compute
We spend a good amount of time discussing Diogo’s essay on the Bitterest Lesson:
His point is that “You get what you optimize for and the bitterest lesson in ML is that the most important part of it isn’t ML at all.” - and picking the right north star, eg upvoting for user preference vs being integrated into tool calls - makes everything else fall in line.
We’re excited to catch up with a freshly dyed Diogo to discuss:
* Why AI can solve extraordinarily hard problems but still fail to automate basic work
* What System One Models are and why Jev is built for software rather than chat
* RLHF, mode collapse, calibration, and the hidden costs of optimizing for human preferences
* Why refusals become a problem when AI is buried inside software dependencies
* Why TypeSafe rejects public benchmarks and optimizes for intelligence per dollar
* The “bitterest lesson”: why the right task and the right data can matter more than compute
* Why TypeSafe thinks of itself as a data lab rather than a model lab
* RLCD vs. RLHF and RLVR as fundamentally different North Stars for AI
* Why reliability and robustness matter more than simple determinism
* Jev’s programming primitives and how intelligence maps into software control flow
* Why developers should decompose AI workflows into small, measurable decisions
* How structured state replaces giant prompts and system messages
* Why Diogo thinks AI should eventually disappear into the background of software
* The “inverse SaaS-pocalypse” and how AI could supercharge existing software
* System One vs. System Two intelligence and the limits of reasoning models
* Dark data, computer use, real-time intelligence, and Jev’s biggest early use cases
* Why Jev could reshape coding agents built around a single-model architecture
* Why Diogo says he wouldn’t pre-train with $1 billion
* The OpenAI journey that led to TypeSafe and why he thinks many neo-labs are approaching AI incorrectly
* Coding agents beyond the KV cache, shared state, sub-agents, and the multi-agent future
Diogo Almeida
* LinkedIn: https://www.linkedin.com/in/diogomda
* X: https://x.com/CompleteSkeptic
* TypeSafe AI: https://typesafe.ai/
Timestamps
00:00:00 Jev Launch Week and the AI Economic Revolution
00:02:50 What Is Jev? System One Models and Programmable AI
00:05:54 RLHF, Mode Collapse, Calibration, and Yann LeCun
00:10:29 Programmatic AI, Refusals, and Safety Alignment
00:17:21 Why TypeSafe Rejects Public Benchmarks
00:20:43 The Bitterest Lesson: Data, Compute, and the Right Task
00:24:59 RLCD vs. RLHF and RLVR
00:28:42 Why Powerful AI Still Hasn’t Automated the Economy
00:39:55 Reliability, Robustness, and Determinism
00:48:11 Model Versioning, LTS, Speed, and Intelligence per Dollar
00:54:04 Inside Jev’s API and Programming Primitives
00:58:28 How to Build with Jev: Structure, Decomposition, and Small Decisions
01:18:28 The Inverse SaaS-pocalypse and AI Disappearing into Software
01:33:21 Computer Use, Dark Data, and Jev’s Biggest Use Cases
01:38:48 How Jev Could Reshape Coding Agents
01:41:00 AI Safety, Frontier Pacing, and the Limits of RLVR
01:48:03 Why Diogo Wouldn’t Pre-Train with $1 Billion
01:55:19 The OpenAI Story Behind TypeSafe
02:01:41 Why Diogo Thinks Most Neo-Labs Are Getting AI Wrong
02:08:00 Coding Agents Beyond the KV Cache and the Multi-Agent Future
Transcript
Introduction: Jev Launch Week and Developer Momentum
Swyx [00:00:00]: Okay, we’re in the studio. A special occasion because this week, Diogo, my good buddy, launched Jev, and it’s been taking over the complete timeline. How do you feel? What’s it like to be you right now?
Diogo Almeida [00:00:16]: Emotionally?
Swyx [00:00:17]: Yeah.
Diogo Almeida [00:00:17]: Never been worse. Like, I’m a ragged corpse of a person right now because there’s so much going on, and I’m like a technical CEO, so I have, like, a lot of fires to fight.
Swyx [00:00:29]: Yeah.
Diogo Almeida [00:00:29]: But mentally, I feel—I say this all the time, and I’ve been saying this kind of for years in my over-under events. Like, I feel like the entire AI field is like one of those, like, carnival house of mirrors, and everyone is just insane and saying the weirdest stuff that doesn’t make sense. And it feels like for just this week, like, I’m on a better in sync with reality and like, oh, people see it now. AI can be so much more than what was once thought.
Diogo Almeida [00:01:06]: And like, yes, we are going to make. Like, an AI-based economic revolution is back on the table, and this is f*****g awesome.
Diogo Almeida [00:01:17]: I’m so jazzed the developers get it. It’s, it’s, Yeah, and I want to show my eternal gratitude to the developers and
Swyx [00:01:25]: Yeah.
Diogo Almeida [00:01:26]: I’m so jazzed about the community and everything. It’s so great.
Swyx [00:01:28]: Yeah, you were saying yesterday that you decided to prioritize the town hall and not a bunch of, like, VIP, investor-type people because you wanted to make sure that they are the people that you get your most, attention, right? The engineers, the developers.
Diogo Almeida [00:01:43]: Yeah, it felt a little like, oh man, I’m talking to, like, really important people right now.
Swyx [00:01:47]: Yeah.
Diogo Almeida [00:01:47]: I probably shouldn’t reveal who.
Swyx [00:01:48]: Yeah.
Diogo Almeida [00:01:48]: But it feels a little bit dirty for me to, I’m, like, perhaps overly genuine in things. Like, it feels, like, dirty if, like, in my gigantic calendar event of people to talk to, the community isn’t one of those.
Swyx [00:02:04]: Yeah.
Diogo Almeida [00:02:04]: And actually, in my ideal world, it would be, like, community all the time. I was thinking, “Should I host a town hall while walking to y
Underwriting Superintelligence: Backing Agents you can Sue — Rune Kvist, AIUC
2026/09/16
AIUC first got our attention with the NFDG backing, and have just announced a $40M series A today, with the most impressive industry advisor list we may have ever seen for an early startup behind AIUC-1, their agent standard backed by real insurance:
From being Anthropic’s first product hire to building the standards, testing, and insurance infrastructure meant to make frontier AI deployable, Rune Kvist is betting that the biggest constraint on AI adoption won’t be capability it will be trust. In this episode, the AIUC cofounder joins swyx and Vibhu to announce a new $40M round and explain why companies like Cursor, Harvey, Lovable, and ElevenLabs are increasingly confronting a problem that gets harder as AI gets better: who is responsible when autonomous systems fail?
We go deep on AIUC-1, the emerging standard for agent security, safety, and reliability; how AI agents are stress-tested for jailbreaks, hallucinations, and data leaks; and why Rune thinks standards and insurance could become critical infrastructure for AI. We also discuss the growing trust gap between governments and frontier labs, AI-enabled cyber and biological risks, why every model can ultimately be jailbroken, what happens when a $20 coding agent causes $200M of damage, whether AI engineers should be certified, and why even after AGI there may be one job the labs can never do themselves: be their own watchdog.
We discuss:
* Why risk, liability, and trust may become the binding constraint on AI adoption
* Rune’s path from reading the Scaling Laws paper to joining Anthropic in its earliest days
* What Anthropic understood about scaling, compute, and the future years before it became obvious
* Why Waymo illustrates the gap between AI capability and real-world deployment
* AIUC’s $40M round and work with Cursor, Harvey, Lovable, ElevenLabs, and other frontier AI companies
* AIUC-1: a standard for AI agent security, safety, and reliability
* How agents are tested for jailbreaks, hallucinations, and data leakage
* Why most AI companies optimize the happy path without seriously stress-testing adversarial cases
* Why AI standards may need to update every quarter instead of every decade
* The emerging trust gap between frontier AI labs and governments
* Cybersecurity, child safety, biological weapons, and the expanding frontier-model risk surface
* Why standards and insurance may need to evolve together
* How Lloyd’s of London can insure AI systems and bring trust to enterprise deployment
* What happens if a $20 Cursor subscription contributes to a $200M plane crash
* The Air Canada chatbot case and how AI failures are beginning to clarify legal liability
* Why copyright may be one of the hardest AI risks to insure
* Evals, mechanistic interpretability, monitoring, and models becoming aware they’re being tested
* The impossible CISO mandate: adopt AI fast, but don’t let anything go wrong
* Why robotics will make AI liability dramatically more consequential
* Whether AI engineers should have Level 1, 2, and 3 certifications
* AIUC’s roadmap across agents, frontier models, robotics, and universal red teaming
* Why AGI could become a question of national sovereignty
* Why the labs can never fully serve as their own watchdogs
* The Big Short problem: how do you stop competing watchdogs from racing standards to the bottom?
Rune Kvist
* LinkedIn: https://www.linkedin.com/in/runekvist/
* X: https://x.com/RuneKvist
AIUC
* https://aiuc.com
Timestamps
00:00:00 AIUC’s $40M Round and the Risk Bottleneck for AI
00:01:07 From Scaling Laws to Early Anthropic
00:07:58 Why Trust, Not Capability, Could Limit AI Adoption
00:12:19 Founding AIUC and Building AIUC-1
00:18:52 How AI Agents Are Audited and Stress-Tested
00:25:26 Frontier Models, Government, and the AI Trust Gap
00:33:32 Cyber, Child Safety, and AI-Enabled Biological Risk
00:38:14 Why Standards and Insurance Belong Together
00:41:45 What Does an AI Insurance Policy Actually Cover?
00:50:44 The $20 Cursor Subscription and the $200M Plane Crash
00:53:53 AI Liability, Monitoring, and Earning Enterprise Trust
00:56:21 From AI Agents to Models to Robotics
00:58:29 Copyright, Adverse Selection, and AI Insurance
01:03:28 Evals, Mechanistic Interpretability, and Eval Awareness
01:08:36 The Impossible Enterprise AI Mandate
01:11:52 Prediction Markets vs. AI Audits
01:14:43 Should AI Engineers Be Certified?
01:19:10 AIUC’s Roadmap, AGI, and Who Watches the Watchdogs?
Transcript
Introduction: AIUC, the $40M Series A, and Risk as the Adoption Bottleneck
Swyx [00:00:00]: Okay, we’re in the studio with Rune from AIUC, the Artificial Intelligence Underwriting Company, with our trusty co-host, Vibhu. Welcome.
Rune Kvist [00:00:10]: Thank you. Thanks for having me. Thank you.
Swyx [00:00:11]: What are you announcing today?
Rune Kvist [00:00:12]: We have raised $40 million, led by Ribbit Capital and First Harmonic.
Swyx [00:00:17]: You first came to my attention when Nat and Daniel invested in you guys. Is the story, like, pretty much the same? Like, what are you today versus what you thought you were back then?
Rune Kvist [00:00:26]: When we raised our seed round, we had a hypothesis that at some point risk was going to hold down adoption. At that point in time, that felt kind of hypothetical, and I think that is now over. Clearly, the moment is now with Mythos and Fable. It’s pretty obvious that literally the binding constraint on adoption is risk. And so for us, it feels like this is a natural continuation of the same hypothesis, but where previously it was speculation, now it feels like fact.
Swyx [00:00:54]: And let’s get a list of the customers that you’re highlighting as part of your Series A.
Rune Kvist [00:00:58]: Totally. Yeah. So we are now working with folks like Cursor, Harvey, Lovable, ElevenLabs.
Swyx [00:01:05]: Yeah. Amazing. Congrats.
Rune Kvist [00:01:06]: Thank you.
Swyx [00:01:07]: So you were famously one of the first hires involved in GTM and product. I’m just kind of curious: what was your path into AI? Just recap.
Rune’s Path Into AI: Scaling Laws, Capital, and Anthropic
Rune Kvist [00:01:18]: Yeah.
Rune Kvist [00:01:19]: Late 2021, I sold a company, my first company, an edtech company. I had a bit of time to think about what was next. I came across the Scaling Laws paper, and that just struck me like lightning. I was just like, “This is a big idea.” In short, the Scaling Laws paper just says the bigger the model, the smarter the model.
Swyx [00:01:38]: So this is the Kaplan one, not the Chinchilla one?
Rune Kvist [00:01:40]: Exactly, the Kaplan one.
Swyx [00:01:42]: Yeah.
Rune Kvist [00:01:42]: And the important thing that clicked for me there was, oh, now capital will understand this. If you put in more money, you get more money out, and so that will kick off a hype cycle. And so you get a sense of predictable returns, which is, in fact, what’s played out. And so I just packed my bags. I’d never been to San Francisco. I’d never been there. I just packed my bags, flew out here to find the people who had written it. And at the time, they had just started a small lab called Anthropic. There were around 40 people at the time or so. Drank a bunch of coffee until I eventually got introduced to Dario. And at the time, they were wrestling with some of these questions of, like, should we deploy our models? Should we make revenue? How should we engage with the rest of the world? They’d just broken off from OpenAI, and it’s been publicly reported that they were kind of concerned with how they were dealing with deployment. So they were wrestling with some of those questions. At this point, this is early fog of war, like early 2022. The hottest product at the time was, like, Jasper. Like, there’s nothing out there. So where value was going to accrue, and what the different parts of the stack were going to be, were all open questions.
Swyx [00:02:48]: I want to highlight to people, you ask these questions because you have a PPE background.
Rune Kvist [00:02:52]: Yes.
Swyx [00:02:52]: I actually was in Singapore in one of the sort of feeder programs for prepping people for PPE. So I had a tutor. We learned, you know, philosophy and politics and economics. But, like, I think your kind of background matters. Machine learning people who read the neural, Scaling Laws paper would not necessarily draw the same conclusions that you did. Whereas any capitalist would read that and go, “Holy s**t.”
Rune Kvist [00:03:19]: Correct.
Swyx [00:03:20]: Right?
Rune Kvist [00:03:21]: Yes.
Swyx [00:03:21]: Who tipped you onto that paper? Because it’s not a paper that you normally read, right, like, in your circles?
Rune Kvist [00:03:26]: Yeah. I think I’d actually, ever since AlphaGo, had some appreciation that AI was a big deal.
Swyx [00:03:36]: Yeah.
Rune Kvist [00:03:36]: But it kind of felt like it raised all these kind of interesting philosophical questions, but it was kind of not clear from afar where exactly that would go. But it was obvious enough that it was like, this is going to be a big thing if we find the kind of right mechanism to kind of get the techno-capital machine to work on this. But it was just not clear. And so I think there was some way in which, like, that became obvious, and also it wasn’t as obvious at the time than it is now, right? Like, it was just like, wow, this is so interesting. But it still felt, coming from kind of a philosophy and economics background, it felt like if this turns out to be true, you’re going to be wrestling with all of the big questions in society. Everything you’ve learned about politics gets thrown out of the window. Everything you’ve learned about economics at least gets challenged. And so what felt interesting was to be at that frontier that has ramifications across everything. So that’s why I sought it out.
Swyx [00:04:32]: I mean, clearly really good insight. For people who don’t know, the PPE progra
Humanity’s Last Invention — Richard Socher of Recursive
2026/09/14
At 1:09:00 we talk about the rise of AI x Finance, and AIE NYC is one month away - our hotel block is 97% sold out, get tix & travel ASAP - we will announce speakers from Bridgewater, Ramp, Coatue, Mastercard, Vanguard, Coinbase, Blackrock, Fidelity, Point72, Capital One, JPMC, Wells Fargo, Bloomberg, A24 (yes the movie studio) Labs, Two Sigma, Apollo Global, and more soon!
From helping pioneer core ideas in NLP to now building AI systems that can automate AI research itself, Richard Socher is betting that the next major step in AI is recursive self-improvement. He is the founder of You.com, AIX Ventures, and now Recursive, which has assembled some of the best open-endedness (& self improving agent) researchers in the world and raised a $4.65B seed round.
In this episode, Richard joins Latent Space to unpack his vision for the “Eureka Machine”: a superintelligence that can improve the process of invention itself, accelerate AI research, and eventually tackle major problems across science, energy, materials, biology, and more.
You can get his book “The Eureka Machine” here!
We go deep on Recursive’s early results, including an AI research system that Richard says outperformed humans and their agents on optimization tasks in less than two days, as well as work on NVIDIA GPU kernels where the system discovered improvements without relying on a team of CUDA experts. Richard also explains why he thinks AI research that currently takes thousands of people and years could eventually be compressed into weeks. These results are summarized in his 20 minute AIE keynote, where we also discuss his 10 dimensions of intelligence:
We also explore the harder questions around increasingly capable AI: reward hacking, whether Anthropic-style constitutions actually work, AI regulation and proposals to “pace” frontier development, open-source models as geopolitical soft power, whether today’s LLM paradigm is enough, and what happens if AI systems eventually begin choosing their own goals. Richard reflects on the rejected research that helped inspire Alec Radford’s GPT, open-endedness, the AI Economist, simulations of entire economies, and his framework for thinking about the upper bounds of intelligence itself.
We discuss:
* The Eureka Machine and Richard’s vision for an AI that can automate invention
* Why Richard is optimistic about superintelligence for science and technology
* Why AI hard-takeoff scenarios may underestimate physical and economic constraints
* The risks of regulating intelligence itself instead of specific AI applications
* Reward hacking and why increasingly intelligent AI makes objective design harder
* Richard’s critique of Anthropic’s constitution and constitutional AI
* Alignment vs. personalization and whose values an AI should follow
* Why open-source AI matters for resilience, competition, and geopolitical soft power
* Why Richard left You.com’s frontier-model work to start Recursive
* Recursive self-improvement and automating the process of AI research
* Whether today’s LLM paradigm is enough — and why Richard is less bullish on world models
* DecaNLP, early prompt-based generalization, and the research that influenced GPT
* Why rejected research can shape entire technological timelines
* Open-endedness, evolutionary approaches, and rainbow teaming
* What happens if AI systems begin setting their own goals
* Why simple objectives like profit maximization can produce dangerous reward hacks
* Recursive’s long-term plan to apply self-improving AI to science
* The compute, hardware, and economic constraints on AI takeoff
* Recursive’s early NanoChat, NanoGPT, and GPU kernel optimization results
* Why automating AI research could reduce years of work to weeks
* Reward engineering and what makes auto-research systems actually work
* The AI Economist and using simulations to test economic policy
* Whether LLMs can realistically simulate people and entire economies
* Benchmark bugs and evaluation harnesses and the difficulty of measuring AI progress
* Recursive’s near-term focus on AI for AI research
* Harness optimization, sandboxing, and web search as core agent infrastructure
* You.com and the search stack for AI agents
* AI in finance, backtesting, and data leakage
* Richard’s three fundamental components and ten “spaces” of intelligence
* The theoretical upper bounds of vision, communication, knowledge, and computation
* Creative intelligence, metacognition, and AI-generated goals
* Survival and replication and why AI does not necessarily need to fear being turned off
* High agency and ambitious goals and Richard’s advice for people building with AI
Richard Socher
* X: https://x.com/RichardSocher
* LinkedIn: https://www.linkedin.com/in/richardsocher/
Timestamps
00:00:00 The Eureka Machine and Superintelligence
00:02:23 AI Optimism, Slow Takeoff, and Regulation
00:07:56 AI Safety, Reward Hacking, and Anthropic’s Constitution
00:11:49 Alignment, Personalization, and Open Source AI
00:15:46 Why Richard Started Recursive
00:20:03 Recursive Self-Improvement and the Founding Team
00:22:55 Are Today’s LLMs Enough?
00:29:03 DecaNLP, GPT, and the Rejected Idea Ahead of Its Time
00:34:38 Open-Endedness and Evolutionary AI
00:36:38 What Happens When AI Chooses Its Own Goals?
00:41:16 Superintelligence for Science
00:42:40 GPUs, Compute, and the Limits of AI Takeoff
00:45:07 Recursive’s Results: AI Beating Humans and Their Agents
00:49:14 Reward Engineering and Auto Research
00:53:12 The AI Economist and Simulating Entire Economies
00:58:07 LLM Simulations, Personas, and Mode Collapse
01:03:38 Recursive’s Roadmap, Agents, Search, and Finance
01:09:13 The Upper Bounds and Spaces of Intelligence
01:30:21 Goals, High Agency, and Advice for Builders
Transcript
Introduction: Richard Socher and the Eureka Machine
Swyx [00:00:00]: We’re here in a studio with Vibhu and myself and Richard Socher. Welcome.
Richard Socher [00:00:06]: Thanks for having me.
Swyx [00:00:07]: We just talked about the Eureka Machine, or we just released a talk, at AI Engineer about the Eureka Machine. Is it — you said it’s your life’s goal. What is the Eureka Machine?
Richard Socher [00:00:16]: The Eureka Machine is the ultimate invention that will afterwards invent most everything for humanity. It’s essentially a superintelligence that can be given any goal, any environment, reward, and then it will try its best to achieve those goals to create the kinds of inventions that humanity would hopefully ask it for.
Swyx [00:00:45]: Yeah, I think we have the book pulled up here that you’ve written.
Richard Socher [00:00:50]: That’s right, yeah. I finished it last year, a little bit before we started Recursive, and now we’re gonna try to build parts of that.
Swyx [00:00:57]: You finished it last year. It’s July. What takes so long?
Richard Socher [00:01:01]: Oh, man, books. Books are incredibly slow.
Richard Socher [00:01:04]: It’s ridiculous. That whole industry is just unfathomably slow.
Richard Socher [00:01:07]: So a lot of the ideas have been out there for a while, but yeah, I’m really glad it’s finally coming out in September this year.
Swyx [00:01:14]: We might have AGI by then. Like, we don’t know.
Vibhu [00:01:18]: Any key takeaway that you’re most excited to put in here?
Techno-Optimism, AI Upside, and Slow Takeoff
Richard Socher [00:01:21]: Yeah. The key takeaway, I think, is that people could and should be much more excited about the positive implications of superintelligence, especially for science, physics, chemistry, biology, but also economics and astrophysics, and all kinds of other engineering tasks. I think there is so much more that can be done with better technology. And right now, I feel like a lot of people need, like, better marketing, not just for the future in general, but also, better marketing for technology and in particular for AI. And this book, should show even the AI skeptics, how much positive upside there is for AI, especially when it comes to inventing, new scientific discoveries.
Swyx [00:02:09]: I think you quoted the techno-optimist manifesto from, Marc Andreessen, which I think was, like, beautiful in its, ambition and clarity and simplicity almost as well.
Richard Socher [00:02:18]: I agree. Yeah. Yeah, you can disagree with him on some things, but, like, I think he’s right on the techno-optimism.
Swyx [00:02:23]: Where do you think optimists get in trouble?
Richard Socher [00:02:26]: Like, you shouldn’t have blind optimism. You should be very clear-eyed, like, especially when with such an omni, like, use type of technology as AI is, you need to think about the potential downside scenarios, especially when people use it for things that you don’t want them to use it for. It’s a little bit like the internet, and I feel like people are trying to regulate AI sometimes because of those potential downsides the way you would regulate the internet, if you were to say, “Well, because there’s bad content on the internet, like torture porn or whatever, like, we should just make it slower. That way, you can’t share the illegal content as quickly, or we should make the hard drive smaller so you can’t store as much illegal content.” But I’m like, “That’s not how you regulate that.” that’s like saying like we should regulate intelligence in the abstract. What you should regulate to avoid those downside scenarios, even as an optimist, are the specific applications. Sure, I don’t want, like, some AI surgeon to, like, practice some RL moves in my brain. It should be fully FDA certified. Sure, I don’t want any random startup to, like, drive on the highway, and cause a major accident. It should, like, have proper certifications before it’s let loose on the highway. But I feel like those downside scenarios, that some optimists sometimes maybe don’t consider enough are fairly easily regulated, compared to, what the doom
🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
2026/08/26
A few years ago, Caltech Prof. and co-founder of Accelerated Understanding, Anima Anandkumar set out to develop the first open-source weather model with AI. Talking to experts in the field, she was met with skepticism. Weather is chaotic, physics simulations are hard, have been developed for decades, and require supercomputers, the data just isn’t there. Despite reservations, Anima went forth and built. Within a year her team had developed FourCastNet, a predictive model that is competitive with the best physics-based simulations available. Thanks to Anima, and her follow up work, anyone can now predict weather accurately over a short timescale using consumer grade GPUs.
In the fifteen or so science episodes we’ve released on Latent.Space, we’ve covered atoms, molecules, materials, biology, and math. Anima is a pioneer in studying physical systems that are continuous. Weather, fusion, and fluid or heat flow are huge areas of science that are extremely difficult to model: they are large, chaotic, and fundamentally multi-scale. This is a field the AI community has somewhat neglected, but one we expect will grow fast. We plan to cover large physical systems more in coming episodes.
One thing you can glean from Anima’s work is that this area of AI resists the scaling ideas that have permeated the rest of the field. The data isn’t there: open source datasets in many of these domains are limited to tens or hundreds of thousands of examples, far from what token-hungry transformers need. Even worse, the resolution that physics demands pushes the context length into the hundreds of billions, so you can’t just throw more tokens at the problem. That isn’t a ceiling though, just a slower road: progress here comes from building in structure and inductive biases. Sorry for all you bitter-lesson-pilled language modelers.
“If each dimension is even a few hundred grid points, which is where industrial scale starts... we’re talking hundreds of billions to even a trillion context length. So forget ever having a transformer for anything of this scale, all of the world’s compute will not be enough.”
The math underneath
To tackle these systems, Anima pioneered a technique known as Neural Operators, one of the most beautiful theoretical developments in AI of the last decade. These allow you to combine data and physical laws to enable multi-scale inputs and outputs. We’re no longer modeling a grid, we’re modeling a function that evolves over many scales. This allows Anima and crew to build in priors based upon physical intuition.
To see how physical priors are still helpful for AI modeling, let’s revisit the problem of weather forecasting on a global scale. The earth is a sphere, which meant that accurate modeling involved using the right basis set — the Spherical Harmonics. Run a weather model on a grid and it blows up fast. Move to the natural basis for the problem and it stays stable far longer, long enough to roll out months ahead instead of days. Anima’s Fourier Neural Operator learns directly in this frequency domain, and its spherical variant powers FourCastNet 3, which models the weather across the whole globe and keeps running stably far into the future.
The physical world is forgiving
Anima explored Neural Operators across other physical domains too, and one striking observation is that the physical world is more forgiving than you’d expect. In fusion, a few thousand samples are enough to predict plasma disruptions, and to do it a million times faster than traditional simulation.
None of this is a rejection of scale, it is a different route to it. Anima ultimately still wants to build a “foundation model for physics”, a model that spans many phenomena and does both simulation and design. You get there by building in the structure the physical world already has, not by waiting for data that will never exist. It is a start, and it will take longer than the token-driven parts of AI, because for the physical world tokens were never the answer.
“All of the things that work with deep learning, let’s take them, but make them a bit more principled.”
Weather is only the beginning
Neural operators and weather modeling were a personal passion of mine, so we’ve spent much of this blog and the episode exploring this work. Anima has done so much more! In the episode, we cover several other recent developments from Anima:
* Anima has a series of works integrating neural networks and automated proof techniques. We talk about TorchLean, a new framework that lets you write PyTorch-style networks inside the proof assistant Lean and formally verify them. This is a major step for proving bounds on neural networks, something that would be really important for someone trying to, e.g., add a neural network as part of the control loop to their fusion reactor!
* Anima was recently appointed to the United Nations Scientific Advisory Board! We talk with her about her goals of bringing evidence-based viewpoints to policy, and how AI in scientific domains can improve people’s lives all over the world.
This episode has something for every AI or science nerd! Elegant math? ✅ Old school harmonic analysis? ✅ Fundamental developments in modern AI? ✅ Practical ways of modeling the physical world? ✅
Give it a watch!
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe
Simulation: the new Scaling Law — Joon Sung Park, Simile AI
2026/08/21
When we first dicsussed the Summer of Simulative AI in 2024 we knew it would be a brief summer, but it has recently come back with a vengeance with SimGym in April and now Simile AI’s $2B Series B, backed by GreenOaks and Index Ventures with prominent backers like Fei-Fei Li and Andrej Karpathy, running tens of millions of simulations for Fortune 100 clients like CVS and 85–99% accuracy vs human focus groups.
Time to catch up on why this Second Summer of simulation is working!
From creating Smallville, the landmark 2023 paper on Generative Agents that showed AI characters could remember, plan, socialize, and develop emergent behaviors, to now building foundation models of human behavior, Joon Sung Park is trying to answer a much bigger question: what if we could simulate the world before making decisions in it? In this episode, the Simile co-founder and CEO joins us to unpack the path from generative agents to digital twins, why today’s frontier models still fail to capture how humans actually behave, and what it would take to eventually simulate all 8 billion people on Earth.
We go deep on Simile’s approach to modeling human behavior: long-form interviews, observational and transaction data, randomized controlled trials, population-level and individual-level models, and post-training on the causal mechanisms behind why people make decisions. Joon explains how his research created digital twins that reproduced human behavior and attitudes 85% as accurately as people reproduced their own responses, why models optimized to be rational can be bad simulations of irrational humans, and why understanding “social physics” may require changing model weights rather than simply prompting frontier LLMs.
We also explore the much larger ambition behind simulation: testing products and policies before deploying them, finding counterintuitive paths toward desired outcomes, modeling emergent behavior across entire societies, and potentially tackling problems like climate change, democratic instability, and UBI. Joon reflects on scaling laws for simulation, the economics of data-center-scale simulated worlds, the connection to Thomas Schelling and psychohistory, why simulation is surprisingly similar to painting, and whether we might already be living in one.
We discuss:
* How Smallville and Generative Agents led to Simile
* Why Joon’s team asked: “What if we can just recreate the world that we live in?”
* Why useful personal agents require deep models of their users
* Memory architectures, Markdown files, and the limits of prompting
* “Social physics” and behavioral foundation models
* Why web data captures what people say more than what they actually do
* Interviews, transactions, observational data, and randomized controlled trials
* Why predicting the future matters less than understanding how to shape it
* How Simile creates representative simulated populations
* Simulation versus prediction and the connection to Foundation’s psychohistory
* How to evaluate simulations instead of simply stacking LLM hallucinations
* Creating digital twins of 1,000 real people and reaching 85% behavioral accuracy
* Why frontier models can struggle to reproduce real human behavior
* Why good simulations need to reproduce human biases and mistakes
* Post-training models on randomized controlled trials
* Population-level versus individual-level simulation
* Scaling laws for human simulation
* The long-term ambition to simulate all 8 billion people on Earth
* Whether simulations could help solve climate change or detect collapsing democracy
* Thomas Schelling and the history of agent-based modeling
* Why future simulations could require an entire data center
* Multi-agent simulations and what happens when simulated people interact
* Replacing expensive human panels with synthetic populations
* Why market research is only the starting point for simulation
* Why Joon sees simulation as surprisingly similar to painting
* Using simulation to study questions like UBI
* Whether we are already living in a simulation
* Why AGI and simulation may be the twin technologies of advanced civilizations
Joon Sung Park
* LinkedIn: https://www.linkedin.com/in/joonspark
* X: https://x.com/joon_s_pk
* Website: https://www.joonsungpark.com
* Simile: https://www.simile.com
Timestamps
00:00:00 Introduction and Joon’s Path from Art to AI
00:01:46 Smallville, Generative Agents, and the Origins of Simulation
00:05:03 “Let’s Just Create a World” and the Future of Personal Agents
00:09:53 Social Physics and Behavioral Foundation Models
00:14:08 Prediction vs. Simulation: How Do You Shape the Future?
00:16:59 How Simile Models Real People and Populations
00:25:35 Evaluating Simulations, Digital Twins, and 85% Accuracy
00:30:23 Post-Training Models to Reproduce Human Behavior
00:40:04 Scaling Laws and Simulating 8 Billion People
00:43:10 From Schelling to Society-Scale Agent Simulations
00:46:13 The Cost and Economics of Simulating the World
00:52:05 Real-World Use Cases, Synthetic Populations, and the Market
00:57:27 The Future of Simulation, Painting, and UBI
01:04:23 Are We Already Living in a Simulation?
01:06:08 Building Simile and Hiring
Transcript
Introduction: Joon Sung Park, Simile, and the Story So Far
Vibhu [00:00:00]: Today, we have Joon in the podcast. Excited to kick this one off. Very exciting company. I wanna kick off and ask you the question, talk us through the story of your life. How have you gotten here?
Joon [00:00:13]: Yeah, for sure. I’m really excited to be here. A story of my life. So I was born in Korea, and I lived there for a good 11 years or so of my life, and then my family moved to Boston. So we moved when I was 11, and my parents were doctors, so they were going through their postdoctoral studies. My dad was a surgeon, so he was doing his sabbatical years at the Boston Children’s Hospital. So I grew up there, not too close to tech. I was very much a music and artsy, painting kind of guy.
Vibhu [00:00:49]: Painting.
Joon [00:00:49]: Exactly. I got into painting a little bit later, in high school, but that’s what I used to do. And then I grew up mostly in the East Coast after Korea. So I lived a good number of years in New Hampshire, and then I went to college in Pennsylvania. And I got into more of this tech scene, in college. So I was originally trained to be an artist. I thought that would be my professional career. So it wasn’t a hobby. It was like, “Hey, let’s make a living out of this.” And then gradually, I got really interested in this idea of, hey, the greatest artist often creates their own medium, and the best medium that we had available today was in computation. So I decided to go deeper into that, and one thing led to another, and we can go deeper into this, but I decided that research was something that I gradually got interested in, and here I am.
Smallville, Generative Agents, and the 2023 Breakout Paper
Swyx [00:01:46]: So there’s a lot that you packed into the research components. You had one of the best papers of 2023, which was the generative agents paper, commonly known as the Smallville paper.
Swyx [00:01:58]: Feel free to call back to anything else that you mentioned, but most people would have heard of you from this. Do you have any statistics on how many people have, like, read it? arXiv gives you something, right? Some stats.
Joon [00:02:10]: Yeah, it’s a good question. How many people have read it, I’m not sure.
Joon [00:02:14]: I know we do keep track of citations, and they are going up quite fast.
Swyx [00:02:23]: Yeah, Google Scholar has 7,200 citations.
Vibhu [00:02:25]: I feel like it made a bigger hit than that, and it was a pretty instrumental paper. It got cited so many times.
Swyx [00:02:34]: It is frequently the answer when people ask, “What is the best paper you’ve read recently?” It’s this one.
Vibhu [00:02:39]: I thought the memory component was pretty underrated. It was a very good early memory system, and one of the biggest papers.
Foundation Models and the Search for Killer Applications
Joon [00:02:47]: Yeah, so maybe I can talk a little bit about how this particular paper came together. So when I got into research, it was back in 2020 when I started my PhD program at Stanford, and that was the year, when we were about to get GPT-3 to be available. So we already had GPT-2, and you could sense that there was this new class of models that was just becoming available in the market, and the team got very intrigued. And the general consensus was, “Well, is this model going to be useful for anything?” “It’s really strange that these models are not trained to do any particular task.” But we decided to take a bet. So a large group of scholars at Stanford, and it was led by one of my co-founders, Percy Liang, and we came together
Swyx [00:03:35]: Who coined foundation models.
Joon [00:03:36]: Who coined the term foundation models. We wrote this paper, where that term came from called Opportunities and Risks of Foundation Models. And during that process, really the thing that I started to think deeply about was, here is a model that is fundamentally new in our ecosystem. The reason why this was new was it wasn’t, again, trained to do anything in particular, but its premise was it could do anything and everything. It was like a stem cell, if you were to take a biology analogy. And I got really interested in this idea that, well, if we were to really think about what are the killer applications that this particular technology would enable, what would that be? Many of my colleagues were using this for simple classification, simple generations. Interesting that these models can do that, but from an interaction perspective, not that interesting. We’ve known how to do that for many decades. And what we came down to was these models are trained on this very broad data from the web, right? So these are human behavioral data. It’s social m
🔬The BioAI Phase Shift - Matthew McPartlon & Neil Patil, Chai Discovery
2026/08/11
This January, four big AI × Pharma tools deals were announced at the huge JPM Pharma conference that takes over San Francisco every year. OpenAI-backed Chai Discovery (now worth $4B) was somehow at the heart despite being all of 2 years old.
The Science team is proud to bring you the first podcast with cofounder Matt McPartlon and product lead Neil Patil to tell the full story!
Editor’s note: not to be confused with Chai AI, which was another top pod of ours.
Pharma suddenly doing big AI tools deals
For the non-pharma people, JPM is JP Morgan’s annual conference for pharma deal-making that takes over San Francisco for a week in January with hundreds of side events, etc. It’s a big thing.
Tools deals for pharma are also a big (new) thing: companies that start as AI for Pharma usually end up building their own drug pipelines instead, and the reason is something like this: convincing pharma to use your tool requires proof that your tool works. Proof means good targets, maybe with good clinical validation. If you have that, then it’s easier to raise money (with a known, if long path to commercialization) or sell (e.g payment in biobucks) for a specific target than it is to sell to lots of companies on a promise that it will work across their portfolios.
The “we’ll just partner / build our own drug” optionality proved to be the only good path up until January. What changed? In short, the tools got good enough for drug design teams to trust.
Good-enough-to-trust unlocks the ability to scale discovery: get more, better candidates into the lab and animal trials faster. More screening for toxicity, better delivery, etc. This means that what you push to the clinic is more likely to succeed.
Tools also unlock new capabilities: mechanisms that are very hard or impossible to develop using lab-based discovery. Designing an antibody that precisely triggers a very specific molecular cascade takes many years of trial and error. Designing bi-specific antibodies (that bind to two different proteins) is similarly difficult. Good design tools can unlock this.
RJ: The fact that the quality of the model has jumped means you’re enabling things you just plain couldn’t do. So it’s a step change. It’s not an efficiency argument at all, or not so much.
Matt: Yeah, exactly. It’s kind of interesting, even for us — it took me a while to believe in the thesis, actually. I talked to Josh for months before Chai started... It’s like, can I beat a mouse, and then can I do what mice can’t do? And then how many levels of interaction can you just keep building on top of that?
Everyone playing in the structural / binding space has an angle here, and some will be better than others, but Chai is pointing to a different unlock: getting good molecules right out of the gate (meaning they don’t then need as much lab work) means that the iteration time is faster. This turns science into engineering: you can design your systems to reduce friction and hill climb towards one-shotting molecules all the way to the clinic.
This, per-se, is not a new thesis: a16z articulated a version of this in 2020. What has changed is that structural models became binding models (how well doesn’t this molecule bind to this molecule, aka “binding affinity). Binding models unlock design, which has been steadily improving. Chai’s observation is that for engineering problems the best product tends to win, and good technology is a necessary but not sufficient condition.
Photoshop for molecules
With that in mind Chai has invested heavily in partnerships that allow them to learn from their Pharma counterparts.
What is kind of cool about working so closely and supporting so many of these partners is we get to really learn about what is the stuff that would be helpful in research. So rather than doing research in a vacuum, based on what would hypothetically be cool, we're able to do informed research based on what our partners have just been organically asking us for help with.
— Neil Patil, (Chai product lead)
This means better UX, such as a molecule editor that is more like a CAD or graphics design program than a chatbot.
Their approach has paid off: since June, Chai has announced three more major deals: Lilly, Novartis, argenx, plus an expansion of their Eli Lily program. This episode is too full of quotable moments for a short blog, so tune in to learn about
* Why protein tokens have the highest downstream value of any token
* Climbing levels of abstraction as models improve
* How Pharma, VC, and research are all just portfolio optimization
* How better tech changes the whole portfolio
* How relentless focus on simplicity leads to scale
Plus much more!
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe
The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
2026/08/03
Watch the full episode on YouTube:
We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiaries of the Inference Inflection.
We return to Baseten at the peak of the 2026 edition of Open Weights debate. Ali has published a viral breakdown of Kimi K3:
And since you last saw him, Philip has spoken at AI Engineer and written the definitive book on Inference Engineering spotted all over SF:
Three years ago, inference engineering barely existed as a category.
Today, it is one of the most critical disciplines in AI. Inference engineering inherently tackles a different question than standard model training: “How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?” Focusing on these creates an entirely new optimization problem.
In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20%, because the errors introduced in different layers could cancel each other out.
Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles.
In this episode, Baseten’s Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API.
We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200%; how quantization errors can cancel one another out; why identical weights can behave differently across clusters; and how Baseten grafted a Kimi vision encoder onto GLM-5.2 without changing the underlying language model.
The conversation then expands beyond LLMs into NVIDIA Dynamo, mega kernels, Rubin, AI-specific chips, local inference, video generation, diffusion versus autoregressive models, and the enormous compute barrier to generating coherent long-form video. Finally, we explore the convergence of training and inference, continual learning through persistent KV cache, and the emerging loop where models help optimize the infrastructure that runs them.
We discuss:
* What happens when a 200,000-token request enters an inference system
* Cache-aware routing and reusing previously computed KV cache
* Why prefill and decode are increasingly handled by different GPUs
* When dedicated deployments become cheaper and more reliable than shared APIs
* How speculative decoding uses a smaller model to accelerate a larger one
* Tool calling, structured outputs, and what LLMs actually do
* What it takes to support a new open model on day zero
* Grafting Kimi’s vision encoder onto GLM-5.2
* Retrofitting inefficient model layers with components from other architectures
* Why models sometimes collapse into repeating the same token
* How hardware, kernels, and race conditions create nondeterministic failures
* Preserving model fidelity while making inference faster
* How quantization errors can cancel each other out
* Why inference optimizations still deliver gains of 20%, 100%, and 200%
* How optimized serving can make a model up to 10× faster
* NVIDIA Dynamo, KV-aware routing, and distributed model serving
* Speculative decoding the speculative decoder
* Why local AI is about making models less dumb while data-center AI is about making them less slow
* Tensor, expert, and pipeline parallelism across GPUs
* Hardware-aware model design, auto-tuning, and the case against mega kernels
* Rubin and why inference is becoming a systems problem
* Whether modern GPUs are evolving into programmable AI ASICs
* Why enormous models like Kimi K3 require GB300-class hardware
* Why open-source video generation still trails Veo, Kling, and other closed models
* The quadratic attention bottleneck behind long-form AI video
* Autoregressive video, real-time generation, and compounding quality drift
* Why future video systems may combine autoregressive and diffusion architectures
* Training for inference and inference for training
* Continuous post-training, deployment, evaluation, and improvement loops
* How GLM-5.2 helped optimize the kernels serving GLM-5.2 itself
* Why faster networking could unlock dramatically faster decoding
* Continual learning, KV-cache compaction, and persistent model memory
Show Notes
* How to build a day-0 API for Kimi K3
* 22580: From GPT2 to Kimi3, Explained
Philip Kiely
* LinkedIn: https://www.linkedin.com/in/philipkiely
* X: https://x.com/philipkiely
* Inference Engineering: https://www.baseten.co/inference-engineering/
Ali Taha
* LinkedIn: https://www.linkedin.com/in/aliestaha/
* X: https://x.com/waterloointern
Timestamps
00:00:00 Introduction and the 200K-Token Prompt
00:03:18 Dedicated Deployments, Speculative Decoding, and Tool Calling
00:11:26 Launching Production-Ready Open Models
00:19:06 Model Retrofits, Failure Modes, and Nondeterminism
00:28:22 Quantization and Canceling Errors
00:32:15 The Race to 10× Faster Inference
00:40:48 Dynamo, Speculation, and Local vs. Data-Center AI
00:50:18 Model Parallelism, Auto-Tuning, and Mega Kernels
01:00:55 Rubin, GPUs vs. ASICs, and Custom AI Chips
01:10:03 Giant Models and the Limits of GPU Memory
01:12:42 AI Video, Quadratic Attention, and Autoregressive Generation
01:21:47 Audio, Images, and Diffusion Models
01:27:32 Training, Self-Optimizing Models, and Continual Learning
01:40:06 Closing Thoughts
Transcript
Introduction: Baseten, Waterloo Intern, and Inference Engineering
Swyx [00:00:00]: Okay, we’re here in the studio with Philip, old friend from Inference Engineering, the book, as well as Baseten and everything that you’ve done, you and I have done before, as well as Ali. Welcome.
Ali [00:00:15]: Pleasure to meet you.
Swyx [00:00:15]: Waterloo intern.
Ali [00:00:16]: Waterloo intern, always.
Swyx [00:00:17]: When did you get “Waterloo intern” as a handle?
Ali [00:00:19]: As a handle? Oh.
Ali [00:00:20]: I think the rebranding happened mid-March. When I saw it was open, I was like, “I have to take it. Up for grabs.”
Philip [00:00:26]: The problem is that Ali is really good at his job and is not gonna be an intern much longer.
Philip [00:00:30]: So we have to figure out who’s gonna get the handle.
Ali [00:00:33]: Well, I’ll pass the torch over to the next intern.
Swyx [00:00:34]: Oh, okay. It can be, like, you just pass it to another Waterloo grad.
Ali [00:00:37]: To another Waterloo intern. No, bruh.
Philip [00:00:39]: Yeah.
Ali [00:00:39]: Intern.
Swyx [00:00:40]: Intern, yeah.
Ali [00:00:40]: And no.
Philip [00:00:41]: You gotta get an intern from Waterloo.
Ali [00:00:42]: Yeah, I’ve gotta get an intern from Waterloo.
Swyx [00:00:44]: Right.
Ali [00:00:44]: But they have to follow the path.
Swyx [00:00:45]: Oh, it could, but it could come from Baseten, so it’s like whoever Baseten gets from Waterloo.
Ali [00:00:48]: Right.
Swyx [00:00:49]: Has the title of Waterloo.
Ali [00:00:50]: It stays in the ecosystem.
Philip [00:00:51]: Exactly.
Ali [00:00:52]: Halfway through the internship, you either get it or you’re out.
Philip [00:00:55]: You should also do, like, a big graduation ceremony where you change the handle.
Ali [00:00:59]: Just say it.
Philip [00:00:59]: For everybody.
Swyx [00:01:00]: You guys are good at ceremonies, clearly. We had a nice launch of the book, very successful. But before we get into all that, I wanna start off with a fun question for you. Okay, you’re an expert inference engineer. What happens when I send a long query, say two hundred thousand tokens into Baseten’s inference? What’s the process of query through GPU model routing, balancing, all that? What is all the stuff that we don’t think about?
Long Context Requests, KV Cache, and Cache-Aware Routing
Philip [00:01:26]: With a long query specifically, the first thing that I’m gonna ask is, “Have you sent me this query before, or at least part of it?” and I really hope you have, because it’s gonna be a lot easier for me and a lot cheaper for you. So the first thing that we’re gonna look at is some cache-aware routing, where we’re going to see, we probably have a number of instances, a number of replicas up serving whatever model you’re hitting. We want to send this one to something with, number one, available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of these two hundred thousand tokens. If you’re doing two hundred thousand tokens, it’s probably coding or a multi-turn agent or something where you would expect to have that cached. If you don’t, we’re gonna have to send it to a prefill worker. We’ve at least on certain models disaggregated prefill and decode, so you’re going to have one set of GPUs that’s solely going to process the input, create the KV cache, and get you your first token, and then that’s going to be passed over to a separate set of GPUs, which is going to run decode. We’re going to iteratively make those tokens. We’re probably going to have some speculator model in front of that. I’m going to assume that you’re doing coding, and because of that, our speculator model, which assumes you’re doing coding, is gonna have a high draft token acceptance rate. If I’m wrong and you’re asking me to summarize every Harry Potter book, it’s gonna be slower. And then we stream that output to you and account for it, charge you, a couple of pennies and say, “Hey, would you like to send another one?”
Swyx [00:03:04]: Except Ba
Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI
2026/07/28
There are roughly 100x more people who use code than who can write code. As code that “just works” becomes easier to generate, this group may be the biggest prize of all — if you can get the agentic interface right.
A key trend we have been tracking over at AINews is the absolute explosion in Codex usage this year, with MAU now up >10x from Jan 2026. Less than two weeks after their July 9th launch, OpenAI said ChatGPT Work and Codex had reached 10M users combined (as we cover in the pod, Codex now powers ChatGPT Work, so all ChatGPT Work users are now users of the Codex harness, even if they aren’t traditional engineers) — showing the early innings of what happens when you graduate from coding agents to knowledge work agents:
We’ve been calling out how coding agents are “breaking containment” to do everything else this year to power every other part of knowledge work - and it started with the org chart, with a major reorg last month that amounted to two of Codex’s most prominent leaders, Greg and Tibo, taking responsibility over product and ChatGPT specifically, completing a “Superapp” consolidation cycle first discussed in March.
With these updates Codex is no longer just a coding tool. In June, OpenAI said knowledge workers already accounting for roughly 20% of Codex’s user base and growing more than 3x as quickly as developers. A product dedicated for knowledge workers was being pulled out of the Codex team.
However, knowledge work has a different set of problems and environments than coding. For decades, knowledge work has been scattered across different primitives like documents for writing, spreadsheets for analysis, slide decks for communication, and specialized applications for everything else. ChatGPT Work now enables users to work across every primitive with agents. Instead of opening an application and manually operating its features, the user can describe an outcome and collaborates with an agent that can assemble the tools, context, and artifact needed to reach it.
From building no-code products at Airtable to leading Productivity Engineering at OpenAI, Akshay Nathan has spent much of his career trying to make the power of software accessible to people who do not write code. In this episode, Akshay joins swyx and Vibhu to unpack the launch of ChatGPT Work, why Codex unexpectedly took off among non-developers inside OpenAI, and the company’s broader plan to bring useful agents from software engineers to knowledge workers and eventually everyone.
We go deep on the shared agent harness behind Codex and ChatGPT Work, why OpenAI brought the experiences together without making them identical, and how persistent computers, artifacts, Sites, plugins, memory, and sub-agents are changing what people can delegate to AI. Akshay explains why some teams are replacing decks and spreadsheets with interactive websites, how agents can gather context across code, Slack, documents, and local files, and what OpenAI learned from personal-agent products like OpenClaw.
Side note: also don’t miss Abhihek’s sandbox track keynote at AIE, which now powers a lot of the sandboxing for ChatGPT Work… and yes was also broken by an unreleased OpenAI model in the recent HuggingFace incident.
Akshay also reflects on how AI is transforming product development itself: why more people will become generalists with a specialty, why ideas and taste become the bottlenecks when almost anyone can build, why LLMs still struggle to generate genuinely grounded new ideas, and why teams must distinguish increased motion from actual progress.
We discuss:
* Why Codex unexpectedly took off among non-developers inside OpenAI
* Why employees felt like using Codex gave them a new superpower
* The product insight that led OpenAI to build ChatGPT Work
* Why Codex and ChatGPT Work share the same underlying agent harness
* How their UX, Git visibility, artifacts, and sandboxing defaults differ
* Why OpenAI merged its agent experiences instead of building separate products
* How AI is blurring the boundaries between engineering, design, strategy, and operations
* Why OpenAI wants the default model configuration to work for most users
* When power users should use deeper reasoning, Ultra, or multi-agent modes
* Artifacts, agentic spreadsheets, and creating high-fidelity work products
* Why interactive Sites may replace decks and spreadsheets
* The challenge of designing a simple interface for an agent that can build almost anything
* Why users should retry tasks that models could not handle three or six months ago
* How AI can gather context for performance reviews without replacing human judgment
* The OpenAI automation that turns internal Slack and document activity into memes
* What reaching ten million ChatGPT Work and Codex users means for the product
* How OpenClaw inspired persistent environments, scheduled tasks, and personal agents
* Using ChatGPT for financial planning, budgeting, workouts, meals, and household management
* The design tradeoffs behind sub-agents and how much of their work users should see
* ChatGPT memory, Chronicle, and long-term context
* Why AI may make more people generalists with deep specialties
* Why ideas and taste become more important when almost anyone can build
* Why LLMs still struggle with the instruction “bring me new ideas”
* Measuring productivity through quality at-bats instead of commits, tokens, or pull requests
* The critical difference between AI-generated motion and meaningful progress
Akshay Nathan
* LinkedIn: https://www.linkedin.com/in/akshaynathan/
* X: https://x.com/akshaynathan_
Timestamps
00:00:00 Introduction and Bringing the Power of Code to Everyone
00:01:33 Joining OpenAI and Preserving a Startup Culture
00:02:40 What OpenAI Learned from Enterprise AI Adoption
00:05:28 Why OpenAI Built ChatGPT Work
00:07:17 Codex vs. ChatGPT Work and the Shared Agent Harness
00:12:07 Why OpenAI Merged Its Agent Experiences
00:16:24 Models, Reasoning Levels, and Choosing the Right Default
00:20:26 Artifacts, Agentic Spreadsheets, and Model–Product Collaboration
00:24:22 Why Sites Could Replace Decks and Spreadsheets
00:30:08 Designing an Agent That Can Build Almost Anything
00:34:28 From Developer Agents to Knowledge Work—and Everyone
00:36:07 Power-User Advice and AI-Assisted Performance Reviews
00:40:41 OpenAI’s Internal AI Memes and the Ten-Million-User Launch
00:44:39 OpenClaw, Personal Agents, and ChatGPT as an Operating System
00:50:24 Sub-Agents, Ultra Mode, and How Much Control Users Need
00:54:39 ChatGPT Memory, Personalization, and Chronicle
01:00:19 How AI Is Reshaping Product Development and Tech Roles
01:03:15 Ideas, Taste, and Why LLMs Struggle to Generate New Ideas
01:04:42 Measuring Productivity, Quality At-Bats, and Motion vs. Progress
Transcript
Introduction: Akshay Nathan, ChatGPT Work, and the No-Code Arc
Swyx [00:00:00]: We’re here in the studio with Akshay from OpenAI. Welcome.
Akshay Nathan [00:00:07]: Thank you.
Swyx [00:00:08]: And with our trusty co-host, Vibhu. So you recently launched ChatGPT Work. You lead Core Product Engineering. It’s been a long journey, into all this. I find it very interesting that you started with no code or low code, with Walrus and Airtable. And to some extent, ChatGPT Work is like the super app of super apps of, well, here is the ultimate no code. You just write a prompt.
Akshay Nathan [00:00:32]: Yeah. It’s funny how things come, full circle. I think for a long time in my career, I started my career working consumer fintech, but then after that, like, there’s this hypothesis that, the things that we were able to do with code, like, as engineers, like, if we could bring that to many more people in a more, accessible way, then that would be truly magical. We were working on a startup. It’s funny, like, before LLMs, before vision LLMs, on how to do automated testing with AI. It was just kinda jank, back then, but doing what we can, and then worked at Airtable for a while on the same thesis that, like, if we can bring a database or the primitives behind a database to people, that’d be really useful to them. But once LLMs came onto the scene, it became clear that, this was the missing piece, like, the missing technology required to, like, bring the magic of code to everyone without them having to know what’s going on underneath the hood. And so, like, I think this launch and a lot of the stuff that we’ve been up to is, like, the manifestation of that.
From Walrus and Airtable to OpenAI
Vibhu [00:01:33]: How was stuff when you joined? So you joined OpenAI 2023. Now we’ve got, so much more stuff, so ChatGPT, Codex app, ChatGPT Work. Have things changed?
Joining OpenAI and What Hasn’t Changed
Akshay Nathan [00:01:44]: I think the more interesting thing is how things haven’t changed. Like, one, I joined I remember when I joined, it was, like, five hundred people. One thing I was worried about was, like, I was looking for something, more early stage and, like, was it gonna feel startup enough? And I joined, and I was like, “This feels even more startup-y than I could ever imagine.” And, like, that really hasn’t changed even till now. I think the, like, level of, like, bottoms-up ambition and, like, the ability of anyone to, like, do anything or have an idea and ship it is really cool. But on the, like, mission side, I think what was really compelling to me is this mission of, bringing frontier intelligence to everyone. Like, building AGI and then bringing it to everyone. And, I think acknowledging back then that, like, that vision is gonna, not be a linear progression. Like, we’re probably gonna, like, try different products and have different things that succeed and don’t. But the vision has stayed the same, and the mission has stayed the same, and we’re starting to see the pieces, fall together, and that’s really cool.
Enterprise Lessons: No One-Size-Fits-All AI
Swyx [00:02:40]: You worke
Podcast reviews
Read Latent Space: The AI Engineer Podcast podcast reviews
4.6 out of 5
103 reviews
★★★★★
BGDem 2026/07/25
Data center locations
Can data centers be placed far away from neighborhoods?
Data center should be located where they won’t compete with communities for electricity and ...
★★☆☆☆
chief_rocka 2025/11/22
Insightful, but Hard to Follow Sans Context
The discussions on this podcast are potentially insightful, but after four episodes, I'm frustrated. The hosts and guests often launch into deep conve...
★★★★★
Willm50 2025/09/06
Immediate Impact
I love this podcast as I I pretty much walk away from every show with something I can apply in my work!
★☆☆☆☆
Thank T. 2025/08/31
The hosts feel like they want to be characters on HBO’s Silicon Valley
I initially subscribed to this podcast because there’s a lack of quality podcasts in the AI, ML, LLM, ML Ops, whatever you’d like to call it, space. U...
★★★★★
Jess Martin 2025/07/16
My main source for AI commentary
The pace and magnitude of change driven by AI engineering is massive. This is how I stay on top of things! Appreciate the in-depth analysis and the br...
★★★★★
seanjames_writes 2025/07/15
Best Applied Series
Its so far been the best way to keep up with everything happening in the AI space from a builder/applied engineering standpoint point.
Cuts through ...
★★★★☆
walterwhitenyc 2025/05/12
Very informative but please get rid of the robotic intro
See title
★☆☆☆☆
AI_Guy_00 2025/04/02
Huge egos
Man, these guys have clearly lost touch with reality. They had an entire podcast the other day where Dharmesh kept calling everyone not in tech “normi...
★★★★★
Watari202 2024/11/02
Thank you
Awesome intros 🎉
★★★★★
Eriklesher 2024/09/28
AI Engineer go-to
If you're an AI engineer or want to become one, this is the best resource for it, everyone being interviewed is actually building interesting stuff an...