Executive Summary
AI
- Lex Friedman frames the year through the DeepSeek moment of January 2025, when DeepSeek R1 reached near state-of-the-art performance with allegedly much less compute and for much cheaper.16:36
- Sebastian argues no company will hold technology nobody else has, because researchers rotate between labs, so the differentiating factor becomes budget and hardware constraints rather than proprietary ideas.18:10
- Nathan expects Chinese open-weight releases to continue for a few years and more open model builders throughout 2026 than in 2025, because US companies will not buy API subscriptions from Chinese firms for security reasons.21:20
- Sebastian says the architecture has barely moved since GPT-2 — mixture of experts, group query attention and RMSNorm are tweaks, not new architectures — so the real gains now come from mid-training, post-training and inference-time scaling.58:32
- Nathan puts DeepSeek's pre-training at about $5 million at cloud market rates and Olmo 3's cluster rental at about $2 million, against serving costs of billions, and expects $2,000 subscriptions after this year's $200 tier.1:07:07
Key Quote
“So winning is a very broad term.”
— Sebastian Rashka17:26
Key Quote
“But on the other side of things, there's a lot of ominous technology from China where there's way more labs than DeepSeq.”
— Nathan Lambert20:01
Key Quote
“And with code, what's nice is you can verify it.”
— Sebastian Rashka39:39