All-In with Chamath, Jason, Sacks & Friedberg · All-In Podcast, LLC
Feldman's half is an argument about architecture age. Every processor before Cerebras followed Moore's law, doubling roughly every 18 months; he claims his chip broke that trajectory and expects to be "way over 2X" in the next 18 months. His reasoning is structural rather than triumphal: a two-decade-old design like the GPU has to buy its gains from smaller geometry and the next fab node, while a new architecture still has unexploited knowledge about the work being presented to it. The demand thesis underneath is that reasoning is inference, and inference is where speed compounds — reasoning models consume enormous token counts internally, so a machine that is 15 times faster run for 24 hours yields what he describes as weeks or months of thinking. He reports a $25 billion backlog and says the buyers were ordering chips before the chips were finished. His framing for the excess is an AWS analogy: the credit-card era where every engineer signed up for everything, followed by enterprises learning to route hard problems to frontier models and ordinary ones somewhere cheaper.
The open-source stretch is the most concrete. Cerebras runs GLM, Kimi and the Qwen family alongside OpenAI's models, plus private models built by GlaxoSmithKline and by its UAE partner G42. Feldman's complaint is about supply, not ideology: he calls OpenAI's OSS 120B a good release and then points out that a company wanting to run open weights on-prem for HIPAA or FINRA reasons is choosing between that one model and Chinese ones. His read on why OpenAI and Amazon build their own silicon is that the hyperscalers learned dependency from Intel and the GPU makers learned it from a small customer base — the goal is not the fastest chip, only not being wholly dependent on someone else's. He also relays that Palo Alto Networks put a recent model against its own software, found bugs it did not know about, and stopped everything for six weeks to patch, which he offers as the argument for staged releases and government red-teaming rather than against them.
Rombach's half traces a line from compression to robotics. Latent diffusion, which he and his co-founders invented as PhD students in Munich, compresses images or video into an efficient representation and trains a transformer on that — the same principle as JPEG and MP3, rendered as a neural algorithm. Stable Diffusion was built on it, Flux after that, and the current direction is multimodal pre-training across image, video and audio combined with action prediction, so one model can generate a film or drive a robot. He is candid about the gap: today each robot has its own action representation and needs a few hours of fine-tuning data per task, and moving that work into context is still a research problem. On Scorsese, he describes the director iterating on images of an eastern European village to get a mental picture out of his head, and argues language is a lossy communication medium where a generated image is not. He declines the obvious pitch that the goal is generating whole movies, saying the interesting outputs come with a human iterating in the loop.
“early in an architecture, you have room to do much better than what was traditionally Moore's law” Andrew Feldman · at 13:15 —
“If they want to run open source right now, it's OSS 120B or Chinese models.” Andrew Feldman · at 18:28 —
Feldman is co-founder and CEO of Cerebras Systems, which builds wafer-scale inference chips and went public before this recording. Rombach is co-founder and CEO of Black Forest Labs, based in Freiburg and San Francisco; he co-invented latent diffusion as a PhD student in Munich, worked on Stable Diffusion, and now leads the team behind the open-weight Flux models.