Recording intelligenceAI-generated brief · check the source for context

The End of Frozen LLMs? (Google’s Hope Explained)

15:39 recording · AUTO · 1 speaker

Watch the original

Brief overview

Google's Nested Learning paper proposes HOPE, an architecture aiming at continual learning without catastrophic forgetting.

  1. Frozen weights, not intelligence, are the real ceilingAfter training ends models stop learning, and attempts to keep updating them hit catastrophic forgetting, where new learning overwrites what the model already knew.
  2. Different update frequencies act as a safety netBecause each Continuum Memory System module updates at a different time, information lost by one module still exists in slower modules that haven't updated yet and can be recovered.
  3. Depth is the sequence of learning updatesThe authors argue stacking static layers isn't what creates depth; HOPE makes depth explicit by having modules learn at different frequencies, creating a hierarchy across time.
Executive Summary AI
  • Today's large language models are trained once and then frozen: they adapt inside a conversation through in-context learning, but that learning disappears when the context does, so the model only ever experiences the immediate present.
  • Google's new paper, Nested Learning: The Illusion of Deep Learning Architecture, introduces an architecture called HOPE that the authors claim could rival transformers while enabling true continual learning, comparing current models to anterograde amnesia, where a person keeps old memories but can form no new long-term ones.
  • The paradigm takes two cues from the brain: neuroplasticity, illustrated by children who have parts of the brain removed to treat severe epilepsy yet later develop largely normal cognition, and brain oscillations, where fast waves support short-term adaptation and slow waves drive long-term consolidation.
  • HOPE is built from Neural Learning Modules, each with its own objective, its own learning rate and its own update frequency; fast modules update every token or few tokens while the slowest waits 16 million tokens, and the sequential stack of them, the Continuum Memory System, replaces the transformer's feed forward network so knowledge never changes everywhere at once.
  • On results, HOPE beats the transformer on needle-in-a-haystack retrieval (though Hope Attention still wins there because attention caches all past tokens uncompressed), achieves the best average performance on language modeling and common sense reasoning at both 760 million and 1.3 billion parameters, and outperforms existing continual learning approaches across multiple benchmarks.
Key Quote
“What if the biggest limitation of today's AI isn't intelligence, but memory?”
— Speaker
Key Quote
“In other words, the models only experience the immediate present.”
— Speaker
Key Quote
“It could rival transformers while enabling true continual learning without catastrophic forgetting.”
— Speaker