Recording intelligenceAI-generated brief · check the source for context

Ego and Ego2 for Lossless Inference Acceleration

48:25 recording · AUTO · 1 speaker

Watch the original

Brief overview

EAGLE and EAGLE-2 make large language model inference up to four times faster without changing outputs.

  1. Predict the next feature, not the next tokenThe top-layer feature path is simple enough for a single transformer layer, while the embedding path needs a super large model.
  2. Feed the sampled token back to kill uncertaintyConcatenating the predicted feature with the embedding of the token just sampled pushed acceptance to roughly 0.8 and speed-up to three times.
  3. Dynamic trees beat static treesPosition one has the highest acceptance and position six the lowest, so EAGLE-2 prunes by draft confidence and reaches about 4.2 times speed-up.
Executive Summary AI
  • The speaker presents two papers from his group, EAGLE and EAGLE-2, whose goal is lossless inference acceleration for large language models, and starts by showing that vanilla autoregressive decoding generates tokens strictly one at a time while modern GPUs are built for parallel computing.
  • The proposed remedy is speculative sampling, a two-mode framework in which a tiny draft model guesses several tokens cheaply and the original large model then checks all of them in parallel, accepting a token when a uniform random value r is below min(1, p_t/q_t) and otherwise resampling from the normalised max(0, p - q) correction distribution.
  • EAGLE's core observation is that the evolution path at the feature level (f_how to f_can to f_I) is far simpler than at the embedding level, so instead of next-token prediction the draft model does next-feature prediction with a single trainable transformer layer on top of a frozen embedding layer, transformer stack and LM head.
  • To fix the remaining feature uncertainty, the predicted feature is concatenated with the word embedding of the token actually sampled in the previous round, lifting acceptance to roughly 0.8 and speed-up to roughly three times, with draft models of 0.24 billion parameters (3.4%) for a 7 billion model and about one billion (1.4%) for 70 billion, trainable on an RTX 3090 in one or two days.
  • EAGLE-2 replaces the static draft tree with a context-aware dynamic one, ranking nodes by draft-model confidence (shown to correlate linearly with the true acceptance rate) and keeping the top eight, which raises MT-bench speed-up from three times to about 4.2 times, reaches 4.6 average accepted tokens, and lets an RTX 3060 costing $600 hit 40 tokens per second — faster than vanilla decoding on a
Key Quote
“each token is generated one by one.”
— Speaker
Key Quote
“So we can see this is a sequential inference.”
— Speaker
Key Quote
“However, we also notice that the tiny model is very small.”
— Speaker