EAGLE and EAGLE-2 make large language model inference up to four times faster without changing outputs.
Predict the next feature, not the next tokenThe top-layer feature path is simple enough for a single transformer layer, while the embedding path needs a super large model.
Feed the sampled token back to kill uncertaintyConcatenating the predicted feature with the embedding of the token just sampled pushed acceptance to roughly 0.8 and speed-up to three times.
Dynamic trees beat static treesPosition one has the highest acceptance and position six the lowest, so EAGLE-2 prunes by draft confidence and reaches about 4.2 times speed-up.
Executive SummaryAI
The speaker presents two papers from his group, EAGLE and EAGLE-2, whose goal is lossless inference acceleration for large language models, and starts by showing that vanilla autoregressive decoding generates tokens strictly one at a time while modern GPUs are built for parallel computing.
The proposed remedy is speculative sampling, a two-mode framework in which a tiny draft model guesses several tokens cheaply and the original large model then checks all of them in parallel, accepting a token when a uniform random value r is below min(1, p_t/q_t) and otherwise resampling from the normalised max(0, p - q) correction distribution.
EAGLE's core observation is that the evolution path at the feature level (f_how to f_can to f_I) is far simpler than at the embedding level, so instead of next-token prediction the draft model does next-feature prediction with a single trainable transformer layer on top of a frozen embedding layer, transformer stack and LM head.
To fix the remaining feature uncertainty, the predicted feature is concatenated with the word embedding of the token actually sampled in the previous round, lifting acceptance to roughly 0.8 and speed-up to roughly three times, with draft models of 0.24 billion parameters (3.4%) for a 7 billion model and about one billion (1.4%) for 70 billion, trainable on an RTX 3090 in one or two days.
EAGLE-2 replaces the static draft tree with a context-aware dynamic one, ranking nodes by draft-model confidence (shown to correlate linearly with the true acceptance rate) and keeping the top eight, which raises MT-bench speed-up from three times to about 4.2 times, reaches 4.6 average accepted tokens, and lets an RTX 3060 costing $600 hit 40 tokens per second — faster than vanilla decoding on a
Key Quote
“each token is generated one by one.”
— Speaker
Key Quote
“So we can see this is a sequential inference.”
— Speaker
Key Quote
“However, we also notice that the tiny model is very small.”