A transformer language model is built from scratch on Shakespeare to show how ChatGPT works underneath.
Attention is a data-dependent communication mechanismTokens emit queries and keys, dot-product them into affinities, and aggregate values from the nodes they find interesting.
Masking makes it a decoder blockSetting future positions to negative infinity before the softmax stops later tokens from leaking the answer to earlier ones.
Depth needs residuals and layer normSkip connections give gradients a superhighway back to the input, and pre-norm layer norm took the validation loss from 2.08 to 2.06.
Executive SummaryAI
The lecture opens with ChatGPT writing a haiku about AI and prosperity, and a mock breaking-news article about a leaf falling from a tree, to show that it is a probabilistic system that answers the same prompt differently each time.
The neural network underneath is the transformer, proposed in the 2017 paper Attention is All You Need — GPT stands for generatively pre-trained transformer — and the plan is to rewrite the speaker's nanoGPT repository, two files of about 300 lines, from an empty file.
The training data is the tiny Shakespeare dataset, a single file of about one megabyte and roughly one million characters, tokenized at the character level into a vocabulary of 65 symbols and split 90% train / 10% validation, then sampled in chunks of block size 8 stacked into batches.
The simplest baseline, a bigram model built on an embedding table, starts at a loss of 4.87 against the 4.17 expected from 65 uniform classes, and training with AdamW for 10,000 steps brings it to about 2.5 with still-garbage output.
Self-attention is then derived step by step — every token emits a query, a key and a value, affinities come from the dot product of queries and keys, scaled by one over the square root of the head size and masked with a lower-triangular matrix — and adding multi-head attention, a feed-forward layer, residual connections and layer norm drives the validation loss from 2.4 to 2.06.
Key Quote
“ChatGPT is a probabilistic system, and for any one prompt, it can give us multiple answers”
— Speaker
Key Quote
“Now the query vector, roughly speaking, is what am I looking for?”
— Speaker
Key Quote
“Attention supports arbitrary connectivity between nodes.”