Recording intelligenceAI-generated brief · check the source for context

Let's build GPT: from scratch, in code, spelled out.

1:56:20 recording · AUTO · 1 speaker

Watch the original

Brief overview

A transformer language model is built from scratch on Shakespeare to show how ChatGPT works underneath.

  1. Attention is a data-dependent communication mechanismTokens emit queries and keys, dot-product them into affinities, and aggregate values from the nodes they find interesting.
  2. Masking makes it a decoder blockSetting future positions to negative infinity before the softmax stops later tokens from leaking the answer to earlier ones.
  3. Depth needs residuals and layer normSkip connections give gradients a superhighway back to the input, and pre-norm layer norm took the validation loss from 2.08 to 2.06.
Executive Summary AI
  • The lecture opens with ChatGPT writing a haiku about AI and prosperity, and a mock breaking-news article about a leaf falling from a tree, to show that it is a probabilistic system that answers the same prompt differently each time.
  • The neural network underneath is the transformer, proposed in the 2017 paper Attention is All You Need — GPT stands for generatively pre-trained transformer — and the plan is to rewrite the speaker's nanoGPT repository, two files of about 300 lines, from an empty file.
  • The training data is the tiny Shakespeare dataset, a single file of about one megabyte and roughly one million characters, tokenized at the character level into a vocabulary of 65 symbols and split 90% train / 10% validation, then sampled in chunks of block size 8 stacked into batches.
  • The simplest baseline, a bigram model built on an embedding table, starts at a loss of 4.87 against the 4.17 expected from 65 uniform classes, and training with AdamW for 10,000 steps brings it to about 2.5 with still-garbage output.
  • Self-attention is then derived step by step — every token emits a query, a key and a value, affinities come from the dot product of queries and keys, scaled by one over the square root of the head size and masked with a lower-triangular matrix — and adding multi-head attention, a feed-forward layer, residual connections and layer norm drives the validation loss from 2.4 to 2.06.
Key Quote
“ChatGPT is a probabilistic system, and for any one prompt, it can give us multiple answers”
— Speaker
Key Quote
“Now the query vector, roughly speaking, is what am I looking for?”
— Speaker
Key Quote
“Attention supports arbitrary connectivity between nodes.”
— Speaker