This opening lecture explains reasoning in LLMs and the four methods the course will teach from scratch.
Reasoning means showing the intermediate stepsHe defines reasoning narrowly as answering questions that require complex multi-step reasoning with intermediate steps, illustrated by the train travelling 60 miles per hour for 3 hours to cover 180 m
More compute at inference buys better answersJust as humans answer better with more time, allotting more computing resources during inference makes the model generate reasoning, and accuracy scales with test time compute.
Trust is why high-stakes clients want reasoning modelsFor a healthcare, pharma or finance application, clients trust the product far more when they can see the step-by-step process that led the LLM to its answer.
Executive SummaryAI
Dr. Rajat Dandekar opens his first YouTube course, Reasoning-based LLMs from Scratch, introducing himself as the twin brother of Dr. Raj, a B.Tech and M.Tech from IIT Madras with a PhD from Purdue University, and co-founder of Vijwara with Rajan Sridhar.
He frames reasoning through system 1, the fast intuitive thinking that he says accounts for about 95% of our thinking, and system 2, the slow, logical, lazy and indecisive thinking used for questions like choosing a movie.
He traces the shift from ChatGPT 3.5, released November 2022, which answered fast but hallucinated and showed no steps, to OpenAI's O1 released on September 20, 2024, and then to DeepSeek R1, which was open source, comparable in accuracy to O1, and achieved reasoning by pure reinforcement learning with its documented aha moment.
Using Roger's five tennis balls and two cans of three, he shows a regular LLM spending three tokens on the answer while a thinking model spends 23, which is how test time compute — the computing resources used during inference — scales model accuracy.
He closes by listing the four methods the course covers: inference time compute scaling, pure RL, supervised fine tuning with reinforcement learning, and pure supervised fine tuning with distillation, the last two with hands-on builds using GRPO on a 5.4 model and distillation.
Key Quote
“I am his twin brother, don't confuse me with him.”
— Dr. Rajat Dandekar
Key Quote
“System 1 helps us think immediately for an answer.”
— Dr. Rajat Dandekar
Key Quote
“AI is not made to think like humans. AI is made to answer like humans.”