Laojin ChuhaiAI · GO GLOBAL
Technical ThinkersAndrej Karpathy · Feb 5, 2025 · 3:31:23

Deep Dive into LLMs like ChatGPT

Tip: use the player's CC button to enable or switch subtitles; English captions are available on these videos.

Why it matters

The 2025 'full edition': over three and a half hours covering the entire pipeline behind ChatGPT — data, tokenization, pretraining, post-training, inference, tool use, hallucinations and mitigations. A full level deeper than the 2023 one-hour talk.

Key takeaways

  • The full pipeline: raw internet data → tokenization → pretraining → supervised fine-tuning → RL post-training.
  • Hallucination is a feature, not a bug: models always produce the statistically most plausible answer.
  • Reasoning models think before answering, with RL rewarding good chains of thought.
  • Capabilities are 'Swiss cheese': hard problems solved, easy ones fumbled.
  • Practical ChatGPT: treat context as memory — start a fresh conversation when you should.

Original video

Speaker
Andrej Karpathy
Channel
Andrej Karpathy
Venue
A 3.5-hour deep dive on his own channel
Date · Duration
Feb 5, 2025 · 3:31:23

Deep Dive into LLMs like ChatGPT

Watch the original on YouTube