Laojin ChuhaiAI · GO GLOBAL
AI News · ClassicsAndrej Karpathy · Feb 5, 2025 · 3:31:23

Deep Dive into LLMs like ChatGPT

Tip: use the player's CC button to enable or switch subtitles; English captions are available on these videos.

Why it matters

The 2025 'full edition': over three and a half hours Karpathy walks the entire pipeline behind products like ChatGPT — data, tokenization, pretraining, post-training, inference, tool use, hallucinations and mitigations. A full level deeper than the 2023 one-hour talk, now covering reasoning ('thinking') models and RL post-training. Block out an afternoon; afterwards there's no generation gap between you and your engineers.

Key takeaways

  • The full pipeline: raw internet data → tokenization → pretraining → supervised fine-tuning → RL post-training.
  • Hallucination is a feature, not a bug: models always produce the statistically most plausible answer; tools and retrieval are the fix.
  • Reasoning models think before answering, with reinforcement learning rewarding good chains of thought.
  • Capabilities are 'Swiss cheese': hard problems solved, easy ones fumbled — never mistake a demo for production.
  • Practical ChatGPT use: treat context as memory — start a fresh conversation when you should.

Original video

Speaker
Andrej Karpathy
Channel
Andrej Karpathy
Venue
A 3.5-hour general-audience deep dive on his own channel
Date · Duration
Feb 5, 2025 · 3:31:23

Deep Dive into LLMs like ChatGPT

Watch the original on YouTube