Technical ThinkersAndrej Karpathy · Nov 22, 2023 · 59:48
Intro to Large Language Models
People:Andrej Karpathy
Tip: use the player's CC button to enable or switch subtitles; English captions are available on these videos.
Why it matters
Karpathy's first public lecture after leaving Tesla: one hour on what LLMs are, how they're trained, and where they're going. No math required. Tens of millions of views and the source text for countless explainers.
Key takeaways
- An LLM is a lossy compression of internet text into weights, replayed by predicting the next token.
- Two training stages: pretraining yields the base model; fine-tuning turns it into an assistant.
- Scaling laws: performance improves smoothly with compute/data/parameters — and is predictable.
- The LLM is the kernel of a new OS: the context window is memory, tools are peripherals.
- Security matters early: prompt injection, jailbreaks and data poisoning are real attack surfaces.
Original video
- Speaker
- Andrej Karpathy
- Channel
- Andrej Karpathy
- Venue
- Public lecture on his own channel
- Date · Duration
- Nov 22, 2023 · 59:48
[1hr Talk] Intro to Large Language Models
Watch the original on YouTube