Week 6: The Attention Revolution & Building Transformers
Forget everything you know about sequential loops. RNNs hit a massive wall of catastrophic memory loss and slow training speeds. We dismantle that bottleneck and step into the architecture that took over the world: the Transformer. You will learn how "Attention Is All You Need" replaced recurrence with a highly efficient database lookup. We dissect the modern Transformer block-by-block and see how it powers everything from Llama 3 to modern robotics. Core Topics: The Self-Attention Mechanism, Queries, Keys, and Values (QKV), Multi-Head Attention, Positional Encodings (Sinusoidal vs. RoPE), Feed-Forward Networks (SwiGLU), Modern Optimizations (RMSNorm, KV Cache, Grouped Query Attention), and Cross-Domain Transformers (ViT, AST, VLA). Week 6 Lab: Rip out last week's RNN and upgrade to a Transformer Decoder. Build Scaled Dot-Product Attention, Multi-Head matrices, and Positional Encodings entirely by hand to see how modern LLMs actually tick.
8 weeks · 59 lectures · free to watch
Start Week 6 →- 3.1Transfer Learning: Standing on the Shoulders of Giants19m
- 3.2Embeddings, Vector Search, and Retrieval Augmented Generation (RAG)13m
- 3.3Babysitting - The Learning Process23m
- 3.4Hyperparameter Optimization28m
- 3.5Deep Learning 102: Mastering the Convolutional Building Block33m
- 3.6Building Blocks - Convolution12m
- 3.7Building Block - Max Pool13m
- 3.8Cross Entropy vs Mean Square Loss4m
- 4.1Cross Validation and Hyperparameters9m
- 4.2Receptive Field of Deep Convolutional Networks4m
- 4.3Weight Initialization17m
- 4.4LeCun's Cake & The Hidden Geometry of Data13m
- 4.5Why High-Dimensional Space is a Lonely Place (The Sea Urchin)20m
- 4.6How FaceID Works: Siamese Networks & One-Shot Learning23m
- 4.7From Siamese to Triplet Networks: How Google Trained FaceNet14m
- 5.1The Deep Learning Story: From Cat Brains to AlphaFold9m
- 5.1-2Why Data, GPUs, and ReLU Changed Everything13m
- 5.2CNN Architectures: Evolution of Depth, Width, and Residuals22m
- 5.3From Fixed Inputs to Infinite Sequences: Introduction to RNNs13m
- 5.4From Vanishing Gradients to LSTMs: Solving the Memory Problem10m
- 5.5Sequence-to-Sequence: Encoder-Decoders and the Vanishing Gradient16m
- 6.1Breaking the Bottleneck: From RNNs to the Attention Revolution7m
- 6.2The Trinity of Transformers: Queries, Keys, and Values Explained24m
- 6.3Inside the Transformer: How Queries, Keys, and Values Create Meaning9m
- 6.4Assembling the Transformer: From Positional Encodings to GPT22m
- 6.5The Evolution of the Transformer: From "Attention Is All You Need" to Llama 39m
- 6.6Everything is a Transformer: Applying Attention to Images, Audio, and Robots9m
- 8.1Why do we need to post train LLMs2m
- 8.2Post Training LLMs29m
- 8.3Supervised Fine tuning14m
- 8.4RL based Fine Tuning16m
- 8.5Pitfalls and Advanced RL8m
- 8.6The Full Pipeline in Practice23m
- 8.7Mathematical Reasoning and Tool Calling - Notebook Walkthrough8m
- 8.8Mathematical Reasoning and Tool Calling - Notebook 2 Walkthrough18m