Week 4: The Hidden Geometry of Data & Self-Supervised Learning
Stop assuming your data is perfect. Real world datasets are corrupted and neural networks fail spectacularly if initialized blindly. You will uncover the "dark matter" of Machine Learning by training on massive unlabeled datasets using Self-Supervised Learning. We bend your mind with high-dimensional geometry, unpack the architecture behind Apple's FaceID, and reveal how Kaiming initialization prevents deep networks from collapsing. Core Topics: Cross-Validation & Hyperparameter Tuning, Receptive Fields, Weight Initialization (Xavier/Glorot & Kaiming/He), The Manifold Hypothesis & High-Dimensional Geometry (The "Sea Urchin" Effect), Metric Learning (Siamese Networks, Triplet Loss, Hard Negative Mining), and the Holy Trinity of Self-Supervised Learning (SimCLR, CLIP, and Masked Autoencoders). Week 4 Lab: Real world data is garbage. We hand you a corrupted dataset so you can let your model's own mistakes flag the bad images. Purge the fakes in Voxel51, augment what is left, and crush your original score.
8 weeks · 59 lectures · free to watch
Start Week 4 →- 3.1Transfer Learning: Standing on the Shoulders of Giants19m
- 3.2Embeddings, Vector Search, and Retrieval Augmented Generation (RAG)13m
- 3.3Babysitting - The Learning Process23m
- 3.4Hyperparameter Optimization28m
- 3.5Deep Learning 102: Mastering the Convolutional Building Block33m
- 3.6Building Blocks - Convolution12m
- 3.7Building Block - Max Pool13m
- 3.8Cross Entropy vs Mean Square Loss4m
- 4.1Cross Validation and Hyperparameters9m
- 4.2Receptive Field of Deep Convolutional Networks4m
- 4.3Weight Initialization17m
- 4.4LeCun's Cake & The Hidden Geometry of Data13m
- 4.5Why High-Dimensional Space is a Lonely Place (The Sea Urchin)20m
- 4.6How FaceID Works: Siamese Networks & One-Shot Learning23m
- 4.7From Siamese to Triplet Networks: How Google Trained FaceNet14m
- 5.1The Deep Learning Story: From Cat Brains to AlphaFold9m
- 5.1-2Why Data, GPUs, and ReLU Changed Everything13m
- 5.2CNN Architectures: Evolution of Depth, Width, and Residuals22m
- 5.3From Fixed Inputs to Infinite Sequences: Introduction to RNNs13m
- 5.4From Vanishing Gradients to LSTMs: Solving the Memory Problem10m
- 5.5Sequence-to-Sequence: Encoder-Decoders and the Vanishing Gradient16m
- 6.1Breaking the Bottleneck: From RNNs to the Attention Revolution7m
- 6.2The Trinity of Transformers: Queries, Keys, and Values Explained24m
- 6.3Inside the Transformer: How Queries, Keys, and Values Create Meaning9m
- 6.4Assembling the Transformer: From Positional Encodings to GPT22m
- 6.5The Evolution of the Transformer: From "Attention Is All You Need" to Llama 39m
- 6.6Everything is a Transformer: Applying Attention to Images, Audio, and Robots9m
- 8.1Why do we need to post train LLMs2m
- 8.2Post Training LLMs29m
- 8.3Supervised Fine tuning14m
- 8.4RL based Fine Tuning16m
- 8.5Pitfalls and Advanced RL8m
- 8.6The Full Pipeline in Practice23m
- 8.7Mathematical Reasoning and Tool Calling - Notebook Walkthrough8m
- 8.8Mathematical Reasoning and Tool Calling - Notebook 2 Walkthrough18m