EMBEDDINGS1.1Ch-1 Text Representations & Embeddings1.Tokenization Techniques: How AI Reads2.Word2Vec: Turning Words into Math3.Negative Sampling: A Shortcut to Speed4.GloVe & Matrix Factorization: The Giant Spreadsheet5.The PyTorch Embedding LayerCH 21.2Ch-2 Order & Position Injection1.Why Order Matters2.Implicit Order: The Memory Loop3.Absolute Positional Encodings: The Page Numbers4.Relative & Rotary Encodings (RoPE)CH 31.3Ch-3 Sequential Recurrence & BPTT1.The Vanilla RNN Cell: The Blender2.Unrolling Over Time: The Comic Strip3.BPTT: Time Traveling Mistakes4.The Telephone Game: Vanishing GradientsCH 41.4Ch-4 Gated Linear Highways1.LSTM Architecture: The VIP Highway2.LSTM Gates: The Traffic Cops3.GRU: The 2-in-1 Shampoo4.Bidirectional RNNs: Reading the Future5.PyTorch ImplementationSELF ATTENTION1.5Ch-5 Dynamic Alignment & Classical Attention1.The Seq2Seq Bottleneck2.Bahdanau Attention: The Flashlight3.Luong Attention: The Puzzle Pieces4.Decoding Strategies: The MazeSELF ATTENTION1.6Ch-6 Self-Attention & Multi-Head Projections1.Self-Attention: The Cocktail Party2.Q, K, V: The Library Search3.Multi-Head Attention: The Detectives4.Building Attention in PyTorchCH 71.7Ch-7 Information Visibility & Masking1.Bidirectional Context: The Open Book Test2.Causal Masking: The Blindfold3.Prefix & Span MaskingCH 81.8Ch-8 Deep Stacking, Normalization & Feed-Forward1.Residual Streams: The Express Elevator2.LayerNorm: The Volume Knob3.The Feed-Forward Network4.KV Caching: The Ultimate Speed Hack
EMBEDDINGS1.1Ch-1 Text Representations & Embeddings1.Tokenization Techniques: How AI Reads2.Word2Vec: Turning Words into Math3.Negative Sampling: A Shortcut to Speed4.GloVe & Matrix Factorization: The Giant Spreadsheet5.The PyTorch Embedding Layer
CH 21.2Ch-2 Order & Position Injection1.Why Order Matters2.Implicit Order: The Memory Loop3.Absolute Positional Encodings: The Page Numbers4.Relative & Rotary Encodings (RoPE)
CH 31.3Ch-3 Sequential Recurrence & BPTT1.The Vanilla RNN Cell: The Blender2.Unrolling Over Time: The Comic Strip3.BPTT: Time Traveling Mistakes4.The Telephone Game: Vanishing Gradients
CH 41.4Ch-4 Gated Linear Highways1.LSTM Architecture: The VIP Highway2.LSTM Gates: The Traffic Cops3.GRU: The 2-in-1 Shampoo4.Bidirectional RNNs: Reading the Future5.PyTorch Implementation
SELF ATTENTION1.5Ch-5 Dynamic Alignment & Classical Attention1.The Seq2Seq Bottleneck2.Bahdanau Attention: The Flashlight3.Luong Attention: The Puzzle Pieces4.Decoding Strategies: The Maze
SELF ATTENTION1.6Ch-6 Self-Attention & Multi-Head Projections1.Self-Attention: The Cocktail Party2.Q, K, V: The Library Search3.Multi-Head Attention: The Detectives4.Building Attention in PyTorch
CH 71.7Ch-7 Information Visibility & Masking1.Bidirectional Context: The Open Book Test2.Causal Masking: The Blindfold3.Prefix & Span Masking
CH 81.8Ch-8 Deep Stacking, Normalization & Feed-Forward1.Residual Streams: The Express Elevator2.LayerNorm: The Volume Knob3.The Feed-Forward Network4.KV Caching: The Ultimate Speed Hack