CHAPTER 11.1Chapter-1-Introduction1.Architecture Ingredients & Pre-Trained Weight MatricesTOKENIZATION1.2Chapter-2-Input-Processing-Tokenization1.Chapter 2.1 - Tokenizing Text2.Chapter 2.2 - Data Sampling3.Chapter 2.3 - Creating Token Embeddings4.Chapter 2.4 - Encoding Positional EmbeddingsEMBEDDINGS1.3Chapter-3-Input-Processing-and-Embeddings-End-to-End-Walkthrough1.Chapter 3.1 - Tokenization & Input Prompt2.Chapter 3.2 - Token Embedding Lookup (We Matrix)3.Chapter 3.3 - Positional Embeddings & Combined Input MatrixSELF ATTENTION1.4Chapter-4-Scaled-Dot-Attention-Core-Mechanics1.Chapter 4.1 - Dot Products2.Chapter 4.2 - Matrix Multiplication3.Chapter 4.3 - attention scores4.Chapter 4.4 - normalization5.Chapter 4.5 - context vector6.Chapter 4.6 - full matrix conversionSELF ATTENTION1.5Chapter-5-Single-Head-Self-Attention1.Chapter 5.1 - Weight Matrices2.Chapter 5.2 - Raw Attention Scores Scaling3.Chapter 5.3 - Masking, Softmax Context VectorSELF ATTENTION1.6Chapter-6-Multi-Head-Attention1.Chapter 6.1 - Multi-Head Self-Attention2.Chapter 6.2 - Output ProjectionCHAPTER 71.7Chapter-7-Pre-FFN-Cleanup1.Chapter 7.1 - 1st Residual Connection2.Chapter 7.2 - 1st Layer NormalizationCHAPTER 81.8Chapter-8-Feed-Forward-Network-FFN1.Chapter 8.1 - First Linear Layer2.Chapter 8.2 - GELU Activation Function3.Chapter 8.3 - Second Linear LayerCHAPTER 91.9Chapter-9-Post-FFN-Cleanup1.Chapter 9.1 - 2nd Residual Connection2.Chapter 9.2 - 2nd Layer NormalizationGENERATION1.10Chapter-10-Next-Word-Generation1.Chapter 10.1 - Logits to Probabilities2.Chapter 10.2 - Temperature & Top-K Sampling3.Chapter 10.3 - Autoregressive Generation Loop4.Chapter 10.4 - Complete Inference WalkthroughCHAPTER 111.11Chapter-11-Training1.Chapter 11.1 - Cross Entropy Loss2.Chapter 11.2 - Generating Text Batches3.Chapter 11.3 - Calculating the Batch Loss4.Chapter 11.4 - Backpropagation5.Chapter 11.5 - The Optimization Step6.Chapter 11.6 - Tensors7.Chapter 11.7 - Weight Optimization (Optimizers)8.Chapter 11.8 - Complete Training PipelineCHAPTER 121.12Chapter-12-Pretrained-Weights1.Chapter 12.1 - Saving Loading Weights2.Chapter 12.2 - Downloading GPT-2 Weights3.Chapter 12.3 - Mapping TF to PyTorch4.Chapter 12.4 - Complete Weight Loading PipelineCHAPTER 131.13Chapter-13-Finetuning1.Chapter 13.1 - Introduction to Finetuning2.Chapter 13.2 - Instruction Dataset3.Chapter 13.3 - Collate Function4.Chapter 13.4 - The Finetuning PipelineCHAPTER 141.14Chapter-14-Output refining1.Chapter 14.1 - Streaming text2.Chapter 14.2 - Reducing repetitive output3.Chapter 14.4 - Generating Response in Answer formatCHAPTER 151.15Chapter-15-Multi turn chatting1.Chapter 15.1 - Chatting2.Chapter-15.2---Chatting Dataset and loader3.Chapter 15.3 - Chat Dataset4.Chapter 15.4 - Collate functionCHAPTER 161.16Chapter-16-FineTuning Improvements1.Chapter 16.1 - FineTuning ImprovementsCHAPTER 171.17Chapter-17-Optional-Alternatives-Modern-Variants1.Chapter 17.1 - Activation Functions (ReLU, GELU, SwiGLU)2.Chapter 17.2 - Normalization (LayerNorm vs. RMSNorm)3.Chapter 17.3 - Positional Encodings (RoPE ALiBi)4.Chapter 17.4 - Attention Variants (GQA FlashAttention)5.Chapter 17.5 - Scaling Speedups (MoE KV Caching)
TOKENIZATION1.2Chapter-2-Input-Processing-Tokenization1.Chapter 2.1 - Tokenizing Text2.Chapter 2.2 - Data Sampling3.Chapter 2.3 - Creating Token Embeddings4.Chapter 2.4 - Encoding Positional Embeddings
EMBEDDINGS1.3Chapter-3-Input-Processing-and-Embeddings-End-to-End-Walkthrough1.Chapter 3.1 - Tokenization & Input Prompt2.Chapter 3.2 - Token Embedding Lookup (We Matrix)3.Chapter 3.3 - Positional Embeddings & Combined Input Matrix
SELF ATTENTION1.4Chapter-4-Scaled-Dot-Attention-Core-Mechanics1.Chapter 4.1 - Dot Products2.Chapter 4.2 - Matrix Multiplication3.Chapter 4.3 - attention scores4.Chapter 4.4 - normalization5.Chapter 4.5 - context vector6.Chapter 4.6 - full matrix conversion
SELF ATTENTION1.5Chapter-5-Single-Head-Self-Attention1.Chapter 5.1 - Weight Matrices2.Chapter 5.2 - Raw Attention Scores Scaling3.Chapter 5.3 - Masking, Softmax Context Vector
SELF ATTENTION1.6Chapter-6-Multi-Head-Attention1.Chapter 6.1 - Multi-Head Self-Attention2.Chapter 6.2 - Output Projection
CHAPTER 71.7Chapter-7-Pre-FFN-Cleanup1.Chapter 7.1 - 1st Residual Connection2.Chapter 7.2 - 1st Layer Normalization
CHAPTER 81.8Chapter-8-Feed-Forward-Network-FFN1.Chapter 8.1 - First Linear Layer2.Chapter 8.2 - GELU Activation Function3.Chapter 8.3 - Second Linear Layer
CHAPTER 91.9Chapter-9-Post-FFN-Cleanup1.Chapter 9.1 - 2nd Residual Connection2.Chapter 9.2 - 2nd Layer Normalization
GENERATION1.10Chapter-10-Next-Word-Generation1.Chapter 10.1 - Logits to Probabilities2.Chapter 10.2 - Temperature & Top-K Sampling3.Chapter 10.3 - Autoregressive Generation Loop4.Chapter 10.4 - Complete Inference Walkthrough
CHAPTER 111.11Chapter-11-Training1.Chapter 11.1 - Cross Entropy Loss2.Chapter 11.2 - Generating Text Batches3.Chapter 11.3 - Calculating the Batch Loss4.Chapter 11.4 - Backpropagation5.Chapter 11.5 - The Optimization Step6.Chapter 11.6 - Tensors7.Chapter 11.7 - Weight Optimization (Optimizers)8.Chapter 11.8 - Complete Training Pipeline
CHAPTER 121.12Chapter-12-Pretrained-Weights1.Chapter 12.1 - Saving Loading Weights2.Chapter 12.2 - Downloading GPT-2 Weights3.Chapter 12.3 - Mapping TF to PyTorch4.Chapter 12.4 - Complete Weight Loading Pipeline
CHAPTER 131.13Chapter-13-Finetuning1.Chapter 13.1 - Introduction to Finetuning2.Chapter 13.2 - Instruction Dataset3.Chapter 13.3 - Collate Function4.Chapter 13.4 - The Finetuning Pipeline
CHAPTER 141.14Chapter-14-Output refining1.Chapter 14.1 - Streaming text2.Chapter 14.2 - Reducing repetitive output3.Chapter 14.4 - Generating Response in Answer format
CHAPTER 151.15Chapter-15-Multi turn chatting1.Chapter 15.1 - Chatting2.Chapter-15.2---Chatting Dataset and loader3.Chapter 15.3 - Chat Dataset4.Chapter 15.4 - Collate function
CHAPTER 171.17Chapter-17-Optional-Alternatives-Modern-Variants1.Chapter 17.1 - Activation Functions (ReLU, GELU, SwiGLU)2.Chapter 17.2 - Normalization (LayerNorm vs. RMSNorm)3.Chapter 17.3 - Positional Encodings (RoPE ALiBi)4.Chapter 17.4 - Attention Variants (GQA FlashAttention)5.Chapter 17.5 - Scaling Speedups (MoE KV Caching)