Chapter-1-Introduction○Architecture Ingredients & Pre-Trained Weight MatricesChapter-2-Input-Processing-Tokenization○Chapter 2.1 - Tokenizing Text○Chapter 2.2 - Data Sampling○Chapter 2.3 - Creating Token Embeddings○Chapter 2.4 - Encoding Positional EmbeddingsChapter-3-Input-Processing-and-Embeddings-End-to-End-Walkthrough○Chapter 3.1 - Tokenization & Input Prompt○Chapter 3.2 - Token Embedding Lookup (We Matrix)○Chapter 3.3 - Positional Embeddings & Combined Input MatrixChapter-4-Scaled-Dot-Attention-Core-Mechanics○Chapter 4.1 - Dot Products○Chapter 4.2 - Matrix Multiplication○Chapter 4.3 - attention scores○Chapter 4.4 - normalization○Chapter 4.5 - context vector○Chapter 4.6 - full matrix conversionChapter-5-Single-Head-Self-Attention○Chapter 5.1 - Weight Matrices○Chapter 5.2 - Raw Attention Scores Scaling○Chapter 5.3 - Masking, Softmax Context VectorChapter-6-Multi-Head-Attention○Chapter 6.1 - Multi-Head Self-Attention○Chapter 6.2 - Output ProjectionChapter-7-Pre-FFN-Cleanup○Chapter 7.1 - 1st Residual Connection○Chapter 7.2 - 1st Layer NormalizationChapter-8-Feed-Forward-Network-FFN○Chapter 8.1 - First Linear Layer○Chapter 8.2 - GELU Activation Function○Chapter 8.3 - Second Linear LayerChapter-9-Post-FFN-Cleanup○Chapter 9.1 - 2nd Residual Connection○Chapter 9.2 - 2nd Layer NormalizationChapter-10-Next-Word-Generation○Chapter 10.1 - Logits to Probabilities○Chapter 10.2 - Temperature & Top-K Sampling○Chapter 10.3 - Autoregressive Generation Loop○Chapter 10.4 - Complete Inference WalkthroughChapter-11-Training○Chapter 11.1 - Cross Entropy Loss○Chapter 11.2 - Generating Text Batches○Chapter 11.3 - Calculating the Batch Loss○Chapter 11.4 - Backpropagation○Chapter 11.5 - The Optimization Step○Chapter 11.6 - Tensors○Chapter 11.7 - Weight Optimization (Optimizers)○Chapter 11.8 - Complete Training PipelineChapter-12-Pretrained-Weights○Chapter 12.1 - Saving Loading Weights○Chapter 12.2 - Downloading GPT-2 Weights○Chapter 12.3 - Mapping TF to PyTorch○Chapter 12.4 - Complete Weight Loading PipelineChapter-13-Finetuning○Chapter 13.1 - Introduction to Finetuning○Chapter 13.2 - Instruction Dataset○Chapter 13.3 - Collate Function○Chapter 13.4 - The Finetuning PipelineChapter-14-Output refining○Chapter 14.1 - Streaming text○Chapter 14.2 - Reducing repetitive output○Chapter 14.4 - Generating Response in Answer formatChapter-15-Multi turn chatting○Chapter 15.1 - Chatting○Chapter-15.2---Chatting Dataset and loader○Chapter 15.3 - Chat Dataset○Chapter 15.4 - Collate functionChapter-16-FineTuning Improvements○Chapter 16.1 - FineTuning ImprovementsChapter-17-Optional-Alternatives-Modern-Variants○Chapter 17.1 - Activation Functions (ReLU, GELU, SwiGLU)○Chapter 17.2 - Normalization (LayerNorm vs. RMSNorm)○Chapter 17.3 - Positional Encodings (RoPE ALiBi)○Chapter 17.4 - Attention Variants (GQA FlashAttention)○Chapter 17.5 - Scaling Speedups (MoE KV Caching)
Chapter-2-Input-Processing-Tokenization○Chapter 2.1 - Tokenizing Text○Chapter 2.2 - Data Sampling○Chapter 2.3 - Creating Token Embeddings○Chapter 2.4 - Encoding Positional Embeddings
Chapter-3-Input-Processing-and-Embeddings-End-to-End-Walkthrough○Chapter 3.1 - Tokenization & Input Prompt○Chapter 3.2 - Token Embedding Lookup (We Matrix)○Chapter 3.3 - Positional Embeddings & Combined Input Matrix
Chapter-4-Scaled-Dot-Attention-Core-Mechanics○Chapter 4.1 - Dot Products○Chapter 4.2 - Matrix Multiplication○Chapter 4.3 - attention scores○Chapter 4.4 - normalization○Chapter 4.5 - context vector○Chapter 4.6 - full matrix conversion
Chapter-5-Single-Head-Self-Attention○Chapter 5.1 - Weight Matrices○Chapter 5.2 - Raw Attention Scores Scaling○Chapter 5.3 - Masking, Softmax Context Vector
Chapter-6-Multi-Head-Attention○Chapter 6.1 - Multi-Head Self-Attention○Chapter 6.2 - Output Projection
Chapter-7-Pre-FFN-Cleanup○Chapter 7.1 - 1st Residual Connection○Chapter 7.2 - 1st Layer Normalization
Chapter-8-Feed-Forward-Network-FFN○Chapter 8.1 - First Linear Layer○Chapter 8.2 - GELU Activation Function○Chapter 8.3 - Second Linear Layer
Chapter-9-Post-FFN-Cleanup○Chapter 9.1 - 2nd Residual Connection○Chapter 9.2 - 2nd Layer Normalization
Chapter-10-Next-Word-Generation○Chapter 10.1 - Logits to Probabilities○Chapter 10.2 - Temperature & Top-K Sampling○Chapter 10.3 - Autoregressive Generation Loop○Chapter 10.4 - Complete Inference Walkthrough
Chapter-11-Training○Chapter 11.1 - Cross Entropy Loss○Chapter 11.2 - Generating Text Batches○Chapter 11.3 - Calculating the Batch Loss○Chapter 11.4 - Backpropagation○Chapter 11.5 - The Optimization Step○Chapter 11.6 - Tensors○Chapter 11.7 - Weight Optimization (Optimizers)○Chapter 11.8 - Complete Training Pipeline
Chapter-12-Pretrained-Weights○Chapter 12.1 - Saving Loading Weights○Chapter 12.2 - Downloading GPT-2 Weights○Chapter 12.3 - Mapping TF to PyTorch○Chapter 12.4 - Complete Weight Loading Pipeline
Chapter-13-Finetuning○Chapter 13.1 - Introduction to Finetuning○Chapter 13.2 - Instruction Dataset○Chapter 13.3 - Collate Function○Chapter 13.4 - The Finetuning Pipeline
Chapter-14-Output refining○Chapter 14.1 - Streaming text○Chapter 14.2 - Reducing repetitive output○Chapter 14.4 - Generating Response in Answer format
Chapter-15-Multi turn chatting○Chapter 15.1 - Chatting○Chapter-15.2---Chatting Dataset and loader○Chapter 15.3 - Chat Dataset○Chapter 15.4 - Collate function
Chapter-17-Optional-Alternatives-Modern-Variants○Chapter 17.1 - Activation Functions (ReLU, GELU, SwiGLU)○Chapter 17.2 - Normalization (LayerNorm vs. RMSNorm)○Chapter 17.3 - Positional Encodings (RoPE ALiBi)○Chapter 17.4 - Attention Variants (GQA FlashAttention)○Chapter 17.5 - Scaling Speedups (MoE KV Caching)