Chapter 17.1 - Activation Functions (ReLU, GELU, SwiGLU)
Upgrading the bouncers (ReLU, GELU, SwiGLU)
Chapter 17.2 - Normalization (LayerNorm vs. RMSNorm)
Faster showers! (LayerNorm vs RMSNorm)
Chapter 17.3 - Positional Encodings (RoPE ALiBi)
Knowing where words belong (RoPE, ALiBi)
Chapter 17.4 - Attention Variants (GQA FlashAttention)
Sharing the brain! (GQA, FlashAttention)
Chapter 17.5 - Scaling Speedups (MoE KV Caching)
Super speed! (MoE, KV Caching)