📄Chapter 0.1 - LLM Architecture ExplorerExpand each section to reveal the structure inside itChapter-1-Tokenizer○Chapter 1.1 BPE○Chapter 1.2 WordPieceChapter-2-Token Embeddings○Chapter 2.1 Learned EmbeddingsChapter-3-Positional Encoding○Chapter 3.1 Sinusoidal○Chapter 3.2 RoPE○Chapter 3.3 ALiBiChapter-4-Layer Normalization○Chapter 4.1 Standard LayerNorm○Chapter 4.2 RMSNormChapter-5-Self Attention○Chapter 5.1 Multi Head Attention○Chapter 5.2 Grouped Query Attention○Chapter 5.3 Multi Query AttentionChapter-6-Residual Connections○Chapter 8.1 Pre Norm vs Post NormChapter-7-Activation Functions○Chapter 6.1 ReLU○Chapter 6.2 GELU○Chapter 6.3 SwiGLUChapter-8-Feed Forward Network○Chapter 7.1 Standard FFNChapter-9-Output and Decoding○Chapter 9.1 Vocabulary Projection○Chapter 9.2 Softmax○Chapter 9.3 Greedy Search○Chapter 9.4 Nucleus Sampling
Chapter-5-Self Attention○Chapter 5.1 Multi Head Attention○Chapter 5.2 Grouped Query Attention○Chapter 5.3 Multi Query Attention
Chapter-9-Output and Decoding○Chapter 9.1 Vocabulary Projection○Chapter 9.2 Softmax○Chapter 9.3 Greedy Search○Chapter 9.4 Nucleus Sampling