Chapter 17.4 - Attention Variants (GQA FlashAttention)
Overview
Multi-Head Attention is awesome, but it takes up way too much computer memory. Scientists invented GQA (Grouped Query Attention) and FlashAttention to make the AI's brain run 10x faster while using way less memory!
🎯 Why we do it
Rationale
In standard attention, every "Query" head gets its very own "Key" and "Value" head. This uses huge amounts of RAM. GQA says, "Hey, why don't 4 Query heads just share 1 Key and 1 Value head?" It's like carpooling! It saves a ton of space.
🛠️ How we do it
Methodology
FlashAttention is even cooler. It's a hardware trick. Instead of saving the attention scores to the GPU's slow memory, it does all the math instantly in the GPU's super-fast brain cache!
# Standard Attention:
# 8 Queries, 8 Keys, 8 Values (Takes up 24 spaces)
# GQA (Grouped Query Attention):
# 8 Queries, but they share 2 Keys and 2 Values! (Takes up 12 spaces)
print("Carpooling activated! Memory saved by 50%!")