← Back to Labs

LLM Context Windows & Attention Limits

Step through QK^T attention matrix calculation, RoPE positional encoding, KV cache allocation, and needle-in-a-haystack limits

Token x_m"retrieved"Q = x·WqK = x·WkV = x·Wvx_2im·θᵢR_m · QR_m · KRoPE Properties:• Rotates pairs of 2D embedding coordinates• Dot product depends ONLY on distance (m - n)• Eliminates explicit positional embeddings
STEP 1 OF 6

Token Vector Projection & RoPE Encoding

Tokens are mapped to dense vector representations, multiplied by Query, Key, and Value weights (Wq, Wk, Wv), and rotated in 2D vector pairs using Rotary Position Embedding (RoPE).

Arrow keys to navigate · R to reset

Tap dots to jump to any step

Read the full article →Take the quiz →