Input text → one predicted next token. Amber steps are matrix multiplies. Tap any step; drag the layer count.
Inside one block: 6 of the 12 steps are matrix multiplies (amber) — the other half are softmax, ReLU, layer-norms and adds. About half and half by step count, yet the matmuls hold essentially all of the model's learned weights.