One sentence, two attention masks. Rows are query positions; columns are key positions. The grid shows permitted connections, not learned attention weights.
Click a row — a word — to see exactly which words it’s allowed to look at.
Each cell represents a query–key pair. A causal mask sets future-position scores to −∞, giving them zero weight after softmax. Bidirectional attention also permits later positions. Both architectures produce contextual representations; this diagram isolates which context is available.