Self-Attention
Self-attention lets each position in a sequence weigh every other position in the same sequence to build context-aware representations.
Prerequisites
Attention Applied Within One Sequence
The Attention lesson explains the general idea: weighing how relevant each part of the input is when producing each part of the output. Self-attention is that same mechanism applied within a single sequence — every word looks at every other word in the same sentence to figure out which ones matter for understanding it, rather than looking at a separate input and output.
Key Idea
Self-attention is why transformers can resolve something like which noun a pronoun refers to — each word can directly attend to any other word in the sequence, no matter how far apart they are.
A Concrete Example
In the sentence "The trophy didn't fit in the suitcase because it was too big," the word "it" is ambiguous on its own — it could refer to the trophy or the suitcase. Self-attention lets the model weigh the relationship between "it" and both candidate nouns, using the rest of the sentence as context, to resolve which one "it" most likely refers to.
"it" (current word)
Ambiguous AloneCould refer to more than one noun on its own.
Weighs relevance to every other word
Same SequenceNot a separate input and output.
"trophy": high relevance
Likely ReferentScored higher given the sentence context.
"suitcase": lower relevance
Less LikelyStill considered, but scored lower.
Context-aware representation of "it"
Resolved MeaningNow carries the resolved reference.
Why This Matters for Transformers
- Every word's representation becomes context-aware — the vector for "it" in one sentence differs from the vector for "it" in a sentence where the ambiguity resolves differently.
- Self-attention processes all positions in parallel, which is a major reason transformers train faster than older sequence models that had to process one position at a time in order.
- This same mechanism, computed multiple times in parallel with different learned weightings — called multi-head attention — lets a transformer track several kinds of relationships between words simultaneously.
Common Mistakes
Confusing self-attention with attention in general
Self-attention is specifically attention applied within one sequence — attention more broadly can also relate two different sequences, like an encoder and a decoder.
Assuming self-attention only looks at nearby words
Self-attention can directly relate any two positions in a sequence regardless of distance, which is a key advantage over older sequence models.
Overcomplicating the explanation with matrix math
The core idea — every word weighing every other word's relevance — can be understood without walking through the underlying linear algebra.
Interview Question
What is self-attention, and how does it help a model resolve something like pronoun ambiguity?
Self-attention is the attention mechanism applied within a single sequence — every word weighs how relevant every other word in the same sequence is, rather than relating a separate input and output. That's what lets a transformer resolve something like which noun a pronoun refers to: the word "it" can directly attend to candidate nouns elsewhere in the sentence and use the surrounding context to figure out which one is more relevant, regardless of how far apart they are in the sequence. It also processes all positions in parallel, which is a major reason transformers train faster than older sequence models that had to process one position at a time in order.
What an interviewer may ask next
- How is self-attention different from attention in general?
- Why can self-attention relate two words that are far apart in a sentence just as easily as two words next to each other?
- How does self-attention contribute to transformers being faster to train than older sequence models?
Explain It in 30 Seconds
Self-attention is attention applied within a single sequence — every word weighs how relevant every other word in the same sentence is to build a context-aware representation. This is what lets a transformer resolve ambiguity like which noun a pronoun refers to, since any two positions can directly relate to each other regardless of distance. It also runs in parallel across all positions, which is a big reason transformers train faster than older sequence models.