Attention
Attention lets a model weigh how relevant each part of the input is when producing each part of the output — the mechanism at the core of the transformer.
What Attention Does
Consider this sentence: "The animal didn't cross the road because it was tired." To understand what "it" refers to, you need to connect that word back to "animal" — even though several other words sit in between. Attention is the mechanism that lets a model make exactly that kind of connection: for every token, it decides how much to focus on every other token when building that token's representation.
"Self-attention" is the specific case where a sequence attends to itself — every token in the sentence can look at every other token in that same sentence, which is what happens inside a transformer.
The Mental Model: Query, Key, Value
A common way to build intuition for attention is the query, key, value framing, borrowed from search:
- Query
- What the current token is "asking" — roughly, "what information do I need from the rest of the sequence?"
- Key
- What every other token "advertises" about itself — roughly, "here is the kind of information I hold."
- Value
- The actual content a token contributes once it has been identified as relevant.
Each token’s query is compared against every other token’s key to produce a relevance score — the attention weight. Those weights determine how much of each token’s value gets blended into the current token’s new representation. In the example sentence, "it" would end up with a high attention weight toward "animal", which is how the model resolves what "it" is referring to.
Query
What This Token AsksWhat information the current token needs.
Compare against Keys
Relevance CheckMatched against what every other token advertises.
Attention Weights
Relevance ScoresHow much to focus on each other token.
Weighted Values
Blended ContentEach value scaled by its attention weight.
New Representation
Context-AwareThe token now carries relevant context from the sequence.
Why This Matters
Because attention lets every token connect to every other token directly, a model doesn't need to carry information forward step by step through a long chain — it can pull in the relevant word from anywhere in the sequence in a single step. That is a big part of why transformer-based models are good at handling longer, more complex context than earlier architectures.
Common Mistakes
Thinking attention is one single lookup
In practice, transformers run many attention computations in parallel ("heads"), each potentially picking up on different kinds of relationships.
Assuming attention understands meaning the way a person does
It is a learned mechanism for weighting relevance based on training data patterns, not comprehension in the human sense.
Confusing attention with the whole transformer
Attention is the core mechanism, but a transformer layer also includes feed-forward processing and normalization around it.
Interview Question
How does attention work, conceptually?
Attention lets each token in a sequence decide how much to focus on every other token when building its own representation. A common way to describe it is with query, key, and value: each token produces a query, compares it against every other token’s key to get a relevance score, and then blends in the values of the tokens it scored as relevant. This lets a model connect related words even when they’re far apart in a sentence, without processing everything strictly in order.
What an interviewer may ask next
- What is self-attention, specifically?
- Why do transformers use multiple attention heads instead of one?
- How does attention help with long-range dependencies compared to earlier architectures?
Explain It in 30 Seconds
Attention lets a model weigh how relevant each token in a sequence is to every other token when building its representation. Using the query, key, value framing: each token asks a question (query), compares it against what every other token offers (key), and pulls in the content (value) of the ones that matter most. That’s how a transformer resolves things like a pronoun referring back to a noun several words earlier.