AI Workspace Loading

We’re preparing your intelligent learning experience. Our AI systems are processing content, optimizing resources, and setting everything up for you.

Preparing Learning Paths...
AI Processing
Smart Automation
Learning Engine
Good things take a moment.

LearnLess.ai

LEARN LESS. UNDERSTAND MORE.
Intermediate7 min read

Attention

Attention lets a model weigh how relevant each part of the input is when producing each part of the output — the mechanism at the core of the transformer.

What Attention Does

Consider this sentence: "The animal didn't cross the road because it was tired." To understand what "it" refers to, you need to connect that word back to "animal" — even though several other words sit in between. Attention is the mechanism that lets a model make exactly that kind of connection: for every token, it decides how much to focus on every other token when building that token's representation.

"Self-attention" is the specific case where a sequence attends to itself — every token in the sentence can look at every other token in that same sentence, which is what happens inside a transformer.

The Mental Model: Query, Key, Value

A common way to build intuition for attention is the query, key, value framing, borrowed from search:

Query
What the current token is "asking" — roughly, "what information do I need from the rest of the sequence?"
Key
What every other token "advertises" about itself — roughly, "here is the kind of information I hold."
Value
The actual content a token contributes once it has been identified as relevant.

Each token’s query is compared against every other token’s key to produce a relevance score — the attention weight. Those weights determine how much of each token’s value gets blended into the current token’s new representation. In the example sentence, "it" would end up with a high attention weight toward "animal", which is how the model resolves what "it" is referring to.

Query

What This Token Asks

What information the current token needs.

matched against

Compare against Keys

Relevance Check

Matched against what every other token advertises.

produces

Attention Weights

Relevance Scores

How much to focus on each other token.

scale

Weighted Values

Blended Content

Each value scaled by its attention weight.

blended into

New Representation

Context-Aware

The token now carries relevant context from the sequence.

One token attending to the sequence

Why This Matters

Because attention lets every token connect to every other token directly, a model doesn't need to carry information forward step by step through a long chain — it can pull in the relevant word from anywhere in the sequence in a single step. That is a big part of why transformer-based models are good at handling longer, more complex context than earlier architectures.

Common Mistakes

  • Thinking attention is one single lookup

    In practice, transformers run many attention computations in parallel ("heads"), each potentially picking up on different kinds of relationships.

  • Assuming attention understands meaning the way a person does

    It is a learned mechanism for weighting relevance based on training data patterns, not comprehension in the human sense.

  • Confusing attention with the whole transformer

    Attention is the core mechanism, but a transformer layer also includes feed-forward processing and normalization around it.

Interview Question

How does attention work, conceptually?

Attention lets each token in a sequence decide how much to focus on every other token when building its own representation. A common way to describe it is with query, key, and value: each token produces a query, compares it against every other token’s key to get a relevance score, and then blends in the values of the tokens it scored as relevant. This lets a model connect related words even when they’re far apart in a sentence, without processing everything strictly in order.

What an interviewer may ask next

  • What is self-attention, specifically?
  • Why do transformers use multiple attention heads instead of one?
  • How does attention help with long-range dependencies compared to earlier architectures?

Explain It in 30 Seconds

Attention lets a model weigh how relevant each token in a sequence is to every other token when building its representation. Using the query, key, value framing: each token asks a question (query), compares it against what every other token offers (key), and pulls in the content (value) of the ones that matter most. That’s how a transformer resolves things like a pronoun referring back to a noun several words earlier.

On this page