Positional Encoding
Positional encoding injects information about token order into a transformer, which otherwise has no built-in sense of sequence.
Prerequisites
The Problem It Solves
Self-attention looks at every token in relation to every other token, all at once, in parallel — which is great for speed, but it means the mechanism has no inherent sense of which token came first. Without extra information, "the dog chased the cat" and "the cat chased the dog" would look identical to a pure self-attention computation, since it's the same set of words. Positional encoding fixes this by adding information about each token's position directly into its representation.
Key Idea
Positional encoding is what lets a transformer process a sequence in parallel while still knowing which word came first, second, third, and so on.
How It Works
Token Embedding
No Order InfoA token's meaning, without knowing its position.
Positional Encoding
Adds PositionEncodes first, second, third, and so on.
Combined Representation
Meaning + OrderWhat actually enters the attention layers.
Self-Attention Layers
Order-AwareCan now tell "dog chased cat" from the reverse.
A positional encoding is combined with each token's regular embedding before the sequence enters the transformer's attention layers. The exact mathematical scheme varies — some use fixed patterns based on sine and cosine functions, others use positions the model learns during training — but the goal is always the same: give the model a consistent, distinguishable signal for "this token is at position 1, this one is at position 2," and so on.
Common Mistakes
Assuming a transformer inherently understands word order
Without positional encoding, self-attention alone treats a sequence as an unordered set of tokens — order has to be explicitly added.
Assuming positional encoding is only relevant to short sequences
How well a positional encoding scheme generalizes to longer sequences than it saw during training is actually a real, active engineering concern, closely tied to a model's effective context window.
Overcomplicating the explanation with the exact mathematical formula
The core idea — giving each position a distinguishable signal — matters more for understanding transformers practically than memorizing the specific sine/cosine formula.
Interview Question
Why do transformers need positional encoding, and what problem does it solve?
Self-attention looks at every token in relation to every other token in parallel, which gives it no inherent sense of order — without extra information, a sentence and its word-scrambled version would look identical to the attention mechanism, since it's processing the same set of tokens. Positional encoding solves this by adding information about each token's position directly into its representation before it enters the attention layers, so the model can distinguish "first token" from "second token" and so on. This is what lets a transformer process an entire sequence in parallel — for speed — while still understanding word order, which matters for practically any language task.
What an interviewer may ask next
- What would happen to a transformer's understanding of a sentence without positional encoding?
- Why does processing tokens in parallel create the need for positional encoding in the first place?
- How does positional encoding relate to a model's context window?
Explain It in 30 Seconds
Self-attention processes all tokens in parallel with no inherent sense of order, so positional encoding adds information about each token's position directly into its representation before attention runs. This lets a transformer understand word order — distinguishing "first token" from "second" and so on — while still getting the speed benefit of processing the whole sequence in parallel rather than one token at a time in order.