Transformers
The transformer is the neural network architecture — built around attention — behind nearly every modern language model.
Prerequisites
Why Transformers Matter
Before transformers, the dominant approach for processing sequences (like sentences) processed one token at a time, in order, carrying forward a summary of everything seen so far. That worked, but it was slow to train and struggled to keep track of relationships between tokens that were far apart in a sequence.
The 2017 transformer architecture replaced that step-by-step processing with attention: every token can directly look at every other token in the sequence at once, regardless of distance. That has two big consequences: the model can capture long-range relationships far more effectively, and — because tokens are processed together rather than one after another — training can be heavily parallelized on GPUs.
How It Works
At a high level, a transformer takes a sequence of tokens through a repeating stack of layers:
Input Tokens
Sequence StartThe tokens the model will process together.
Embeddings + Position
Meaning + OrderToken meaning combined with position information.
Attention
All at OnceEvery token looks at every other, regardless of distance.
Feed-Forward Network
Per-Token ProcessingFurther transforms each token after attention.
Repeated Layers
Stacked DepthThe same block repeated many times.
Output
Final RepresentationReady for whatever task comes next.
- Embeddings
- Each token is converted into a vector that represents it numerically before any processing happens.
- Positional Encoding
- Information about each token's position in the sequence, added to its embedding — attention alone has no built-in sense of order.
- Attention
- The mechanism that lets each token weigh how relevant every other token is when building its representation.
- Feed-Forward Network
- A smaller network applied to each position after attention, adding further capacity to transform the representation.
Encoder, Decoder, or Both
Transformers are used in three general shapes, depending on the task:
- Encoder-only
- Reads the entire input at once to build a rich representation — well suited to understanding tasks like classification or search.
- Decoder-only
- Generates a sequence one token at a time, only attending to tokens already produced — the shape most modern LLMs use.
- Encoder-decoder
- An encoder processes the input and a decoder generates the output from it — a natural fit for tasks like translation or summarization.
Common Mistakes
Assuming "transformer" means "LLM"
Transformers are used well beyond language — in vision, audio, and multimodal models. LLMs are one major application of the architecture.
Thinking attention processes tokens in order
Attention itself has no inherent sense of sequence — that is exactly why positional encoding has to be added separately.
Overstating parallelization at inference time
Training benefits heavily from parallelization; generating a response token by token at inference is still sequential.
Interview Question
Why are transformers important for modern AI?
Transformers replaced step-by-step sequence processing with attention, letting every token relate directly to every other token in a sequence regardless of distance. That improved the model's ability to capture long-range relationships and, because tokens can be processed together rather than strictly in order, made training far more parallelizable on GPUs. That combination of quality and trainability at scale is a big part of why transformers became the default architecture behind modern language models.
What an interviewer may ask next
- What problem did attention solve compared to earlier sequence models?
- Why does a transformer need positional encoding?
- What is the difference between an encoder-only and a decoder-only transformer?
Explain It in 30 Seconds
A transformer is a neural network architecture built around attention, which lets every token in a sequence directly relate to every other token instead of processing them strictly in order. That made it much better at capturing long-range relationships in text, and because tokens can be processed together, training scales well on parallel hardware. It's the architecture behind nearly every modern language model.