Encoder
An encoder reads an input sequence and produces a contextual representation of it.
Prerequisites
What an Encoder Does
An encoder reads an entire input sequence — all at once, not one token at a time in order — and produces a representation of it that captures the meaning and relationships of every token in context, using self-attention. That representation, not the raw input text, is what a downstream component actually works with.
Input Sequence
Read All at OnceNot processed one token at a time in order.
Self-Attention Layers
Relates Every TokenCaptures meaning and relationships in context.
Contextual Representation
What Gets UsedWhat a downstream component actually works with.
Key Idea
An encoder's job is understanding, not generating — it produces a rich representation of the input, but doesn't itself produce new output text.
Where Encoders Are Used
- Embedding models — many embedding models are essentially encoders, turning text into a vector representation used for similarity comparison and search.
- Encoder-decoder architectures — some models pair an encoder (to understand the input) with a decoder (to generate output), useful for tasks like translation.
- Encoder-only models — used for tasks like classification, where you need to understand input deeply but don't need to generate new text.
- Decoder-only LLMs — most modern chat-oriented LLMs are decoder-only, meaning they don't use a separate encoder stage at all; the decoder handles both understanding and generating from the same sequence.
Common Mistakes
Assuming every LLM has a separate encoder
Many popular chat-oriented LLMs are decoder-only — there's no separate encoder stage; understanding and generation happen in the same component.
Assuming an encoder generates text
An encoder produces a representation of the input for something else to use — generating new output text is the decoder's job.
Assuming encoders and embedding models are unrelated
Many embedding models are structurally encoders — the same encoding idea, applied to produce a vector representation used for search rather than for a decoder to continue from.
Interview Question
What does an encoder do in a transformer, and how does it differ from a decoder?
An encoder reads an entire input sequence at once and produces a contextual representation of it, using self-attention to capture how every token relates to every other token — its job is understanding the input, not generating new text. A decoder, by contrast, generates output — typically one token at a time, using previously generated tokens and, in an encoder-decoder architecture, the encoder's representation as context. Many modern chat-oriented LLMs are actually decoder-only, meaning there's no separate encoder stage at all; the decoder handles both understanding and generating within the same component. Encoders show up on their own too — many embedding models are structurally encoders, turning text into a vector representation for search rather than passing it to a decoder.
What an interviewer may ask next
- Why don't most modern chat-oriented LLMs use a separate encoder?
- How does an encoder relate to embedding models?
- What kind of task would benefit from an encoder-decoder architecture specifically?
Explain It in 30 Seconds
An encoder reads an entire input sequence at once and produces a context-aware representation of it using self-attention — its job is understanding, not generating text. It's used in encoder-decoder architectures for tasks like translation, in encoder-only models for tasks like classification, and structurally underlies many embedding models. Most modern chat-oriented LLMs are actually decoder-only, with no separate encoder stage at all.