Learn AI
Build your understanding from AI fundamentals to production AI systems.
Prefer a guided sequence? Explore Learning PathsTransformers
The attention-based architecture behind nearly every modern language model.
6 conceptsTransformers
The attention-based architecture behind nearly every modern LLM.
Intermediate · 5 minAttention
Attention lets a model weigh how relevant each part of the input is when producing each part of the output.
Intermediate · 5 minSelf-Attention
Self-attention lets each position in a sequence weigh every other position in the same sequence to build context-aware representations.
Intermediate · 5 minEncoder
An encoder reads an input sequence and produces a contextual representation of it.
Intermediate · 4 minDecoder
A decoder generates an output sequence, typically one token at a time, using the encoded context and what it has generated so far.
Intermediate · 4 minPositional Encoding
Positional encoding injects information about token order into a transformer, which otherwise has no built-in sense of sequence.
Intermediate · 4 min