Tokens
A token is the small unit of text — a word piece, character, or symbol — that a language model actually reads and generates.
What Is a Token?
A language model doesn't read raw text the way you do. Before anything reaches the model, a tokenizer splits the text into tokens — small chunks that might be a whole word, part of a word, a punctuation mark, or a number.
Text
Raw InputNot what the model actually reads directly.
Tokenizer
Splits TextBreaks text into small, model-readable chunks.
Tokens
What the Model SeesWords, word pieces, punctuation, or numbers.
Model
Reads TokensNever sees the original raw text.
For example, a common word like "the" is usually its own token, but a less common or made-up word might get split into several pieces — "tokenization" could become "token" + "ization". Punctuation and whitespace are typically tokens too.
Warning
Tokenization is not universal. Different models use different tokenizers, so the same sentence can split into a different number of tokens depending on which model you ask.
Why Tokens Matter
- Context window — the maximum number of tokens a model can consider at once is measured in tokens, not words or characters.
- Latency — generating a response token by token means more tokens generally takes more time.
- Cost — many providers charge based on the number of input and output tokens processed.
- Input/output limits — both the prompt and the response count against the same token budget.
Common Mistakes
Assuming one word equals one token
Longer, rarer, or made-up words are frequently split into multiple tokens; short common words are often a single token.
Assuming token counts are the same across models
Each model family typically has its own tokenizer, so identical text can produce different token counts in different models.
Forgetting the response counts too
The context window and any per-request limits apply to the prompt and the generated output combined, not just what you send in.
Interview Question
What is a token, and why does tokenization matter?
A token is the small unit of text a model actually processes — a whole word, part of a word, or a punctuation mark, produced by a tokenizer before the text ever reaches the model. It matters because a model's context window, its response latency, and often its cost are all measured in tokens, not words or characters. Because tokenization differs by model, the same piece of text can use a different number of tokens depending on which model's tokenizer processes it.
What an interviewer may ask next
- Why might tokenizing the same sentence produce a different number of tokens on two different models?
- How does the number of tokens relate to the context window?
- Why would token count matter for the cost of running an application?
Explain It in 30 Seconds
A token is the small chunk of text — a word, part of a word, or a symbol — that a tokenizer produces before feeding it to a language model. Tokens matter because a model’s context window, response time, and often its cost are measured in tokens rather than words or characters, and tokenization differs from model to model, so you can’t assume a fixed word-to-token ratio.