AI Workspace Loading

We’re preparing your intelligent learning experience. Our AI systems are processing content, optimizing resources, and setting everything up for you.

Preparing Learning Paths...
AI Processing
Smart Automation
Learning Engine
Good things take a moment.

LearnLess.ai

LEARN LESS. UNDERSTAND MORE.
Beginner4 min read

Temperature

Temperature controls how random or deterministic a model's output is during generation.

What Is Temperature?

At each step of generation, a model doesn't just pick one next token — it computes a probability for every token in its vocabulary. Temperature is a setting that changes how those probabilities get turned into an actual choice.

Token Probabilities

Raw Scores

A probability for every token in the vocabulary.

adjusted by

Temperature

Reshapes Probabilities

Sharpens or flattens the distribution.

feeds

Sampling

Makes the Pick

Turns probabilities into one actual choice.

selects

Chosen Token

One Result

The single token that gets generated.

A low temperature sharpens the differences between probabilities, so the model almost always picks the highest-probability token — output looks focused and repeatable. A high temperature flattens the differences, giving lower-probability tokens a real chance of being picked — output looks more varied and creative, but also less predictable.

Key Idea

Temperature doesn't change what the model knows — it changes how confidently it commits to its top guess versus exploring less likely ones.

How It Works in Practice

  • Temperature near 0 — the model almost always picks the single most likely next token. Good for factual Q&A, extraction, classification, and code generation, where you want the same input to produce a consistent output.
  • Temperature around 0.7–1.0 — a common default range that balances coherence with variety. Good for general conversation and everyday assistant tasks.
  • Temperature above 1.0 — output becomes noticeably more varied and unpredictable, and at high enough values can turn incoherent. Useful for brainstorming or creative writing where variety matters more than reliability.

The exact numeric range and default differ by provider and model, but the underlying idea — a knob between "always pick the top choice" and "sample more broadly" — is the same everywhere.

A Real-World Example

Ask a model the same factual question twice at temperature 0, and you should get the same answer both times, because the model keeps choosing its highest-probability token at every step. Ask the same question twice at temperature 1.2, and you may get two differently worded (and occasionally differently correct) answers, because the model is more willing to pick a lower-probability token somewhere along the way — and small differences early in generation compound as later tokens are generated based on what came before.

Code Example

Illustrative — the exact parameter name and default vary by SDK and provider:

temperature_example.py (illustrative)
# Low temperature: deterministic, repeatable — good for extraction/classification
response = client.generate(
    prompt="Extract the invoice total from this text: ...",
    temperature=0,
)

# Higher temperature: more varied — good for brainstorming
response = client.generate(
    prompt="Suggest five taglines for a coffee shop.",
    temperature=1.0,
)

Common Mistakes

  • Assuming temperature makes the model "smarter"

    Temperature changes how the model samples among its existing predictions — it does not add knowledge or reasoning ability.

  • Using a high temperature for factual or structured tasks

    Extraction, classification, and structured-output tasks usually want low or zero temperature for consistent, reliable results.

  • Assuming temperature 0 guarantees identical output every time

    Depending on the provider and underlying hardware, extremely low temperature is usually very consistent but is not always perfectly deterministic.

  • Assuming the same temperature value behaves identically across models

    The exact effect of a given temperature value depends on the model and provider — treat it as a relative knob, not an absolute setting.

Interview Question

What does temperature control in an LLM, and when would you change it?

Temperature controls how the model samples from its predicted probabilities for the next token. A low temperature makes it almost always pick the highest-probability token, producing focused, repeatable output — good for factual answers, extraction, or code. A higher temperature flattens those probabilities so less likely tokens get picked more often, producing more varied, creative output at the cost of consistency. It's a sampling setting, not a measure of intelligence — it doesn't change what the model knows, just how confidently it commits to its top guess.

What an interviewer may ask next

  • Why might you want temperature 0 for a data-extraction task?
  • Why can two requests at high temperature produce very different answers to the same question?
  • Is temperature 0 guaranteed to produce identical output every time? Why or why not?

Explain It in 30 Seconds

Temperature controls how randomly a model samples its next token from the probabilities it computes. Low temperature sticks to the highest-probability choice, giving focused, repeatable output. High temperature gives lower-probability tokens more of a chance, giving more varied but less predictable output. It doesn't make the model smarter — it just changes how much it explores beyond its top guess.

On this page