Inference
Inference is using a trained model to produce predictions or outputs on new data.
What Is Inference?
Once a model is trained, using it is called inference: feeding it new input it hasn't seen before and getting back a prediction or generated output. Every time a chat application calls an LLM to answer a question, that single call is one inference request.
Key Idea
Training happens once, offline, over a large dataset. Inference happens every time the model is actually used — potentially thousands of times a second in production.
Why Inference Is an Engineering Concern
- Latency — every inference call takes real time, and that time is on the critical path of whatever product feature depends on it.
- Cost — most providers charge per inference call, typically based on the number of input and output tokens processed.
- Throughput — a production system needs to handle many concurrent inference requests, which is a different problem from training a single model once.
- Consistency — the same input can produce a different output on different calls, depending on settings like temperature, which is a property specific to inference, not training.
This is exactly why concepts like caching, streaming, rate limiting, and model routing exist — they're all ways of managing the cost, latency, and scale of inference in a production system.
Common Mistakes
Assuming inference is free or instantaneous
Every inference call has real latency and, with most providers, real cost — this adds up quickly at production scale.
Confusing inference with training
Inference doesn't change the model or "teach" it anything — it just produces an output from the model as it already is.
Ignoring inference cost when designing a feature
A feature that makes several LLM calls per request multiplies both latency and cost — this needs to be part of the design, not an afterthought.
Interview Question
What is inference, and why does it matter as an engineering concern rather than just a machine learning concept?
Inference is using an already-trained model to produce a prediction or generated output on new input — as opposed to training, which is the one-time process that produced the model in the first place. It matters as an engineering concern because every inference call has real latency and, with most providers, real cost, and a production system needs to handle many of these calls concurrently and reliably. That's exactly why patterns like caching, streaming, rate limiting, and model routing exist — they're all ways of managing the cost, latency, and scale of inference in production, not concerns that show up during training.
What an interviewer may ask next
- Why can the same input produce a different output on two different inference calls?
- How does inference cost scale differently from training cost?
- What engineering patterns exist specifically to manage inference at scale?
Explain It in 30 Seconds
Inference is using an already-trained model to produce a prediction or output on new input, as opposed to training, which is the one-time process that created the model. It matters in production because every inference call has real latency and cost, and a system needs to handle many of them concurrently — which is why patterns like caching, streaming, and rate limiting exist specifically to manage inference at scale.