Every LLM API exposes temperature and top-p, most people set both to "feel right," and half the advice online is wrong. Here's what each one actually changes about the model's next-token choice — and why tuning both at once fights itself.

Open any LLM playground and you're met with two sliders — temperature and top-p — usually with a vague tooltip about "creativity." So people nudge both, get output that feels a bit different, and move on without ever knowing what they changed. That's a shame, because these two parameters are precise, they do different jobs, and setting both aggressively at the same time produces behavior that's hard to reason about. Understanding them takes about five minutes and pays off every time you tune a prompt.

What the model is doing at each step

A language model doesn't decide on a whole sentence. At every step it produces a probability for every possible next token — thousands of numbers saying "given everything so far, here's how likely each token is to come next." The word "the" might get 18%, "a" 9%, some rare token 0.001%, and so on across the whole vocabulary. Sampling is the act of picking one token from that distribution. Temperature and top-p are two different ways of shaping how that pick happens — and neither one changes the model's underlying knowledge, only how boldly it commits to its own top guesses.

Temperature: reshaping the whole distribution

Temperature is a single number applied to the model's raw scores before they become probabilities, and it stretches or compresses the entire distribution. The parameter comes straight from the softmax function that turns scores into probabilities:

"The softmax function [...] converts a vector of K real numbers into a probability distribution of K possible outcomes. [...] A higher temperature results in a more uniform output distribution, while a lower temperature results in a sharper distribution concentrated on the most likely outcomes."

— Wikipedia, "Softmax function" (CC BY-SA 4.0)

In plain terms: low temperature (near 0) makes the model greedy — it piles probability onto the single most likely token and picks it almost every time, giving focused, repeatable, sometimes robotic output. High temperature (above 1) flattens the distribution so unlikely tokens get a real chance, giving varied, surprising, sometimes incoherent output. Temperature 0 is effectively deterministic: always take the top token. It touches every token in the vocabulary, including the long tail of nonsense ones.

Top-p: trimming the tail before you sample

Top-p — also called nucleus sampling — takes a different approach. Instead of reshaping probabilities, it throws away the unlikely options entirely, then samples from what's left. You set a cutoff like 0.9, and the model keeps adding tokens to the candidate pool, most-likely first, until their probabilities sum to 90%. Everything below that line is discarded and can't be chosen at all. The clever part is that the pool size adapts to context: when the model is confident (one token holds most of the probability) the pool is tiny; when it's genuinely uncertain across many plausible words, the pool is large. Top-p caps the weirdness — it can't pick from the tail because the tail was cut — while still allowing variety when variety is warranted.

Why tuning both at once fights itself

Here's the practical rule most tooltips omit: temperature and top-p are two levers on the same quantity — output randomness — and cranking both makes the result hard to predict. High temperature inflates the tail's probabilities; top-p then tries to cut that inflated tail; the interaction is genuinely unintuitive. The common guidance is to pick one to tune and leave the other at its neutral default. Adjust temperature or adjust top-p, not both. For most work, holding top-p at its default and moving temperature alone gives you a single, interpretable dial: lower for factual, structured, repeatable tasks; higher for brainstorming and creative drafting.

Neither one fixes a bad prompt

It's tempting to reach for the sliders when output is wrong, but sampling parameters only control randomness — they can't add information the prompt didn't supply or remove a contradiction the prompt contains. If the model is hallucinating a fact, lowering temperature makes it hallucinate the same fact more consistently; it doesn't make the fact true. Randomness tuning is the last mile after the prompt is right, not a substitute for a clear instruction. Fix the prompt first; reach for temperature to control style and repeatability second.

Test the effect instead of guessing

The honest way to tune these is empirically — change one knob, hold everything else, and compare outputs. Run the same prompt at two temperatures and put the results side by side with a prompt diff tool to see exactly how much the wording actually shifted, rather than trusting a vibe. Keep an eye on length and cost while you experiment: a token counter shows how many tokens each variant spends, since higher temperature often rambles. And before you blame the sampler for inconsistent behavior, run your instructions through a system prompt linter — a surprising amount of "randomness" is really an ambiguous prompt leaving the model free to wander. Set one knob deliberately, measure the difference, and the two mysterious sliders become two precise tools.

← All articles