Core AI Technology

Temperature

Definition
Temperature is a core control parameter for text generation in LLMs. It describes how strongly the probability distribution over candidate tokens is reshaped before sampling. At temperature 1, the trained distribution stays unchanged. Lower values sharpen it: the model produces more consistent, predictable output because common tokens dominate. Higher values flatten it: less likely tokens gain real weight, making results more varied but also more error-prone. Two practical reference points have emerged: 0.2 for deterministic tasks such as classification, extraction, or JSON output, and 0.7 for more creative text. The limit value 0 corresponds to greedy decoding, where the most probable token is selected every time. Temperature differs from top-p (nucleus sampling): temperature reshapes the distribution, top-p restricts the candidate set; the two can be combined. Reproducibility matters in production: the value is fixed per task type and documented in the prompt or agent protocol, usually together with a seed. The name comes from thermodynamics, because the formula applies the Boltzmann distribution — the temperature indicates how strongly the system is excited. In practice the value stays small, since high temperatures measurably raise the error rate in multi-step pipelines. For companies, temperature is a simple but effective lever to consciously steer the trade-off between consistency and flexibility at every stage of an agentic workflow.</definition> <parameter name="relatedTerms">["llm", "top-p-sampling", "chain-of-thought", "structured-outputs", "prompt-engineering"]
Category
Core AI Technology

Deep Dive: Temperature

Temperature is a core control parameter for text generation in LLMs. It describes how strongly the probability distribution over candidate tokens is reshaped before sampling. At temperature 1, the trained distribution stays unchanged. Lower values sharpen it: the model produces more consistent, predictable output because common tokens dominate. Higher values flatten it: less likely tokens gain real weight, making results more varied but also more error-prone. Two practical reference points have emerged: 0.2 for deterministic tasks such as classification, extraction, or JSON output, and 0.7 for more creative text. The limit value 0 corresponds to greedy decoding, where the most probable token is selected every time. Temperature differs from top-p (nucleus sampling): temperature reshapes the distribution, top-p restricts the candidate set; the two can be combined. Reproducibility matters in production: the value is fixed per task type and documented in the prompt or agent protocol, usually together with a seed. The name comes from thermodynamics, because the formula applies the Boltzmann distribution — the temperature indicates how strongly the system is excited. In practice the value stays small, since high temperatures measurably raise the error rate in multi-step pipelines. For companies, temperature is a simple but effective lever to consciously steer the trade-off between consistency and flexibility at every stage of an agentic workflow.</definition> <parameter name="relatedTerms">["llm", "top-p-sampling", "chain-of-thought", "structured-outputs", "prompt-engineering"]

Tech Stack
OpenaiAnthropicGoogle

Production-Ready Guardrails