Q: What is temperature, and when should I change it?
Temperature controls how random the model's word choices are. Low temperature: the model almost always picks the most probable next token, so outputs are consistent and focused. High temperature: less probable tokens get a chance, so outputs are more varied and creative, and eventually incoherent.
The scale typically runs 0 to 2, default around 1. In practice: use 0 to 0.3 for code, data extraction, and anything where correctness matters; the default for general work; 0.7 to 1.2 when you want variety, like brainstorming.
One important 2026 update: OpenAI's newest models (the GPT-5 family and o-series) no longer accept temperature at all. You control them through settings like reasoning effort instead. Temperature still works as described on Claude, Gemini, open source models, and older OpenAI chat models.
Two things engineers get wrong. Temperature 0 doesn't make outputs fully deterministic: small variations remain. And temperature is not a quality dial: turning it down makes the model more repeatable, not more accurate. A wrong answer at temperature 0 is wrong every time.