AI Engineering Program — go from software engineer to production AI engineer · Live training with Kirill Eremenko · Watch the program breakdown→AI Engineering Program — go from software engineer to production AI engineer · Live training with Kirill Eremenko · Watch the program breakdown→AI Engineering Program — go from software engineer to production AI engineer · Live training with Kirill Eremenko · Watch the program breakdown→

Q: Why do LLM outputs differ every run — and how do I make an LLM follow instructions reliably?

Because generation is probabilistic by design. At every step, the model chooses the next token from a probability distribution, and that choice involves randomness. Same prompt, different run, different path. This is normal, and your engineering has to assume it.

That's half the question. The other half is the one that actually frustrates people: the model ignoring your instructions. You write "only answer from the provided documents," and it cheerfully answers something else anyway. Every engineer building their first real app hits this.

What reliably helps, in order of impact. Make instructions explicit and specific: "If the answer is not in the context below, reply exactly: I don't know" beats "try to stick to the context." Put critical rules in the system prompt, not buried mid-conversation. Emphasis works: models really do weight IMPORTANT, capitalization, and repetition of the critical rule at the end of the prompt. Give an example of the behavior you want (one good example beats three paragraphs of description). And restructure your content so the model can follow the rule: well-organized context makes obedience easy, a wall of messy text makes it hard.

And know where prompting stops: a well-written prompt gets the model to follow instructions most of the time, but "most of the time" is not a guarantee. For the last mile, you verify outputs in code (checks, retries, guardrails, evals) instead of trusting words to do a program's job. Reliability is an engineering property, not a prompting trick.

← Back to the full FAQ