AI Engineering Program — go from software engineer to production AI engineer · Live training with Kirill Eremenko · Watch the program breakdown→AI Engineering Program — go from software engineer to production AI engineer · Live training with Kirill Eremenko · Watch the program breakdown→AI Engineering Program — go from software engineer to production AI engineer · Live training with Kirill Eremenko · Watch the program breakdown→

Q: What is a context window, and what happens when you exceed it?

The context window is the maximum amount of text a model can hold in one call, measured in tokens. Everything counts against it: your system prompt, the conversation history, any documents you've included, and the model's own answer.

Current frontier models have windows from around 200,000 tokens up to a million or more. That sounds infinite. It isn't, for two reasons.

First, the hard limit: exceed the window and the API rejects your call, or your app has to cut something to make the request fit.

Second, the window fills much faster than people expect, because of what has to fit in it. Remember that the model has no memory: your app resends the entire conversation history with every message, and all of it counts against the window, every turn. And if you're using a reasoning model (or thinking mode), the model's internal reasoning consumes tokens from the budget too, even though nobody ever sees them. A long conversation about a few large documents, with reasoning turned on, can eat hundreds of thousands of tokens before you've done anything unusual.

Now the takeaway that ties it together: don't treat the limit as the target, because quality degrades well before it. Models pay the most attention to the beginning and end of the context and can lose track of information buried in the middle (researchers call this "lost in the middle"). A model with a million-token window can still miss a fact on page 200. So the engineering discipline is to curate the context, not fill it: the relevant chunks, the trimmed history, the instructions that matter. More signal, less filler. That's half the reason RAG exists, and since you pay for every token in the window on every call, curation is also where your API bill gets decided.

← Back to the full FAQ