AI Engineering Program — go from software engineer to production AI engineer · Live training with Kirill Eremenko · Watch the program breakdown→AI Engineering Program — go from software engineer to production AI engineer · Live training with Kirill Eremenko · Watch the program breakdown→AI Engineering Program — go from software engineer to production AI engineer · Live training with Kirill Eremenko · Watch the program breakdown→

Q: What is prompt caching, and when does it save money?

Prompt caching means the provider remembers the beginning of a prompt it has recently seen. When you send it again, the repeated part is processed at a steep discount (typically 50-90% off) and faster.

It matters because of how conversations work: your app resends the system prompt and full history with every message. Without caching you pay full price for text the model has already read a hundred times. With caching, the repeated prefix is nearly free. For chat apps, agents, and RAG systems, this is often the difference between a reasonable API bill and a shocking one.

One rule follows: put stable content first (system prompt, reference documents) and changing content last (the newest user message). Caches match from the start of the prompt, so an early change breaks the cache for everything after it.

Each provider documents the details: OpenAI, Anthropic, Gemini.

← Back to the full FAQ