AI Engineering Program — go from software engineer to production AI engineer · Live training with Kirill Eremenko · Watch the program breakdown→AI Engineering Program — go from software engineer to production AI engineer · Live training with Kirill Eremenko · Watch the program breakdown→AI Engineering Program — go from software engineer to production AI engineer · Live training with Kirill Eremenko · Watch the program breakdown→

Q: How do I monitor an LLM app and control API costs in production?

Monitor two layers: the classic one and the AI one.

The classic layer is what any web app watches: errors, latency, uptime. Your hosting platform and tools like CloudWatch cover it. The AI layer is what's new: token usage and cost per request, output quality over time, and traces of what the model actually did. Tools like Langfuse and LangSmith capture this: every call, its cost, its latency, and the full trace when something went wrong. Quality gets watched by running evals on samples of production traffic, because models change, user behavior shifts, and quality can degrade with no error ever thrown.

On costs, the levers in order of impact:

  1. Model right-sizing. Route routine requests to cheap models and reserve frontier models for the steps that need them. This is usually the biggest saving.

  2. Context discipline. Trim histories, cap retrieved chunks, summarize old conversation. You pay for every token in every call.

  3. Prompt caching. Structure prompts so the stable part is cached.

  4. Limits and alerts. Per-user rate limits, spending caps, and billing alerts, so a surprise looks like a notification instead of an invoice.

The mental model that keeps bills sane: in an LLM app, cost is not a finance topic, it's an architecture property. The decisions in your harness decide the invoice.

← Back to the full FAQ