Sort the levers by the evidence each one needs

Most guides list the same levers: caching, batching, prompt trimming, model routing, budgets. The list is correct. What it leaves out is sequence, and sequence decides whether a savings project ends with a smaller invoice or a quality incident.

A useful ordering asks one question of each lever: can it change what the user receives? Provider discounts cannot. The same model answers the same prompt at a lower price. Trimming a prompt changes the input, so it needs a regression check. Swapping the model or serving a stored answer changes the output itself, so it needs an evaluation on your own traffic before it carries production load.

Before any layer, build the ledger. Record owner, route, model, input and output tokens, cache state, retries, and outcome for every request. Without it you cannot find the largest cost pools, and you cannot prove afterward that a change saved money.

  • Layer 1, provider discounts: output unchanged. Evidence needed is billing data.
  • Layer 2, request shape: input changes. Evidence needed is a regression check.
  • Layer 3, model routing and response caching: output changes. Evidence needed is an evaluation on your traffic.
  • Layer 4, governance: keeps the first three from eroding.

Layer 1: discounts that leave the output unchanged

Batch processing is the plainest discount available. Anthropic's Message Batches API offers a 50% discount on all usage compared to standard prices, with most batches finishing in less than an hour and expiring if they do not complete within 24 hours. OpenAI's Batch API documents a 50% cost discount compared to synchronous APIs, with each batch completing within 24 hours, and its batch rate limits sit in a separate pool from standard limits. Any workload that can wait qualifies: evaluations, backfills, nightly summaries, classification jobs.

Prompt caching discounts repeated prefixes. On Anthropic, a 5-minute cache write costs 25% more than base input tokens, a 1-hour write costs 2x, and a cache read costs 10% of the base input price on most models, less on some newer ones. On OpenAI, caching is enabled by default for supported models, and reused tokens bill at a cached-input rate discounted up to 95%.

The write premium sets a break-even. With Anthropic's 5-minute cache at the standard read rate, two requests sharing a prefix cost 1.35x the base price for that prefix instead of 2x, so the cache pays off on the first reuse. The 1-hour cache is different. One write plus one read costs 2.1x against 2x uncached, so it needs at least two reads to come out ahead. A prefix that is written and never read costs more than no caching at all.

Two documented details catch teams out. Prompts under the minimum length are processed without caching, and no error is returned. Anthropic's minimum ranges from 512 to 4,096 tokens depending on the model, and OpenAI's is 1,024 tokens for GPT-5.6 and later. The discounts also combine: Anthropic states that prompt caching and batch discounts can stack, though cache hits inside a batch are best-effort because requests run concurrently.

  • Batch: 50% off at both Anthropic and OpenAI, in exchange for asynchronous delivery.
  • Anthropic cache: 1.25x to write for 5 minutes, 2x to write for 1 hour, 0.1x to read on most models.
  • OpenAI cache: automatic on supported models, cached input discounted up to 95%.
  • Check the usage fields for cache reads. A silent miss looks identical to a working request.

Layers 2 and 3: change the request, then change the model

Layer 2 reduces tokens per request. Remove duplicated system context, retrieve fewer and better chunks, cap output length at what the product displays, and stop retries from multiplying a failed call. Agent loops deserve their own audit, because each step resends the growing conversation and the tool definitions. These changes alter the input, so run your regression set before and after. They also interact with layer 1: a stable prefix caches, and a prompt rewritten on every call does not.

Layer 3 changes the answer. Routing sends each request to the cheapest model that passes your quality bar. Exact and semantic response caching return a stored answer instead of generating a new one. The potential savings are large, and so is the risk of degrading the product quietly, which is why this layer comes last.

The safe procedure is the same for both. Group traffic into workload families, define a success check per family, run the cheaper path in shadow mode, and compare quality and effective cost before enforcing one cohort at a time. Count escalations, retries, and evaluator calls in the cost of the cheaper path. Headline percentages in vendor guides were measured on someone else's traffic. Your operating point has to come from yours.

Layer 4: governance keeps the savings

Optimization decays. Prompts grow, new features ship on the most expensive model by default, providers reprice, and a cache that hit reliably last quarter stops hitting after a template change. Nobody notices until the invoice arrives.

Governance is the layer that catches this. Give each team or product a budget with alerts well below the limit. Forecast month-end spend from the current run rate. Track cost per successful outcome, so growth in valuable usage reads differently from waste. Re-run the layer 1 checks whenever a prompt template or a provider price changes.

The layers also tell you where the controls belong. Discounts and token reductions live in application code and provider settings. Routing, response caching, budgets, and attribution work best at a shared layer every request passes through: an LLM gateway such as LiteLLM, a FinOps-focused gateway such as FrugalAI, or an internal build. Know which layer you are working on before you pick the tool.

Frequently asked questions

What is LLM cost optimization?

It is the practice of reducing what you spend on model inference without lowering the quality of the output. In practice that means four kinds of work: claiming provider discounts, sending fewer tokens, routing evaluated workloads to cheaper models, and governing spend so the savings last.

What is the first step in LLM cost optimization?

Measure, then claim the discounts that leave output unchanged. Build a request ledger, find the largest cost pools, move work that can wait to a batch API at 50% off, and structure prompts so stable prefixes hit the provider's prompt cache.

Does prompt caching always save money?

No. On Anthropic, a 5-minute cache write costs 25% more than normal input and a 1-hour write costs 2x, so a prefix that is written but rarely read raises cost. Prompts below the model's minimum cacheable length are also processed uncached without an error. Check the cache read fields in the usage data.

Sources and further reading

  1. Anthropic prompt caching (verified 2026-10-03)
  2. Anthropic batch processing (verified 2026-10-03)
  3. OpenAI Batch API (verified 2026-10-03)
  4. OpenAI prompt caching (verified 2026-10-03)

FrugalAI uses primary documentation and published research where possible. Product capabilities and prices can change; verify vendor details before procurement or production changes.