The New Default. Your hub for building smart, fast, and sustainable AI software

See now
Abstract, minimalist illustration of LLM model routing: one path splitting into two, leading to small and large blocks that represent model tiers.

LLM Cost Optimization in 2026: Model Routing, Context Discipline, and Prompt Caching

Maciej Korolik
|   Oct 10, 2026

LLM cost optimization means getting the output quality you need from the fewest, cheapest tokens.

The biggest savings come from architecture: which model answers which request, and how much of each prompt the system reuses instead of paying for again. Rewording prompts barely registers by comparison.

The bill is now a budget-owner problem too. GitHub Copilot switched every plan to usage-based billing on June 1, 2026, and Claude Fable and GPT-5.5 launched at roughly twice the per-token price of the models they replaced.

Below, we trace where tokens leak and which fixes hold up in production.

Executive Summary

The most effective way to cut LLM costs is to treat token efficiency as a system design problem: route each request to the cheapest model that can handle it, and send it only the context it needs. Input prices between model tiers now differ by up to 100×, and in the multi-agent coding system one 2026 study measured, most tokens went to re-reading context the agents passed back and forth.

That makes untuned architectures overspend in predictable places, which is good news, because predictable waste can be measured and removed. Quick fixes like capping output length or moving every call to a cheaper model save money on paper and then come back as quality regressions and retries.

Architectural changes such as routing, caching, and selective retrieval take weeks, and when a baseline and evaluations back them, the savings hold as volume grows.

Why Are LLM Token Costs Rising in 2026?

AI token costs are rising because reasoning models and agents consume far more tokens per task, and vendors have started passing that cost on to customers.

Reasoning models generate internal "thinking" tokens before answering, and those are usually billed as output tokens, the most expensive kind. Agents compound this.

In Mike Loukides' description for O'Reilly Radar, a single agent request can become many model calls, each carrying the task's accumulated context. He estimates the combined effect of reasoning and agents at "a factor of hundreds."

Operationally, that makes per-task cost hard to predict, even within one system. A January 2026 Concordia University study, "Tokenomics: Quantifying Where Tokens Are Used in Agentic Software Engineering", ran 30 software tasks through the ChatDev multi-agent framework on GPT-5.

Reasoning tokens alone ranged from 17,280 to 40,000 per task, and code review consumed an average of 59.4% of all tokens, while writing the initial code took 8.6%.

For the business, AI spend now shows up as a line item someone has to defend. Copilot's monthly allotments of AI Credits are now metered against input, output, and cached tokens. At the API level, Anthropic's current pricing puts Claude Fable 5.1 at $10 per million input tokens and $50 per million output tokens, twice Opus 4.8's $5/$25. Teams that built features assuming prices would keep falling now face margins they can't forecast.

As Drew Breunig put it in "Fable & The End of the Free Lunch", there was a time when it barely paid to improve your coding harness or context strategy, because the next model would arrive at the same price and paper over the problems. With top-tier prices going up, teams now have to decide which work goes to which model.

Where Do Tokens Leak in LLM Applications and AI Agents?

Tokens leak wherever a system sends the model more than the current step needs, or asks it to recompute something it has already worked out.

These are the patterns we see most often:

  • Bloated context windows: full documents, entire chat histories, and raw tool outputs loaded into every call "just in case."

  • System prompts repeated in conversation history, re-sent with every turn and sometimes duplicated when the conversation is rebuilt.

  • Naive RAG pipelines that return whole documents, or dozens of chunks, when three to five precise passages would answer the question. PostHog's LLM cost guide cites "3-5 RAG chunks instead of 50" as a typical correction.

  • Uncached repeated inputs: the same system prompt, tool definitions, and reference material processed at full price on every request.

  • Tool bloat. An agent that sees dozens of tool definitions pays to read all of them on each call and is more likely to pick the wrong one.

  • Unbounded review loops, where agents pass full artifacts back and forth for review and revision. In the Tokenomics study, input tokens made up 53.9% of all usage, which the authors attribute to a "communication tax" between agents.

  • Verbose output. Output tokens cost five times as much as input tokens on every current Claude model. A classifier that writes a paragraph to explain a one-word label is expensive.

Almost every item on this list is a design decision, which is why trimming adjectives from a system prompt rarely moves the needle.

Quick Fixes or Architectural LLM Cost Optimization?

Quick fixes deliver fast savings but often trade away quality. Architectural optimization takes longer, protects quality, and keeps paying off as volume grows.

Approach

Type

Typical effect on cost

Risk to quality

Effort

Use it when

Lower max_tokens across the board

Quick fix

Caps output spend

High: truncates answers that need to be long

Minutes

Only for tasks with short, bounded outputs (classification, extraction)

Switch every call to a cheaper model

Quick fix

Large per-token drop

High: hard tasks fail, and retries erode the savings

Hours

Never as a blanket policy; use routing instead

Structured outputs and per-task output limits

Quick fix with design

PostHog reports 30–60% of output cost in some cases

Low when scoped per task

Days

Extraction, classification, form filling

Prompt caching

Architectural

Cache reads billed at 2.5–10% of base input price on Claude

Very low

Days

Any workload with a stable prompt prefix

Batch processing

Architectural

50% off input and output on Anthropic's Batch API

None; adds latency

Days

Offline jobs nobody is waiting on

Model routing across tiers

Architectural

Depends on traffic mix; tier prices differ up to 100×

Medium: misroutes need detection and fallback

Weeks

Mixed workloads with many simple requests

Selective, authorized retrieval

Architectural

Smaller prompts on every RAG call

Low, and often improves accuracy

Weeks

Knowledge assistants and enterprise search

Structured state and bounded loops

Architectural

Stops input from growing turn after turn

Medium: over-trimming drops needed context

Weeks

Long conversations and multi-agent workflows

Semantic caching

Architectural

Skips the model entirely on cache hits

High: wrong or stale answers

Weeks

Read-only, stable, FAQ-style answers

Before you build a multi-model cascade, test your current model at a lower reasoning-effort setting, which most frontier APIs now expose. Staying on one model also keeps a single prompt cache, since caches don't carry over between models.

How Do You Measure Where LLM Tokens Go Before Optimizing?

Start by building a cost map per workflow, broken down by stage and token type, before you change anything. A single "total tokens per day" number hides everything you need to know.

Instrument every model call with metadata: the feature, the workflow stage (retrieval, planning, execution, review, retry), the model, and the input, output, cached, and reasoning token counts.

Tools like PostHog's AI observability estimate cost from token counts and matched pricing. PostHog itself recommends reconciling the totals against your provider's console, because negotiated rates and new models can fall outside its price data.

Then add quality signals next to cost, such as task success rate and how often a human has to correct the output. The number to manage is cost per successfully completed task.

A cheap call that needs three retries and a human fix costs more than a call that gets it right the first time. The Tokenomics paper makes the point that the distinct cost profiles of design, coding, and review let teams predict spend from the type of work involved.

Watch input tokens across a single trace or session. If they climb with every turn, you've found a context leak.

How Does Model Routing Across Small, Medium, and Large LLMs Cut Costs?

Model routing sends each request to the cheapest model tier that can reliably handle it, so the most expensive model works only on the tasks that need it. In practice, three functional tiers cover most systems.

Tier

Example models (Anthropic list prices, per million input/output tokens)

Typical work

Small

Claude Haiku 5.5 ($0.10 / $0.50 for prompts up to 100K tokens); open-weight models

Classification, intent detection, extraction, routing decisions, short summaries

Medium

Claude Sonnet 5.5 ($2 / $10), Claude Opus 5.5 ($4 / $20)

Most RAG answers, routine coding, standard agent steps

Large

Claude Fable 5.1 ($10 / $50)

System design, ambiguous high-stakes reasoning, long-horizon planning

Prices from Anthropic's pricing page, checked October 10, 2026.

The gap is easy to underestimate. Take one million classification requests a month, each with 1,000 input tokens and a 20-token answer. On Fable 5.1, that's about $10,000 in input and $1,000 in output, before any reasoning tokens (thinking is always on for Fable). On Haiku 5.5, the same traffic costs about $110. That 100x gap is meant to be closes with routing.

Routing can be rule-based (by feature, endpoint, or input length) or model-based (a small classifier that scores complexity and risk). Glean's guide on token efficiency in agentic systems suggests routing on three factors: complexity, the cost of a wrong answer, and how much context the task needs. Each route should also see only the tools it needs.

The same pattern works at the human-in-the-loop level. Breunig describes using the frontier model to shape a design, then handing a written brief to a model he estimates at roughly one-ninth the cost for the execution. Use the expensive model to decide what to do, and cheaper models to do it.

Remember that routing pays off only if the savings exceed the cost of the classifier, the retries from misroutes, and the engineering time to maintain it.

How Does Authorized Selective Retrieval Reduce RAG Token Costs?

Authorized selective retrieval sends the model only the specific passages that are both relevant to the question and permitted for the user asking it. That reduces tokens per call and compliance exposure at the same time.

A precision retrieval pipeline combines four steps:

  1. Hybrid search. Keyword search (such as BM25) catches exact terms, product codes, and names. Vector search catches paraphrases. Together they surface better candidates than either does alone.

  2. Reranking. A reranker scores candidates by relevance, and only the top few go into the prompt.

  3. Permission filtering at retrieval time. Documents the user isn't authorized to see never enter the context. That keeps tokens off the bill and protects against data leaking through the model's answer.

  4. Just-in-time loading. Agents keep references (file paths, IDs, queries) and fetch content only when a step needs it. Anthropic's context engineering guidance describes Claude Code doing exactly this with tools like grep, head, and tail instead of loading whole files.

Smaller context also tends to mean better answers. Anthropic describes "context rot" as the number of tokens in the context window grows, the model's ability to recall information from it declines.

It treats context as a finite resource with diminishing marginal returns. For teams in regulated industries, permission-aware retrieval also supports the GDPR principle of data minimization: you process only the personal data the task needs.

How Does Prompt Caching Work, and What Breaks It?

Prompt caching stores the processed form of a stable prompt prefix, so repeat requests pay a fraction of the normal input price for that prefix. It's the lowest-risk option in this guide, because the model sees identical content and only the price changes.

On Anthropic's API, writing to the cache costs 1.25× the base input price for a five-minute lifetime, or 2× for one hour. Reading from it costs 0.1× on most models, 0.05× on Claude Opus 5.5 and Sonnet 5.5, and 0.025× on Claude Fable 5.1. These discounts stack with the 50% Batch API discount. Independent research backs this up: a January 2026 evaluation of long-horizon agent tasks, "Don't Break the Cache", found prompt caching cut API costs by 41–80% and improved time to first token by 13–31% across OpenAI, Anthropic, and Google.

For an illustrative example, take a 20,000-token system prompt plus tool definitions on Claude Sonnet 5.5, sent 100,000 times a month. That's 2 billion input tokens. Uncached, it costs about $4,000. Served from the cache at $0.10 per million, it costs about $200, plus occasional cache writes.

Beware: any change in the prefix can break the cache. A timestamp in the system prompt, unsorted JSON, a tool list that varies by request, or a per-user greeting placed before the shared instructions will each force a full-price rewrite. The right order fixes this. Put stable content first (tools, system instructions, reference documents) and volatile content last.

Don't confuse prompt caching with semantic caching, which returns a stored answer whenever a new question looks similar enough. It skips the model entirely, so it saves more, but it can be wrong.

"Cancel order" and "cancel subscription" can look alike to an embedding model, and a cached answer about last month's price is wrong today. Use semantic caching only for stable, read-only answers, and shadow-test it before launch.

How Do You Keep Agent Context and Loops From Growing Out of Control?

Replace raw transcripts with structured state, and give every iterative loop an explicit exit condition. Agents get expensive mainly because each turn re-sends everything that came before.

Long conversations hurt quality as well as cost. A study of more than 200,000 simulated conversations, "LLMs Get Lost in Multi-Turn Conversation", found an average 39% performance drop in multi-turn settings compared with single-turn ones. Most of the drop came from increased unreliability, with only a minor loss in aptitude. An agent that carries a compact state of the task (its objective, constraints, decisions so far, and next action) stays on track better than one replaying every message.

For loops, three controls cover most cases:

  • A hard cap on review or revision cycles, after which the system stops or escalates to a human.

  • A minimum-improvement threshold, so iteration stops once a pass no longer measurably improves the result.

  • Diffs instead of full artifacts. A reviewer checks the changed lines and the open issues, and skips the unchanged rest of the document.

Runtime supervision helps here. An ICLR 2026 paper, "Stop Wasting Your Tokens", reports that a lightweight supervisor agent cut the Smolagents framework's token consumption on the GAIA benchmark by an average of 29.68% without lowering its success rate. On the platform side, Anthropic now offers server-side compaction, which summarizes earlier context in long sessions, and context editing, which clears old tool results.

What Do Token Optimization Use Cases Look Like in Production?

Four common workloads show how these techniques map to real systems, each paired with the change that pays off most for it.

Customer support assistant with a large policy prompt

Answer customer questions using a long, stable block of policies, product facts, and tone guidelines. Use it in every inbound chat or ticket, at high volume.

The shared prefix often dwarfs the user's question, so caching it cuts most of the input bill. In the illustrative example above, that's about $4,000 a month reduced to about $200.

To make this work, you need a frozen prompt prefix with no timestamps or per-user content ahead of the cache breakpoint.

Agentic coding and code review pipeline

Agents can plan, write, and review code changes across a repository, working inside your delivery workflow, from ticket to pull request.

Review loops dominate cost. Passing diffs instead of whole files, capping review rounds, and putting a frontier model only on planning and final review keep spend tied to the size of the change.

For this to work, you need good tests, so output from cheaper models gets checked automatically.

Internal knowledge assistant over sensitive documents

Answer employee questions from wikis, contracts, and tickets. This might work in enterprise search, HR, legal, and operations.

Permission-filtered, reranked retrieval shrinks every prompt and keeps unauthorized content out of answers. But, you need document-level access control lists synced to the index.

Nightly enrichment and classification jobs

Tags, summarize, or extract fields from large volumes of records. Perfect for back-office pipelines that run overnight or on a schedule.

These jobs suit a small-tier model, structured outputs, and batch processing, which together reduce cost on several fronts at once. But you will need to build a tolerance for results that arrive hours later instead of seconds.

What Are the Risks and Trade-offs of LLM Cost Optimization?

Every cost technique saves money by removing something, whether context, model capability, or answer freshness, so the work is in removing only what the task didn't need. Plan for these frictions:

  • Integration. Routing and observability add components to your architecture. Every LLM call needs metadata, and every route needs a fallback when a smaller model fails. Without good tracing, a misroute looks like a random quality bug.

  • Quality regressions. Over-trimmed context drops information whose value appears only several turns later, so PostHog advises tuning on real traces. Changes need an evaluation set that runs before and after.

  • Compliance and data residency. Model choice is also a data decision. Anthropic notes that Claude Fable 5.1 requires 30-day data retention and isn't available under zero data retention unless Anthropic expressly authorizes it. Breunig observed that Fable's access controls and retention requirements pushed companies to think harder about where their traces go. In HealthTech, FinTech, and other regulated sectors, which model and endpoint may process which data is a compliance question first and a cost question second.

  • Maintenance. Model prices and capabilities change several times a year. Routing rules and cache layouts need an owner.

  • Knowing when not to optimize. If your monthly API bill is smaller than a week of engineering time, a routing layer probably won't pay for itself yet. Turn on caching, set sensible output limits, and revisit when volume grows.

Key Takeaways

  • Token efficiency is an architecture problem; rewording prompts rarely moves the bill.

  • Measure cost per completed task, broken down by workflow stage, before you change anything.

  • Model tiers can differ 100× in input price, so frontier models should handle only the work that needs them.

  • Prompt caching is the lowest-risk lever, but a single changing byte in the prefix breaks it.

  • In multi-agent pipelines, re-reading context can cost more than generating output, so pass state and diffs instead of transcripts.

What Does Sustainable Token Efficiency Require?

Lasting token efficiency comes from architecture you can measure: per-task cost data, an expensive model used only where its judgment pays for itself, and context windows sized to the step at hand. Evaluations keep each saving honest, because a cheaper call that degrades answers is a cost moved somewhere else.

Model prices will keep changing several times a year, and teams with observability and routing already in place can respond with a configuration change instead of a rebuild. If your AI costs are growing faster than the value you get from them, Monterail's AI engineering team can audit your existing AI infrastructure and implement production-grade token management where the numbers justify it. Get in touch to start with a diagnostic.

LLM Cost Optimization FAQ

Maciej Korolik
Maciej Korolik
Senior Frontend Developer and AI Expert at Monterail
Linkedin
Maciej is a Senior Frontend Developer and AI Expert at Monterail, specializing in React.js and Next.js. Passionate about AI-driven development, he leads AI initiatives by implementing advanced solutions, educating teams, and helping clients integrate AI technologies into their products. With hands-on experience in generative AI tools, Maciej bridges the gap between innovation and practical application in modern software development.