AI FinOps · 5 min read
Cutting LLM costs without cutting quality
Most LLM bills are dominated by a few avoidable patterns. Routing, caching, prompt hygiene and attribution can bring costs down sharply while your eval scores stay flat.
When an organisation’s LLM bill starts to hurt, the first instinct is to switch everything to a cheaper model. That usually trades a cost problem for a quality problem. There are better levers — and you should pull them in this order.
1. Attribute before you optimise
You can’t cut what you can’t see. Tag every request with the team, feature and model behind it, and send the usage to a dashboard. Most organisations discover that a small number of features drive most of the spend.
2. Trim the prompt
Long system prompts, oversized retrieval contexts and full conversation histories are silent cost multipliers. Retrieve fewer, better chunks. Summarise old turns instead of resending them. Remove instructions nobody can trace to a real problem.
3. Cache what repeats
Many requests are near-duplicates: the same FAQ, the same document summary, the same classification. Cache responses for identical inputs, and use your provider’s prompt caching for long, stable prefixes.
4. Route by difficulty
Not every request needs your most capable model. Send simple, well-defined tasks — classification, extraction, short rewrites — to smaller, cheaper models, and reserve the large ones for complex reasoning. Use your eval set to prove that each route holds quality.
5. Set budgets and alerts
Give each team a budget and alert them when they approach it. Spend that people can see is spend they manage.
The rule
Every optimisation should be checked against your evaluation set. If the score stays flat and the cost goes down, ship it. If the score drops, you’ve found the real price of that saving — and can decide deliberately whether it’s worth it.