SaaS Tech Watch
saas

FinOps for AI: unit economics, caching, and model routing as finance controls

Every figure states its provenance: measured (we ran it) · reported (vendor says) · derived (we calculated).

Cloud FinOps taught engineering teams to treat compute as a budget line, not background cost. AI inference is now forcing the same lesson on product teams. In 2025 the inference bill moved from “engineering expense, tracked loosely” to “a named line item, with a review meeting.” This piece is the playbook we consider table stakes: unit economics before shipping, caching tiers in the architecture, model routing as a cost lever, and budgeting practices that keep finance and engineering pointing at the same numbers.

Nothing here requires a new vendor. It requires instrumentation and a willingness to look at the bill per feature, not per account.

Start with unit economics, before shipping

“How much will this feature cost” is unanswerable at the granularity of monthly bills. The right unit is cost per user-visible operation:

  • Cost per document summarised
  • Cost per draft email generated
  • Cost per agent task completed
  • Cost per search query answered

Compute this from three numbers the team already has or can measure: tokens per operation (input + output), price per token at the chosen model tier, and expected operations per active user per period.

A worked example, using made-up illustrative numbers chosen to make the arithmetic obvious — these are not any vendor’s prices: if a feature averages 2,000 input tokens and 800 output tokens per call, and the model costs $3 per million input tokens and $15 per million output tokens, each call costs about $0.018. At 20 calls per active user per month, that is $0.36 per active user. Whether that is acceptable depends on the seat price — which is exactly the conversation this arithmetic enables.

The discipline is doing this calculation before building, not after launch.

The four levers, in order of impact

Once a feature is live, cost is governed by four levers. In our experience the ordering matters:

LeverEffect sizeImplementation costWho owns it
Model routing (task → cheapest capable model)LargeLow–mediumPlatform
Caching (semantic / prefix / response)Medium–largeMediumPlatform
Prompt and context disciplineMediumLowProduct eng
Batch and off-peak schedulingSmall–mediumMediumPlatform

A team that has done none of this is almost always leaving its largest cost reduction on the model-routing row, not the caching row.

Model routing is the finance control

Most AI features as shipped route every request to one model, usually the strongest available. Production usage does not justify that. A share of requests — classification, extraction, short-form generation — is handled well by a smaller, cheaper model; a share genuinely needs the frontier tier.

Routing rules do not need machine learning. A simple decision tree on task type, expected output length, and whether the user is on a paid tier captures most of the saving. The failure mode is silent quality regression; routing decisions need quality evals alongside cost dashboards, or the finance win becomes a product loss.

Routing also separates “cost at capability ceiling” from “cost per request.” A team can serve 90% of traffic on a cheap model and 10% on a frontier one, and both numbers can be true in the same dashboard. Finance cares about the blended cost; engineering cares about the tail behaviour. Both views need to exist.

Caching tiers, concretely

Caching takes most of the engineering effort once the routing wins are taken. Three tiers: response caching (identical prompts return stored responses; high hit rate in FAQ bots and common summaries), prefix caching (shared system prompts cached at the provider, cutting input-token cost), and embedding caching (for RAG, cache embeddings of stable documents).

Each tier has a correctness story. Response caching is dangerous where staleness matters. Prefix caching depends on provider support. Embedding caching depends on document change rates. None is free. All are cheaper than naive per-call full-context inference.

Budgeting practices that survive the quarterly review

Three budgeting habits separate teams that get surprised from teams that do not:

  1. Cost per feature, per week. The monthly invoice is too coarse. A weekly per-feature cost review catches drift within one sprint, not within one quarter.
  2. Budget alerts at the feature level. A global token-budget alert tells us “the account is burning”; a feature-level alert tells us which feature. Set them separately.
  3. A margin model, shared with finance. Engineering owns cost per operation; finance owns margin per seat. A one-page shared model with both, refreshed quarterly, prevents the post-mortem where each side was optimising for a different number.

What to model before shipping an AI feature

Before committing to build, a one-page cost model should answer:

  • What is the unit operation?
  • What is the average cost per operation under expected usage?
  • What is the same cost at 10x expected usage, and at 100x?
  • Which of the four levers (routing, caching, prompts, batching) are available if cost drifts above budget?
  • What is the kill-switch story — can the feature be disabled per-tenant, per-user, or globally without redeploying?

If any of these cannot be answered, the feature is not ready for cost review. That is separable from whether it is ready for users.

The bottom line

FinOps for AI is not a separate discipline; it is FinOps applied to a more granular meter. Teams managing it well measure cost per feature per week, model before they ship, and treat routing and caching as engineering primitives. Teams getting surprised treat the inference bill as a single number and discover what drove it only after finance escalates.

Related reading

from the desk ▸