AI Costs Need an Owner Before Production

The FinOps NPCs
AI Costs Need an Owner Before Production

Every organization adopting AI gets the same cast of characters. They are not villains. They are people doing their jobs inside a financial system that was not built for this kind of spending.

Traditional software costs are easier to place in the mental model. A service has an account. A team owns it. A budget gets approved. Usage may grow, but the shape of the expense is familiar.

AI costs do not behave so neatly. They can spread across teams, vendors, model providers, API keys, and corporate cards. They can rise with usage in ways that are hard to predict. They can look small while a workflow is a pilot and become a serious operating expense once people depend on it.

The accounting problem often begins before anyone realizes there is an accounting problem.

The Avalanche Skeptic

The Avalanche Skeptic carries the pilot results like a lawyer carries favorable precedent.

The pilot was cheap. The pilot worked. Therefore the rollout will be cheap and will work.

The Skeptic is not wrong about the pilot. They are wrong about what the pilot proves.

Ten users for six weeks is not twenty-five thousand users for a year. A small test can show that a workflow is possible. It can show that users will try it. It can reveal obvious failures. It cannot automatically tell you what happens when usage expands, inputs become less predictable, or the workflow becomes something people rely on every day.

The pilot is a photograph of a river before the flood.

Scale changes more than the bill. It changes support needs, quality expectations, retry behavior, monitoring, and the cost of bad output. A result that is acceptable during an experiment may create expensive rework once it sits inside a busy process.

The Skeptic wants the rollout to be a continuation of the pilot. FinOps needs to treat it as a new financial question.

The Token Optimizer

The Token Optimizer treats cost reduction as an optimization problem with a clear solution.

Use cheaper models. Shorten prompts. Cache responses. Reduce token spend.

Each move can be correct in isolation. The problem is the total system. It can be tuned to produce worse answers slightly cheaper, which is not the same as reducing cost.

The Optimizer measures the bill from the model provider but not the cost of degraded output. The extra round-trips count. So does rework. So does the user who switches to a different tool because the answers got worse. So does the engineer who has to repair a workflow that saved a few cents on each request.

A cheaper response is useful only if it still does the job.

That does not mean model cost should be ignored. It means token spend is one measurement, not the whole result. Compare the cost of the workflow with the work it replaces or creates. If a cheaper model produces an answer that needs a second pass every time, the first request was not cheap. It was only incomplete.

The Optimizer is good at finding local savings. Someone else has to protect the outcome.

The Budget Keeper

The Budget Keeper discovers, three months into the rollout, that nobody owns the AI budget.

Not in the sense that nobody is responsible. In the sense that the budget is distributed across six teams, three vendors, and a corporate credit card engineering uses for API access.

The Keeper's job is to assemble one number from fragments. The number is always larger than anyone expected, and it is always three months old by the time it reaches someone who can act on it.

This is not a spreadsheet failure. It is an ownership failure. If nobody names the workflows, accounts, vendors, and people responsible for the spend, the organization cannot tell whether a change in cost comes from growth, a new model, a retry loop, a vendor expansion, or a workflow nobody remembered approving.

By the time the bill becomes visible, the system has already normalized the spending.

A useful budget practice starts before production. For each significant workflow, record who owns it, which provider it uses, what drives usage, what a normal month should look like, and what happens when cost moves outside the expected range. The numbers will be imperfect. They will still be more useful than a delayed surprise.

Name the owner first

The pattern across all three NPCs is simple: AI cost does not fit the old infrastructure model.

It does not sit in one account. It does not scale linearly. It does not map cleanly to the categories finance uses for software spending. The NPCs are doing their jobs inside a system designed for predictable software costs and a world where AI spending was not yet a line item.

The counter is not necessarily a new tool. It is a new practice: name an owner for every significant AI workflow before it goes to production.

Not after the bill arrives. Before.

That owner does not need to control every part of the organization. They do need to know what the workflow costs, what value it is meant to provide, and who has authority to change or stop it. Without that person, every cost discussion becomes archaeology.

>