The Real Cost of AI Isn’t the Token

Token prices are falling, but AI bills can still rise. CFOs should manage cost per business outcome, not just cost per token.

Your AI provider cuts model prices by 40%. Good news.

Does your AI process now cost 40% less?

Probably not.

This is where AI economics becomes more interesting for finance. Tokens are easy to count. They appear on invoices, make model comparisons possible, and seem to provide a universal cost unit.

But companies do not consume tokens for their own sake. They want to resolve a customer request, process an invoice, qualify an opportunity, complete a reconciliation, or produce a report.

> 30-second takeaway
> A token measures technical consumption, not business value. To manage AI economics properly, CFOs need to move from cost per token toward cost per completed outcome, including models, tools, integrations, retries, exceptions, and human intervention.

The price can fall while the bill keeps rising

Stanford’s 2025 AI Index measured a striking decline in inference prices. Between November 2022 and October 2024, the cost of running a model at roughly GPT-3.5-level performance on MMLU fell from about US$20 to US$0.07 per million tokens — more than a 280-fold decline.

Meanwhile, usage is expanding rapidly.

The State of FinOps 2026 reports that 98% of respondents now manage AI spend, up from 31% two years earlier. AI cost management is also the number-one skill set those teams say they need to develop.

There is no contradiction.

When something becomes cheaper to use, organizations usually find more ways to use it.

The mistake is assuming that lower unit prices automatically mean lower total costs.

One AI outcome can hide a lot of consumption

Consider a customer service request.

At first glance, you might measure the cost of the model that reads the email and drafts a response.

But the actual workflow may also have to:

  • retrieve the customer record;
  • query a CRM or ERP;
  • load supporting documents;
  • call another model for a specialized task;
  • validate the result;
  • update a business system;
  • retry a failed step;
  • route an exception to a person.

The FinOps Foundation notes that agentic workflows can multiply consumption through context, tool calls, reasoning models, and retries. It also stresses that tokens represent only one layer of the broader AI cost stack.

> The management point
> Model cost is only the beginning. The useful economic unit sits at the end of the process: what did the successfully completed outcome actually cost?

Exceptions may be your most expensive line item

Suppose an agent handles 900 out of 1,000 cases automatically.

A 90% success rate sounds excellent.

But what happens to the remaining 100?

If each exception requires fifteen minutes of human intervention, corrections across two systems, and occasional full retries, those cases may represent a disproportionate share of the total cost.

That is where “average cost per call” starts to mislead.

A workflow is not economical simply because its successful executions are cheap. It is economical when successes, failures, retries, and exceptions collectively produce an acceptable cost per outcome.

Move from tokens to unit economics

The shift is simple:

Total AI workflow cost ÷ successfully completed outcomes = cost per outcome

The outcome depends on the business.

It might be an invoice processed, a service request resolved, a reconciliation completed, a lead qualified, or a report produced and approved.

This metric does not replace token tracking.

It gives token tracking a purpose.

| Easy to measure | CFOs should also measure |
|---|---|
| Tokens consumed | Cost per successful outcome |
| Model cost | Total workflow cost |
| Number of calls | Completed outcomes |
| Budget consumed | Value or capacity created |
| Usage | Exception rate |
| Average cost | Cost of failures and retries |

The distinction is simple: measure what the meter records, but manage what the business produces.

No one would evaluate a factory using only the price of electricity.

AI deserves the same economic discipline.

An AI budget is not the same as AI control

The problem becomes harder as agentic AI spreads across departments.

Finance may see aggregate spend without knowing which process generated it, which part created value, or how much came from retries, poor model selection, or excessive consumption.

In an August 2026 analysis, the FinOps Foundation made this exact point: Finance cannot set the AI envelope alone using gross spend because that number may contain unmeasured leakage and says little about which costs are recoverable.

In other words:

A budget is not yet a control.

Control requires attribution.

Five questions every CFO should ask

You do not need to become an expert in tokens or model architecture.

Start with five questions:

  1. What business outcome is this AI system supposed to produce?
  2. What does one successfully completed outcome cost us?
  3. What percentage of executions requires human intervention?
  4. Which exceptions and retries generate the most cost?
  5. Does rising AI consumption actually produce more revenue, capacity, savings, or customer value?

If the team can answer the first question but not the next four, the problem probably is not token pricing.

It is measurement.

What you can do Monday morning

Choose one AI use case already in production.

Take last month’s total cost.

Instead of dividing it by users, requests, or tokens, divide it by the number of completed business outcomes.

Then add the human time spent handling exceptions.

> 30-minute exercise
> Compare cost per outcome over the last three months. If it is rising, ask why: more value created, a more expensive model, more retries, more exceptions — or simply more consumption?

You will start seeing AI the way Finance ultimately needs to see it:

not as another technology bill,

but as a new unit of business economics to manage.

Sources