The CIO's Real Job on AI Cost is Managing Demand, not Cutting Spend.

06 Aug 2026
Bailey Caldwell
[
Head of Strategy & GTM
]
Share
The CIO's Real Job on AI Cost is Managing Demand, not Cutting Spend.

McKinsey QuantumBlack put a useful stake in the ground. In The cost of intelligence, the firm argues AI cost is a demand management problem rather than a procurement problem. CIOs who spend their weeks shopping for rate cards and swapping models are working the easy half of the equation. The hard half is knowing which teams, agents, and workflows are actually generating value for the tokens they burn, and having the controls to shape that demand while it is happening rather than after the invoice arrives.

We agree with the framing, and we would go one step further. Demand management has to be a system, not a slide. Most enterprises don't have that system, and the gap is starting to show up in the numbers.

The market has moved faster than the tooling

The FinOps Foundation's State of FinOps 2026 puts AI spend management at 98 percent of respondents, up from 31 percent two years ago. FinOps for AI is now the top forward-looking priority in the report. In the same window, Flexera found that AI waste drove cloud spend up roughly 29 percent. Gartner expects more than 40 percent of agentic AI projects to be canceled by the end of 2027 for unclear ROI and unmanaged cost. ICONIQ's 2026 State of AI puts AI product gross margins at 53 percent this year, tracking toward 59 percent by 2027, against the 80 to 90 percent that classic SaaS was built on.

All of that lands as one CIO problem. Budget authority for AI is arriving on your desk without the demand-side tooling to govern it. Procurement discounts move the number a little, and model swaps move it a little more. However, neither lever will tell you which workflow is worth funding, and more importantly, neither can stop a runaway agent.

McKinsey names the problem and points at the levers, then leaves the operating model to the reader. There is no per-team, per-workflow, per-outcome attribution in the piece, no mention of runtime enforcement, no picture of what a working demand-management layer looks like in practice. That is the part we want to fill in, because we have been building it.

The four things a real AI demand-management layer has to do

Strip the buzzwords out, and a working demand-management layer has to do four things. Every serious CIO conversation we are in this year comes back to some version of this list.

1. Attribute. Every AI action needs to tie back to identity, cost, feature, customer, and workflow at the event level. Cloud tagging maxed out around 72 percent accuracy after three years of enterprise effort, and AI moves faster and fans out further than any workload FinOps ever tried to tag. If you can't decompose a single user action into the model calls, tool calls, third-party APIs, and human review steps it triggered, you can't manage its demand; you can only pay its bill.

2. Forecast. Once every event is attributed, you can predict spend by team, workflow, and agent before the provider invoice explains it to you. Most native dashboards skip this step. They show you what happened, not that your cost per call is climbing faster than usage, or that a prompt change quietly regressed the price of a feature last Tuesday.

3. Enforce. Guardrails only work as software that runs when an agent makes a decision. A governance PDF does not stop a runaway loop, and a monthly finance review does not stop a bad workflow from spending real money for another 29 days. Enforcement has to live in the runtime, next to the call, with hard limits, budgets, circuit breakers, and policy checks that block spend before it happens.

4. Correlate to outcome. Cost per token is an infrastructure metric. Cost per outcome is an economic metric. Demand management is only real when the CIO, the CFO, and the product owner can look at the same view and answer one question. For every dollar of AI we spent this week, what did the business get back? Without that answer, cutting AI cost is guesswork.

Why provider dashboards and cloud FinOps tools do not close this gap

Provider dashboards stop at the api_key_id. They can't decompose a call into per-response sub-passes, tool invocations, or business outcomes, because the provider does not know what the call was for. Cloud FinOps tools were built for allocated cloud bills, so they ingest AI as one more line item to normalize instead of a stream of per-event economic records. Model routers optimize the token price, which is useful, but a router doesn't see the tool call it just triggered as an economic event. And tagging, as noted, has already been tried at scale and didn't get there.

Every one of these approaches works bill-up. Demand management has to work event-up.

How Revenium handles each of the four

We built Revenium as an AI Economic Control System. It is the system of record for AI usage, cost, and unit economics, with real-time controls attached. That maps to the four requirements above.

Attribute with Tool Registry. Tool Registry meters every model call, tool call, external API, and human review step as a first-class priced event, joined to agent identity, workflow, customer, and feature. In one fintech pre-approval workflow we instrumented recently, the token line on the provider bill showed roughly two cents per call. The real cost per call once the tool APIs and downstream steps were included was closer to $1.67. That is a 72x gap, and it is invisible without event-level attribution.

Forecast with AI Insights. AI Insights runs continuously against your Revenium data. It detects configuration waste, failure-driven cost, concentration risk, cost-to-outcome misalignment, better-alternative models, unexpected growth, and attribution gaps, and it ranks findings by recoverable dollars with direct links to the transactions that need to change. In 2.17.0, it also explains spikes on its own and alerts when cost per call is climbing faster than usage.

Enforce with Cost Controls and unified Guardrails. Cost Controls block budget-breaking calls in real time, before the money is spent. Unified Guardrails in 2.17 pull budgets, hard limits, and policy checks into one runtime layer that the agent has to pass through to spend. This is the piece that turns demand management from a report into a control loop.

Correlate to outcome with AI Outcomes. AI Outcomes ties every agent execution to a business result and calculates ROI at the transaction level. It deliberately separates execution status from business outcome, so a technically successful workflow that produced nothing of value does not hide inside a green dashboard. When the CIO, CFO, and product owner sit down, they look at cost per outcome, not cost per token, and the AI budget conversation becomes grounded in data.

What the CIO actually owns

McKinsey is right that the CIO's real job on AI cost is managing demand. We would add that dashboards will not do it. Event-level attribution, continuous forecasting, runtime enforcement, and outcome correlation, wired into the same system, will. That is the layer we have been building, and every enterprise running AI in production will need some version of it.

If you want to see it, book a Revenium walkthrough. If you would rather start from code, our open-source SDKs are at github.com/revenium and the free developer tier is at app.revenium.ai/sign-up.

Table of Contents
Ship With Confidence
Sign Up
Ship With Confidence

Start with visibility. Scale with control.

100,000 transactions free. No credit card required.