Cloud Spend Took Years to Get Out of Control. AI Spend Will Do It in Months.

27 Aug 2026
John Rowell
[
CEO, Co-founder
]
Share
Cloud Spend Took Years to Get Out of Control. AI Spend Will Do It in Months.

You already run cloud cost. You have tags, showback, a reserved instance strategy, and someone whose job is to argue with engineering about instance sizing. So when AI spend shows up on the bill, the plan is obvious. Point the same discipline at it and work through the backlog.

That plan is sound and it is also about two years too slow.

A FinOps lead at a Fortune 100 retailer ran the comparison against their own history and told us what came out. "I looked at how long it took us to get to a million dollars of spend in the cloud and it's measured in years... about 15 months... I think we got less than maybe another quarter to get there with AI spend." Fifteen months against one quarter, inside the same company, with the same finance team and the same approval process.

The cloud playbook works. The problem is that it was built during a period when you had time to build it.

One API key, one week

The compressed timeline shows up at ground level as a single API key. At one large software company, a key issued to a single engineer ran well past its expected usage and burned through roughly $93,000 in a week before anyone caught it. The GenAI lead who walked us through it was not describing a breach or a bad actor. One key, issued to one person for legitimate work, misconfigured in a way that nobody noticed until the number was large enough to notice on its own.

Compare that to how the same failure behaves in cloud. Someone leaves an oversized instance running, and it costs you a few hundred dollars a month while it waits to be found. The AI version of that mistake reaches five figures before the weekly standup.

The visibility gap is what turns the mistake into the number. A security company we work with got a bill from Bedrock roughly a month after the spend happened, and by then there was nothing to do but pay it. Their COO was blunt about what would have changed the outcome.

"We could have cut that off immediately on day two had we had any visibility into this."

Day two versus day thirty is the entire story. Nothing about the underlying mistake changed. Only the lag did.

Why AI compounds faster than cloud ever did

Three things make AI spend behave differently, and none of them are about price per token.

The first is that consumption has no natural ceiling. Infrastructure is bounded by what you provisioned. A model call is bounded by how many times something decides to call it, and agentic workloads decide to call it a great deal.

One infrastructure software company’s FinOps lead put the difference in terms of what the balance sheet can survive.

"In cloud you could spend a lot, but you could still afford it...With AI, you just won't be able to afford it. It will bankrupt your company so quickly."

The same person told us that a ten million dollar forecast would not constrain them, because with the budget available they could find twenty million or fifty million worth of things to spend it on.

The second is that adoption is not gated the way cloud was. Provisioning infrastructure required a ticket and a person who knew how. Calling a model requires a key.

At one cybersecurity company, AI access was open to any business unit when we spoke with them, which removed the onboarding friction and also removed the natural checkpoint where somebody asks what this is for.

The third is growth rate. The GM of the finance product at an HR and payroll platform told us their AI line was growing 80% month over month and was on track to consume the R&D budget.

At a customer support software company, the run rate had reached roughly twelve million dollars, and the engineering lead there described the growth rate rather than the total as the thing they could not get under control. A monthly total tells you what happened. A compounding rate tells you what is about to.

The retailer's FinOps lead had already drawn the operational conclusion.

"We started eight, nine years after we migrated to the cloud to actually start trying to manage and control it. And I don't think we've got that luxury with AI, because if we wait even six months, it will be out of control."

Five signals to check in your own environment this week

None of these require a purchase or a project. They require an afternoon.

Can you break spend down by team, or only by bill? One university’s director of AI infrastructure runs thirty-four separate AWS accounts, one per project, because that is what attribution looks like when the platform will not do it for you. At a fleet management company, the AI strategy lead found the Azure portal so hard to pull usable telemetry from that they went to their AI committee asking for full observability rather than better reporting.

How long is your detection lag? Measure it honestly. Take your last unexpected spike and count the days until somebody named it. A FinOps team of one at a large SaaS company described seeing a weekend spike and then waiting while engineering dug through workloads trying to work out what caused it. If your answer is anything past forty-eight hours, that window is your real exposure.

Is any single credential uncapped? Go find the keys issued to individuals for experimentation. Check whether any of them has a hard limit attached. This is the exact shape of the one-week incident above.

Do you know your month-over-month growth rate? If you can’t state it from memory, you can’t forecast the quarter, and neither can your CFO.

Do you have a cap anywhere, or only alerts? A FinOps practice lead at one of our partners described the governance at a client where crossing ten thousand dollars a month for a single user puts you on a list without cutting anything off. A founder who previously ran security at an enterprise spend management company was sharper about why the gap exists at all.

"All of the foundational model providers natively have the budgets and alerts. What none of them have is the ability to cap it... The foundational model providers are not going to do it because it's in their best interest not to do it."

Where we fit

If those five checks surfaced something, the ones worth closing first are attribution and the cap.

Cost Controls lets you define rules that block AI requests once configured usage limits are reached, which is the piece the providers leave out.

Set Budgets & Alerts covers thresholds and notification routing if you want a warning before enforcement.

Attribute Spend to Teams & Departments is how you stop running one AWS account per project to answer the question of who spent what.

Predict & Surface Anomalies shortens the detection lag.

AI Insights runs your usage through detectors and returns findings ranked by potential monthly savings, each with a suggested action, so the first week produces a list rather than a dashboard.

Run the five checks before you talk to anyone, including us.

Table of Contents
Ship With Confidence
Sign Up
Ship With Confidence

Start with visibility. Scale with control.

100,000 transactions free. No credit card required.

Start for free
Send
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.