We've gotten pretty good at looking at the cloud bill. Maybe too good. There are dashboards, reports, allocation models, tags, budgets, forecasts, anomaly detection, showback, chargeback… We have entire teams whose job is to figure out where the money went.
But as I’ve been digging into my scale-to-zero evangelism I’m finding there's a funny thing about the cloud-cost conversation… we've gotten much better at seeing waste than we have at preventing it.
I guess one thing that surprises me is that people haven’t gotten rid of the waste, they just aren’t surprised it’s there anymore… Harness's 2025 FinOps in Focus report surveyed 700 developers and engineering leaders at large organizations in the US and UK.
The report estimates that 21% of enterprise cloud infrastructure spending is wasted which is a about $44.5 billion in 2025. But when you dig into the report the really interesting numbers aren't even the $44.5 billion.
They're the ones underneath it.
- Only 43% of respondents said they had real-time visibility into idle cloud resources.
- Only 39% had real-time visibility into unused or orphaned resources.
- And only 33% had visibility into over- or under-provisioned workloads.
In other words, we've built this enormous industry around cloud cost management and we're still struggling to answer the most basic of questions: What's running right now that doesn't need to be?
I’ve said this before, it’s not that developers don’t care about cost. The report says 62% of developers said they want more control over and responsibility for managing cloud infrastructure costs.
But what they aren’t doing is thinking about how to scale-down… how to honor rule number 1 of the flat-stack manifesto: do not wake the beast. To be honest, even with AI dev tools (or maybe because of AI dev tools) engineers are going to follow the same full-stack, idle compute complexity that we’ve been doing for the last 20 years because that’s the engineering solution… And FinOps is just a way for Finance to make meetings uncomfortable…
The Harness report found that organizations take an average of 31 days to identify and eliminate cloud waste and then It takes about 25 days to detect and rightsize overprovisioned resources.
Think about that for a moment.
You get the wrong architecture in there, maybe too big of an instance, maybe you’re scaling up when you don’t need to… and it takes two months to start to reduce the bill. Half a financial quarter showing that cost on the books…
That's not really a visibility problem anymore, it’s kind of the opposite… it’s like the old bluegrass song about the steamboat running aground in full view of town… everyone sees it and you can’t do anything about it.
And the report shows that only 32% of organizations have fully automated practices for enforcing cost savings — things like automatically shutting down idle resources or rightsizing them.
The reason they often don’t have right-sizing is because the architecture NEEDS that compute… just not all the time.. but if you scale it down when you actually need it, it’s just easier to not scale down…
So, yeah, we’ve got a lot better visibility into the compute costs, but the pattern is awfully familiar… something wastes money… somebody notices… somebody investigates… somebody creates a ticket… somebody changes something. Two months later that cost item goes down… but another one has popped up.
We've spent years getting better at making continuously running infrastructure cheaper, when what we should be doing is not running it in the first place. I’m not saying it’s easy, I’m not even saying it’s obvious, but I am saying it’s important. It’s important to the bottom line of the enterprise, it’s important to control data center sprawl, it’s important to just have less complicated stuff looming over us all the time.
And I think the FinOps movement has done something important. It made engineering organizations take financial consequences seriously and it gave people common language while exposing that invisible waste.
And the FinOps in Focus report makes a good point that organizations need to shift cost awareness left, so developers can understand the financial consequences of architectural decisions before those decisions become expensive.
I mean that’s pretty much what I’m evangelizing but… Don't just move the dashboard left, give the engineer the ability to control this stuff sooner and easier. Use all that finops data to help shape the next generation of architecture into the flat-stack model so the need to right-size never comes up in a conference call again.
Think about compute… we need compute but we don’t need the whole warehouse running… a gateway can accept the request, answer from a cache when it can, and wake the expensive data infrastructure only when the request actually requires it.
Airbrx is one example of this pattern for data workloads, and yes, I’d love for you to give it a try, but the important idea isn't about Airbrx, it's about the separation of availability from compute.
The availability always-on stuff stays cheap, the big heavy lifting is only fired up when you need it. Visibility is still critical (you should see the log details we keep on the data gateway…)
Build it to scale down from the beginning, don’t just wait for the report to tell you what’s wrong. The report is great, keep the reports by all means, but don’t confuse the report with the solution.
This is the promise of flat-stack architecture… Don't just measure waste. Design it out.