Where we startedThe product had scaled faster than cost ownership

Spend had risen slowly enough to look like normal growth. A budget review forced the question: what exactly was driving it? We had the billing data, but it arrived as one total. We could not reliably connect a meaningful share of the bill to a team, service or design decision.

The diagnosisDifferent costs needed different fixes

We separated the bill into three groups. Waste included idle environments, orphaned storage and retention nobody had chosen. Poor fit covered services and capacity models that no longer matched the workload. Missing ownershipmeant resources that would return after any cleanup because no team was responsible for them.

I set two boundaries for the program: do not make delivery slower and do not make production harder to understand. That ruled out broad environment freezes, deferred maintenance and cutting observability to improve the bill.

What the teams changedFix the bill and the operating model

The first change did not save money. We introduced tagging and allocation so meaningful spend traced back to a service and team. Once engineers could see their own number, cost reviews became engineering conversations instead of debates over a consolidated bill.

Each team then worked its own largest drivers. We changed architecture where the service model did not fit the load, added lifecycle and retention rules, scheduled non-production capacity and right-sized from measured utilization. We also removed existing waste, but treated that as cleanup rather than the program's main result.

Team reviews included cost next to delivery and reliability. This was important for scale. A central group can clean an account once. It cannot make every future architecture decision for every product team.

This took engineering time that could have gone to features. I made that tradeoff visible and agreed priorities with product leaders instead of asking teams to hide the work between other commitments. Some fixes also added operational complexity, so we only kept changes whose savings justified the new burden.

We did not reduce the signals needed to operate production. A cheaper system that is harder to debug would have moved cost from the cloud bill into incidents and engineer time.

The result40% lower cost, with ownership attached

Infrastructure cost fell by 40% while teams continued to ship and operate reliably. New spend did not stop, because the products were still growing. The difference was that it appeared with a team, a service and a reason attached. Drift became visible while it was still small.

What I learned

I would keep cost ownership with the teams making design decisions. The technical fixes were available before the review. What was missing was visibility and responsibility.

I would establish attribution from the first production service. Tagging is easy at creation and tedious across hundreds of resources. I would also add a rough operating-cost expectation to design reviews before the product reaches meaningful scale.

The 40% mattered. The more durable result was a team that could explain what it cost to run its product and act before waste became a program.