Cost pressure
Cost optimization
Cloud cost work, FinOps if you prefer the term, is not a procurement exercise. I map where spend is actually growing across AWS, Azure, GCP, Cloudflare and DigitalOcean - infrastructure, AI usage, APIs, CI, environments, and observability - and turn it into an ordered reduction plan with clear trade-offs.
How longTwo weeks to the findings. Acting on them runs longer, with your team or with me.
The work
What this looks like
Marked up the way I would send it back.
-
Kubernetes, three environments EUR 4,200 / mo
Ask Satyam
Staging and preview run the same node sizes as production. Nobody chose that, it was copied when the cluster was created and never revisited.
-
Cross-region data
egressEUR 1,800 / moCut Satyam
Egress is never the problem, it is the receipt. Something is reading across regions on a request path, and finding what takes an afternoon. Paying the bill monthly is the expensive way to avoid that afternoon.
-
Observability, 90-day retention on everything EUR 2,600 / mo
Ask Satyam
When did anyone last open a debug log older than a week? Retention is a compliance question for some of your data and a habit for the rest, and the two are being billed at the same rate.
-
LLM API,
one model for every callEUR 3,100 / moCut Satyam
Classification, extraction and summarization are not the same workload and do not need the same model. Route by task and the bill usually falls by half without a measurable change in output quality. Measure that last part rather than trusting it.
-
Reserved instances
bought in 2023EUR 900 / moCut Satyam
Reservations outlive the architecture that justified them. This one is paying for a shape the system no longer has.
Same platform, same team. The bill above falls by roughly a third before anything is rearchitected, and the part that needs rearchitecting is now a decision rather than a surprise.
How it runs
How the two weeks go
For teams whose cloud, AI or API spend has grown faster than the product has, and who want it reduced without the blunt across-the-board cut that quietly costs a quarter of delivery.
-
01
Find where it is actually growing
Not the biggest line, the fastest-growing one. Infrastructure, AI usage, third-party APIs, CI minutes, environments and observability, each attributed to the team and the change that caused it. Most bills have three real drivers and a long tail that is not worth touching.
-
02
Order the cuts by what they cost you
Every reduction has a price in risk, latency or team time. That price goes next to the saving, in writing, before anything is switched off. The plan comes back sequenced, so you can stop after the first three if the number is already enough.
-
03
Leave the model behind
The spend model, the attribution and the alerts stay with your team. If the same drift is back in a year because nobody could see it happening, the engagement did not work.
What you end up with
- An ordered reduction plan with the trade-off written next to each line
- The three changes worth making first, and what they are expected to save
- A spend model your team can keep running after the engagement
Where this has been done
AWS since 2013, GCP since 2016, DigitalOcean since 2019, Cloudflare since 2020 and Azure since 2023, from a regulated bank's pipelines to consumer data platforms holding more than 300 million people. Most of the saving came from building the data pipelines rather than renting them as managed services.
Bring me the line item nobody can explain.
Book a discovery callThirty minutes, and a straight answer on whether this is the right path.