Azure Cost Optimization
Control cloud cost without flying blind
How Azure cost optimisation actually works, across VMs, reservations, savings plans and ongoing monitoring.
A while back I was going through Cost Management for a tenant I’d never touched before, and found a Standard_D8s_v5 sitting at 12% average CPU. Nobody had looked at it in months. That VM wasn’t an isolated case, it’s basically the default state of every Azure subscription I’ve opened that didn’t have someone actively watching it. So let’s talk about why that happens and what actually fixes it.
Here’s the thing: it’s rarely a pricing problem. It’s a visibility problem. Consumption lives in Cost Management, but the decisions that actually drive it, sizing, redundancy, region, scaling, get made in architecture meetings that never look at the bill. By the time someone in finance asks why the invoice moved, the VM that caused it has already shipped and everyone’s moved on to the next thing.
I’ve opened enough Cost Management blades to know the fix isn’t a one-off cleanup. It’s wiring cost into the same feedback loop you already use for uptime and security, something you check continuously with tooling, not something you rediscover once a quarter. Let me walk you through how I actually do it.
Cost is a side effect of your technical decisions
Every architecture choice you make has a price tag attached, even when nobody wrote it down:
- Sizing and SKU choice. That D8s_v5 I mentioned isn’t cheap just because it’s stable. It’s an unclaimed D4s_v5 with extra steps.
- Redundancy tier. GZRS storage costs roughly double LRS. Correct call for a production ledger, waste of money on a dev container.
- Data movement. Egress and cross-region traffic are the charges you forget to model until a monthly bill jumps for no obvious reason.
- Retention. Log Analytics and Application Insights bill per GB ingested and per day retained. Leave a verbose diagnostic setting at the default 90 days and it quietly compounds every month.
- Idle by design. Dev/test VMs running 24/7 for a team that works 9-to-5 are paying for roughly 120 unused hours a week, per resource.
None of this shows up as a line item called “bad architecture.” It shows up as a slightly bigger number in Cost Management that nobody has time to chase down. The only way I’ve found to actually catch it is making usage visible at the resource level, not just the invoice level.
Build one baseline instead of five spreadsheets
A cost baseline that holds up connects, in one place:
- Azure cost by workload, environment and resource owner, using consistent tags
- Reservations and Savings Plans coverage versus pay-as-you-go spend
- actual utilisation (CPU, memory, storage IOPS) per resource, not just provisioned size
- growth assumptions from your roadmap, not last year’s average
- technical and regulatory constraints on what you’re allowed to change (data residency, retention minimums, HA SLAs)
In practice that means Cost Management + Billing exports feeding a Log Analytics workspace or a Power BI model, tagged consistently enough that you can pivot by owner without a manual reconciliation step. If you want a quick way to check whether that tagging discipline is holding, this Resource Graph query is one I run a lot, it surfaces every VM missing a cost-centre tag in seconds:
Resources
| where type =~ 'microsoft.compute/virtualmachines'
| where isempty(tags['cost-center'])
| project name, resourceGroup, subscriptionId, location
Run that across your management group weekly and untagged spend stops being a mystery. Pair it with Azure Advisor’s cost recommendations (idle resources, underutilised VMs, orphaned public IPs, unattached disks) and you’ve got a baseline that keeps itself honest instead of going stale the week after you build it.
Pay-as-you-go, reservations and savings plans aren’t the same discount wearing different clothes
I still see these three treated as interchangeable, but they behave completely differently, and picking the right one depends on how stable the workload actually is.
- Pay-as-you-go is your baseline: full flexibility, resize or delete anytime, no discount. Right for anything still in flux, new workloads, migrations, short-lived environments.
- Reserved Instances (RI) commit you to a specific VM family and region for 1 or 3 years. Best discount out there, up to roughly 72% versus PAYG, but the least room to move: change the workload to a different VM series or region and the reservation just stops applying. You can resize within the same series, that’s called instance size flexibility, but you can’t switch families. I only reach for RIs on things I’m confident are genuinely stable: domain controllers, core databases, always-on systems.
- Azure Savings Plans for compute don’t lock you into a SKU at all. You commit to an hourly spend amount at the subscription or billing-account level instead, and the discount follows automatically whether the VM grows, shrinks, changes family, or the workload moves to App Service or an eligible container service. Discount’s a bit lower than an RI, roughly 65% versus 72%, but you’re not locked into one instance. This is the one I default to for anything still evolving.
Short version: an RI is tied to the VM, so scaling outside its scope loses the discount. A Savings Plan is tied to spend, so size and family stay free to change. Committing to a discount before rightsizing just locks in the waste at a slightly lower price. Rightsize first, commit second, and when you’re not sure, start with the Savings Plan.
This needs ongoing monitoring, not a one-off review
VM size isn’t a decision you make once at deployment. It’s something that should keep tracking actual load.
- Azure Monitor VM insights gives you the 30-day CPU, memory and disk curve. Never crosses 30%? The SKU’s oversized, no matter what commitment discount is sitting on top of it.
- Azure Advisor turns that same telemetry into rightsizing and shutdown recommendations automatically, so you’re not defining your own thresholds from scratch.
- Autoscale on VM Scale Sets and App Service Plans scales up and back down with load instead of you provisioning permanently for peak demand. That’s the difference between sizing for Black Friday and paying for Black Friday all year.
- Start/stop automation for dev/test VMs, via Azure Automation, Logic Apps, or the Start/Stop VMs solution, actually enforces the after-hours idle time from the list above instead of leaving it as something you meant to get around to.
Scaling up is the easy part, everyone does that the moment an app slows down. What I see teams skip is scaling back down once load drops, and that’s exactly what stays undone unless monitoring and automation are doing it for you.
Make it a cycle, not a project
A one-off cleanup buys you a quarter of savings and then quietly erodes, because new projects and new workloads keep changing the baseline underneath you. What keeps it under control long-term is pretty boring, and that’s kind of the point:
- Budgets and action groups in Cost Management, scoped per subscription or resource group, so an overrun triggers an alert on your side before it triggers a finance escalation.
- Anomaly detection, Cost Management’s built-in alerts or a scheduled query against your cost export, to catch a misconfigured autoscale rule within days instead of at month-end.
- Tag enforcement via Azure Policy, not tagging as a polite suggestion. A deny or append policy on
cost-centerandenvironmentat the management group level is what keeps that Resource Graph query above actually meaningful. - A monthly cost review with an owner per workload, looking at utilisation, commitment coverage and rightsizing recommendations together, not three separate reports nobody cross-references.
None of this is exotic. It’s the same operating discipline you already apply to uptime and patching, just pointed at the invoice instead. The part I actually like about it: the payoff isn’t just a smaller number at the bottom of the bill, it’s being able to say, with data, why a workload costs what it costs, and change that on purpose instead of by accident.
Go open your own Cost Management blade this week. I’d genuinely be surprised if you don’t find your own version of that D8s_v5.
What should change in your Microsoft environment?
In 30 minutes, we'll clarify together:
- Clarify the current situation and objective
- Structure possible solution paths
- Define the most sensible next step
Prefer to write? hello@clouddream.team
30 minutes · Microsoft Teams · No obligation



