The most embarrassing cloud bill I've ever looked at was my own.
It was a small AWS account I use for testing, a throwaway sandbox where I break things on purpose. Nothing in it mattered. I assumed it cost me about the price of a couple of coffees a month.
Then one month the bill came in at roughly $55, and almost all of it sat under one line: EBS snapshot storage.
That made no sense. The biggest disk I had ever attached to a test server in that account was 100 GB. Most were 30. I had terminated nearly all those servers weeks earlier. Yet AWS was billing me for over a terabyte of snapshots.
I spent an evening figuring out where all that data came from. The answer turned out to be boring, and that's exactly why it's worth writing down.
Reading the Bill Backwards
I started in Cost Explorer and grouped the costs by usage type. One entry stood out: the one containing EBS:SnapshotUsage (AWS adds a region prefix, so yours may look slightly different).
So the money wasn't going to servers or running disks. It was going to copies of disks that no longer existed.
Next I opened EC2 → Snapshots and found 47 of them. I tried adding up the Size column to match the bill, and it didn't match. The total was much bigger than what I was being charged for.
That's when I learned the first thing nobody tells you: the Size column shows the size of the original volume, not how much the snapshot actually stores. You're billed on stored data, so the number on your screen and the number on your invoice measure two different things.
What a Snapshot Actually Holds
The closest comparison I can give is Git history, roughly.
The first snapshot of a volume copies every block that's in use. Every snapshot after that stores only the blocks that changed since the one before it. That's what "incremental" means, and it's why snapshots are cheap when a disk sits mostly idle.
It also explains the second thing that confused me. I deleted six old snapshots and the bill barely moved.
When you delete a snapshot, AWS only frees the blocks that nothing else depends on. If a newer snapshot still needs a block from an older one, that block stays. Deleting the oldest snapshot in a chain often frees far less than you'd expect, because its data lives on inside the later ones.
The other factor is churn, meaning how much of the disk gets rewritten. An idle volume produces tiny snapshots. A volume running a database, piling up logs, pulling Docker images or applying system updates rewrites blocks all day. Two disks of identical size can cost completely different amounts to snapshot.
Also Read: How Cloud Logs Can Quietly Fill Your Server Disk
Where the Space Was Hiding
Once I understood the mechanics, the 47 snapshots sorted themselves into a few groups. These are the usual suspects, and I've seen the same ones in other people's accounts since:
| Where it hides | Why it stays | How to spot it |
|---|---|---|
| Orphaned snapshots | Deleting a server or volume never touches its snapshots | The source volume ID no longer exists |
| Snapshots behind old AMIs | Deregistering an AMI leaves its backing snapshots behind | Description starts with "Created by CreateImage" |
| Schedules with no expiry | Something creates snapshots on a timer but nothing ever deletes them | Dozens of near-identical snapshots at regular intervals |
| "Just in case" snapshots | Taken before a risky change, then forgotten | Empty or vague descriptions, dates clustered around one event |
| Copies in other regions | The console shows one region at a time | The bill lists a region you never open |
Most of my terabyte was the first two rows. I had built a few custom images while learning, deleted the servers, deregistered the images, and assumed that cleaned everything up. It didn't.
The Home-Lab Version of the Same Mistake
This isn't only a public-cloud problem. The same thing happened to me on Proxmox in my home lab.
Before a big update, I took a snapshot of a VM that runs a small database and named it before-upgrade. The update went fine and I forgot about the snapshot completely.
About three weeks later I noticed the LVM-thin pool was sitting at 91% full. Nothing was wrong with any single VM. The snapshot had been quietly holding onto every block the database rewrote since that day.
On thin-provisioned storage this matters more than it does in the cloud. If the pool hits 100%, the guests using it can freeze or throw disk errors. Nobody sends you a bill first. The disk just stops accepting writes.
The commands I now run whenever I suspect this:
qm listsnapshot 101
lvs
zfs list -t snapshot -o name,used,creationThe first lists the snapshots on a VM (101 is my VM ID). lvs shows the thin pool's Data% column, and the ZFS command is for setups running on ZFS. Once I confirmed the snapshot was no longer needed, qm delsnapshot 101 before-upgrade cleaned it up.
Also Read: Why Moving a VM Can Change Its Network Behavior
How I Cleaned Up My AWS Account
Here's the process I followed. I'd suggest doing it in this order, because the earlier steps keep you from deleting something you'll regret.
1. Find out what's costing money.
Cost Explorer → group by Usage type → look for the snapshot line. If it's a big share of your bill, keep going.
2. List your snapshots, and check every region.
This is the step I got wrong the first time. Snapshots are regional, and the console only shows the region you're currently viewing.
aws ec2 describe-snapshots --owner-ids self \
--query 'Snapshots[].[SnapshotId,VolumeId,VolumeSize,StartTime,Description]' \
--output tableTo count snapshots across all regions at once:
for r in $(aws ec2 describe-regions --query 'Regions[].RegionName' --output text); do
echo -n "$r: "
aws ec2 describe-snapshots --owner-ids self --region $r --query 'length(Snapshots)'
done3. Sort them into piles.
I used three: snapshots whose source volume is gone, snapshots older than I could realistically ever roll back to, and snapshots backing an image. To check the first and last:
aws ec2 describe-volumes --query 'Volumes[].VolumeId' --output text
aws ec2 describe-images --owners self \
--query 'Images[].BlockDeviceMappings[].Ebs.SnapshotId' --output textIf a snapshot's volume ID isn't in the first list, it's an orphan. If its ID shows up in the second list, an image still depends on it. AWS won't let you delete a snapshot that's in use by a registered image anyway, but it's better to know why.
4. Tag before you delete.
I didn't delete anything immediately. I tagged the candidates with a review date and waited two weeks:
aws ec2 create-tags --resources snap-0abc123 --tags Key=review-after,Value=2026-10-15If you'd like a safety net, AWS also has a Recycle Bin feature for snapshots that lets you recover accidental deletions within a retention window you set.
5. Delete, then be patient with the numbers.
Remember the shared-blocks behavior from earlier. The bill won't drop evenly with each delete. Some deletions free a lot, others almost nothing, and the real drop shows up once whole chains are gone. Mine went from about 1.1 TB of billed snapshot storage down to roughly 180 GB.
6. Automate the retention so it doesn't come back.
In the EC2 console, look for Lifecycle Manager (Data Lifecycle Manager). You create a policy that snapshots tagged volumes on a schedule and automatically deletes anything beyond the number you choose to keep. AWS Backup can do the same job if you already use it. Either way, the point is that something is responsible for deleting, not just creating.
Four Things I Believed That Were Wrong
- "Terminating the server cleans everything up." It removes the server and, depending on settings, its disk. Snapshots are separate objects and stay until someone deletes them.
- "The Size column tells me what I'm paying for." It shows the source volume size. Your bill is based on stored data, so the two rarely agree.
- "Deregistering an image frees its storage." It removes the image entry. The snapshots that backed it are left behind.
- "If I can't see them, they aren't there." I checked one region and declared the account clean. Snapshots can exist in any region you've ever touched.
A Few Questions I Get Asked
Are snapshots the same as backups?
Not on their own. By default they live in the same account and region as the thing they protect, so an account compromise or a careless "delete everything" script can take out both. For anything that matters, I want a copy in another region or account too.
How many should I keep?
There's no universal answer. As a starting point, I'd use something like 7 daily and 4 weekly, then adjust based on how far back you'd realistically need to roll. If you truly need to keep old ones for compliance or long-term safety, look at the snapshot archive tier. As of my last check it costs roughly a quarter of the standard price with a 90-day minimum, but confirm current pricing for your region before relying on that.
Will deleting old snapshots break my running servers?
A running instance doesn't depend on its own past snapshots, so deleting them doesn't affect it. That's why the step-4 waiting period matters. The real risk is deleting the one you'd have needed to roll back to.
Ten Minutes You Won't Regret
I still think about how ordinary the whole thing looked. No alert, no error, no outage. Just a storage line growing a few cents at a time on a bill I wasn't reading closely.
If you take one thing from this, open the snapshot list in whichever cloud or hypervisor you use and sort it by date, oldest first. Whatever is sitting at the top is usually where the surprise is. Look up its source, check whether anything still needs it, and decide on purpose instead of by default.


comments