A while back, I was running a vulnerability scan for a client on a cheap $10-a-month VPS he used as a jump box. Nothing crazy, just Nmap and a couple of scripts against a range of his own internal test machines, logging everything to a file so I could review it later.
Halfway through the scan, his actual production website, hosted on that same droplet, went down.
My first thought was that I'd somehow broken something with the scan itself. I hadn't. What actually happened was dumber and far more common than that: the log file from the scan had quietly grown to fill the entire disk; the disk hit 100%, and the moment that happened, the database couldn't write anymore. The site didn't crash from an attack or a bug. It crashed because there was, quite literally, no more room left on the machine.
That was the moment "the cloud has unlimited resources" stopped being something I vaguely believed and became something I actually respected.
Also Read: What Happens When a Website Server Goes Down
The Cloud Feels Infinite Until It Isn't
Cloud marketing does a really good job of making everything sound elastic and endless. Spin up a server, scale whenever, pay for what you use. All true, mostly.
But every resource behind that promise- CPU, RAM, disk space, network bandwidth, even your account's own permission to spin up more servers- has a hard number attached to it somewhere. Someone set that number, whether it's a physical limit on the hardware or an artificial quota your provider put on your account to protect their own infrastructure.
You don't really feel that ceiling exists until you hit it. And when you do, the way each resource fails is completely different, which is exactly why it confuses people the first time.
What Actually Happens When Each Resource Maxes Out
I've now watched all four of these fail in real client environments, so here's what each one actually looks like in practice, not in theory.
- When CPU maxes out, things don't crash right away. They just get slow. Painfully slow. Every request queues up behind the last one, response times climb, and if it's bad enough, health checks start failing, and your load balancer might pull the server out of rotation entirely, thinking it's dead.
- When RAM maxes out, this is the one that actually causes hard crashes, not just slowdowns. On Linux, once memory is exhausted, something called the OOM killer (Out of Memory killer) steps in and starts forcibly killing processes to free up space, usually whichever one is using the most memory, which is often your actual application. I had this happen inside a Docker container running a Node app for a client. The container kept restarting itself, and
docker logsshowed nothing useful until I checkeddmesgand saw the kill message sitting right there in plain text. - When disk maxes out, this is what happened to my client's droplet. Once a disk hits 100%, a lot of software doesn't fail gracefully. Databases can't write new rows, log files stop mid-line, and some applications throw completely unrelated-looking errors because nobody ever coded them to expect "there is zero space left."
- When network bandwidth maxes out, usually on a smaller VPS with a capped monthly transfer allowance, requests either start timing out, get throttled down to a crawl, or you get hit with an overage bill you didn't budget for, depending on the provider.
- When your account's own quota maxes out, this is the one almost nobody thinks about until it happens to them. I once had a client try to launch a new EC2 instance on AWS during a genuine traffic emergency, and got hit with an error along the lines of "vCPU limit exceeded." Not a hardware problem. Not a billing problem. AWS simply caps how many virtual CPUs your account is allowed to run at once by default, and he'd never needed to raise that limit before, so he'd never known it existed.
Also Read
If you want the flip side of this, where things actually go right during a spike, this connects well: How Cloud Servers Scale When Traffic Suddenly Explodes
A Quick Reference: What Full Actually Looks Like
| Resource | What Happens When It's Full | First Thing To Check |
|---|---|---|
| CPU | Everything slows down, requests queue, health checks may fail | top or htop |
| RAM | OOM killer force-kills processes, apps restart randomly | free -m |
| Disk | Writes fail, logs cut off mid-line, databases throw odd errors | df -h |
| Bandwidth | Requests time out, traffic throttled, or surprise overage charges | Provider's usage dashboard |
| Account quota | New resources refuse to launch, even if you're willing to pay | Provider's service quota/limits page |
Step-by-Step: Checking If You're Actually Close to a Limit Right Now
You don't need anything fancy installed to get a rough picture of where you stand. Here's what I run on almost every server I touch, even ones that seem perfectly healthy:
- Check disk space with
df -h. Anything above 80% used is worth paying attention to, not just 100%. - Check memory with
free -m. Look at the "available" column specifically, not just "free," since Linux caches aggressively and that can look scarier than it actually is. - Check live CPU and RAM usage with
toporhtop, and watch it for a minute rather than glancing once. - Check your log directory size directly with
du -sh /var/log/*, since runaway logging is the single most common silent disk killer I've personally seen. - Check your provider's quota or limits page. On AWS this is under Service Quotas, on Google Cloud it's under IAM & Admin > Quotas, and on Azure it's under Subscriptions > Usage + quotas. Most people have never opened this page once.
If any of these are sitting close to their ceiling and you're not actively monitoring them, that's the gap that eventually turns into a 2 AM phone call.
The Mistake I Made Myself
Here's the part that's actually embarrassing to admit. I've told clients for years to watch their disk usage, and then I went and filled my own client's disk with a log file I forgot I'd left running verbose.
The lesson that actually stuck wasn't "watch your disk," because I already knew that. It was that monitoring only counts if it's actually running, not something you meant to set up eventually. I now set a basic disk alert on every single server I touch, personal projects included, in the first ten minutes of setting it up, before I install anything else. Cloud providers make this easy; DigitalOcean, AWS CloudWatch, and Google Cloud Monitoring all let you set a threshold alert for free or close to it, and none of them take more than a few minutes to configure.
A Few Mistakes Worth Avoiding
- Letting logs grow unchecked, especially debug-level or verbose logging left on by accident during testing.
- Assuming "the cloud" means infinite capacity, when really every account has quotas sitting quietly in the background.
- Never checking provider quota pages until you're already blocked mid-emergency, when raising a limit can take anywhere from minutes to a couple of business days depending on the provider and the size of the request.
- Running production and testing tools on the exact same server, which is exactly what caused my own incident.
- Ignoring memory limits inside containers, since Docker and Kubernetes both let you cap memory per container, and hitting that cap silently kills the process instead of warning you first.
None of these are rare, exotic problems. They're the small, boring oversights that quietly sit there until the one busy day they decide to matter.
Final Thoughts
"Full" doesn't look the same twice. Sometimes it's a slow website. Sometimes it's a container quietly restarting itself all night. Sometimes it's a perfectly good server that simply can't accept one more login because the disk behind it has nothing left to give.
The fix, almost every time, isn't some clever piece of engineering. It's knowing where your actual ceilings are before you're the one finding out about them live, in front of a client, at the worst possible time. A five-minute check now saves you the exact kind of night I had with that scan log.


comments