I got a 3 AM alert once that just said: disk at 100%. No context, no warning trend, just a flat wall.
The site was a mid-size WooCommerce store I manage for a client, running on a 40GB Vultr VPS. Nothing had changed in weeks. No new plugin, no traffic spike, no deploy. I logged in half-asleep expecting a database issue or some runaway backup job.
It wasn't the database. It wasn't a backup. It was logs.
Specifically, it was almost 30GB of Nginx access logs, PHP error logs, and a Docker container that had been quietly writing debug output to a JSON file since the day it was deployed — three months earlier. Nobody had touched it because nobody had a reason to. It just sat there, growing, one request at a time, until it didn't fit anymore.
The Part Everyone Skips: Disks Don't Fill Up Loudly
Here's what makes this problem sneaky: there's no dramatic moment where you'd notice it happening.
A crashed service throws an error. A memory leak eventually kills a process. But a log file growing 2MB an hour doesn't do anything visible — until the exact second it hits 100%, and then everything breaks at once. Writes fail. The database can't checkpoint. Sometimes even SSH sessions won't open because there's no room to write a temp file.
By the time you get an alert, the log file didn't just start being a problem. It's been building for weeks, sometimes months, completely silently.
I checked the disk usage trend afterward (Vultr keeps basic historical graphs), and the climb had been steady and linear since deployment. Nobody looked, because nobody thought to. That's really the whole story of this problem — it's not complicated once you see it, it's just invisible until it isn't.
Also Read: Where Does Your File Go After You Upload It to the Cloud
Where the Logs Were Actually Hiding
Once the site was back up (a quick truncate got us breathing room fast — more on that below), I went looking for exactly what had eaten the disk. Running:
du -sh /var/log/* | sort -rhtold me most of what I needed. But cloud servers have more log sources than people usually account for, and this is the part that trips people up. On that one box, the disk was being eaten from four separate directions at once:
- Nginx access logs at
/var/log/nginx/access.log— never rotated properly because a logrotate config had silently failed after an OS update. - Docker's default logging driver (
json-file), which by default has no size limit at all unless you explicitly configure one. Everyconsole.logfrom that Node container was piling up forever. - journald, systemd's own logging system, which keeps growing until you either cap it or run
journalctl --vacuum-size. - PHP-FPM's slow log, left on from a debugging session someone had done back in July and never turned off.
None of these individually would have filled a 40GB disk. Stacked together, running unattended for months, they absolutely did.
The Docker One Deserves Its Own Warning
Out of everything on that list, the Docker logging default is the one that catches the most people off guard, so it's worth sitting on for a second.
If you run a container without specifying logging options, Docker just... writes forever. There's no automatic cap. A chatty app that logs every request, every warning, every retry attempt will produce a log file that grows for as long as the container lives — which, on a production server, can mean years.
The fix is one line in your run command or Compose file:
docker run --log-opt max-size=10m --log-opt max-file=3 your-imageOr in docker-compose.yml:
logging:
driver: "json-file"
options:
max-size: "10m"
max-file: "3"That caps it at roughly 30MB total for that container, rotated automatically. It takes thirty seconds to add and it should honestly be a default step in every deployment checklist, not something people discover the hard way.
Getting the Disk Back Without Breaking Anything
If you're in the middle of a full-disk emergency right now, don't delete log files with rm. Deleting a file that's still open by a running process doesn't free the space immediately — the process keeps writing to the deleted inode until it's restarted, and you'll be confused why df -h still shows the disk full even after the file "disappeared."
What actually works:
truncate -s 0 /var/log/nginx/access.logThis empties the file in place without breaking the file handle the running process is writing to. Same trick works for any actively-written log. It buys you space instantly, and it's safe.
Once you're not in crisis mode, that's when the real fixes go in — logrotate configs, Docker log limits, and a journald cap:
journalctl --vacuum-size=200MThis tells systemd's journal to trim itself down to 200MB and stay there going forward.
Setting It Up So You Never Get Paged for This Again
After that 3 AM alert, I put a standing routine in place across every server I manage, and it's genuinely reduced disk incidents to almost zero since:
First, confirm logrotate is actually running, not just configured. A config file existing means nothing if the cron job or systemd timer behind it silently failed:
cat /var/lib/logrotate/statusSecond, cap every Docker container's logging explicitly. No exceptions, no "I'll add it later."
Third, set a journald size limit in /etc/systemd/journald.conf (SystemMaxUse=200M) rather than leaving it uncapped.
Fourth, add basic disk monitoring that alerts at 70%, not 95%. Most providers — DigitalOcean, Vultr, Linode, AWS CloudWatch — offer this for free or close to it. The whole point is catching the slope, not the cliff.
Fifth, actually check du -sh /var/log/* every so often, even when nothing's wrong. It takes fifteen seconds and it's the single habit that would have caught this months earlier.
Mistakes I See Repeated Constantly
Assuming logrotate "just works" forever. Package updates, permission changes, and OS upgrades can quietly break a rotation config, and nothing tells you until the disk is already full.
Leaving debug-level logging on in production. It's meant to be temporary. It rarely is.
Never checking Docker's logging defaults, because nobody thinks of container logs as "real" logs the way they think of syslog or Nginx logs.
Alerting only at 100%. By the time you're notified, you're already in an outage, not a warning.
Deleting logs and moving on without asking why they grew that fast. Sometimes a bloated log is just noise. Sometimes it's a bot hammering an endpoint, a misconfigured retry loop, or a scraper that's been abusing your server for weeks — and the log was the only evidence of it.
Final Thoughts
What got me about this whole incident wasn't that logs filled the disk — that part makes sense once you know how logging defaults work. What got me was how ordinary the whole buildup looked the entire time. No errors. No warnings. Just a number quietly climbing in the background while everything on the surface looked completely fine.
Also Read: Why Cloud Servers Suddenly Become Unresponsive
That's really the lesson worth keeping: logs are one of the few things on a server that grow purely from the server doing its job correctly. Nothing has to go wrong for them to become a problem — they just need to be left unwatched long enough. A five-minute check now and then is genuinely all it takes to make sure you find out about it before your users do.


comments