A server that's down is like a phone that's been disconnected. Nothing rings.
A server that's unresponsive is like a phone that rings all day while the person is buried under paperwork and can't pick up. The machine is powered on and the network path works. But the programs that are supposed to answer you, like SSH or your web server, can't get enough of something to do their job.
That "something" is usually memory, disk speed, CPU time, or a limit you didn't know existed. The trick is figuring out which one, and you can do that before you log in to anything.
The question I ask first: what still answers?
Before touching the dashboard, I run three quick checks from my own laptop. Replace the IP and domain with yours.
ping -c 4 203.0.113.10
nc -vz 203.0.113.10 22
curl -I --max-time 10 https://yourdomain.comThen I read the results like this:
Ping works, but SSH and the website time out. The machine is alive at the operating-system level, but the programs on it are starving. Ping is answered by the kernel itself, so it keeps replying long after everything else has slowed to a crawl. This is the classic "alive but frozen" pattern.
One warning: many providers block ping by default. AWS security groups, for example, don't allow it unless you add a rule. If ping fails on a server that normally ignores it, that tells you nothing.
The port opens, but SSH hangs before asking for a password. Run ssh -v user@203.0.113.10 and watch where it stalls. If you see "Connection established" and then nothing, the network handshake succeeded but the SSH service isn't responding. That almost always means the machine is too starved to spawn a new login process.
Nothing answers at all, not even the port. Now it's either a network or firewall problem, a provider-side event, or a fully hung machine.
Everything answers, but the site returns errors. The server is fine. That's an application problem, which is a different hunt.
This ten-second habit has saved me more time than any monitoring tool I've installed.
The usual suspects, in the order I've actually met them
1. Memory that ran out slowly, not suddenly
This was my Sunday afternoon problem. I was building a Docker image on a 1 GB VPS with a small swap file, and the server didn't crash. It got slower and slower, then stopped responding to keystrokes.
Everyone talks about the Linux OOM killer, the part that force-kills processes when memory runs out. But before it steps in, there's an ugly middle stage where the system keeps shuffling data between RAM and swap on disk. The disk becomes the bottleneck, everything waits in line, and your SSH session is just one more thing in the queue.
The tell, once you're back in, is vmstat 1. Look at the si and so columns (swap in, swap out). If those numbers keep jumping while wa (I/O wait) is high, the server was thrashing, not dead.
I wrote about what each maxed-out resource looks like in What Happens When Cloud Resources Are Fully Used, so I won't repeat all of it here. The point for this article is that memory trouble often looks like a freeze long before it looks like a crash.
2. A job that started overlapping itself
This one surprised me. A backup script ran from cron every 15 minutes and normally finished in about two. One week it started taking longer because the data had grown. Eventually a run took more than 15 minutes, and cron launched another copy anyway.
Then there were three copies, then six. Each one slowed the others down, which made the pile-up worse. Nothing was "wrong" with the script. Nobody had told it to wait for itself.
If you're at a prompt and the server is crawling, run ps aux --sort=-%cpu | head and look for the same command listed many times. The fix is one word in front of your cron command: flock -n /tmp/backup.lock your-script.sh. That makes a new run quietly skip if the last one is still going.
3. The burst credits ran out
This one is invisible unless you know to look. Many cheap cloud instances are burstable. AWS T-series instances, for example, earn CPU credits while idle and spend them under load. When the credits hit zero, the CPU is throttled to a low baseline, and the server suddenly feels like it's running through mud.
Disks can do the same thing. Small AWS gp2 volumes work on a burst balance for I/O, and once that's spent, disk speed drops sharply. The machine looks frozen because every read and write is waiting.
The tell is in the provider's graphs, not on the server. In CloudWatch, look at CPUCreditBalance for the instance or BurstBalance for gp2 volumes. If either drops to zero right when the trouble started, you've found it.
4. A noisy neighbour
On shared VPS plans, other customers share the physical hardware with you. If someone next door is hammering the CPU, your virtual machine gets less time than it expects. Inside the server, run top and check the %st value. That's steal time, the share of CPU your VM wanted but the hypervisor gave to someone else. A number that stays high for minutes isn't your fault, and no amount of tuning inside the server fixes it.
5. A limit you didn't know you had
Servers have quiet ceilings beyond CPU and RAM. Two I've bumped into:
- Open file limits. Your application logs "Too many open files" and stops accepting new connections while the rest of the machine looks healthy.
- The connection tracking table. On servers with a firewall and lots of short-lived connections, the kernel can run out of room to track them. New connections get dropped while existing ones keep working.
dmesg -T | grep -i conntrackwill show "table full" messages if this is what happened.
Both look like "the server ignores new visitors but everything else seems fine", which is a very confusing symptom the first time.
6. Something on the provider's side
Sometimes it genuinely isn't you. A host machine gets maintenance or has a hardware problem, and your VM freezes or restarts. Check the provider's status page and your account email before going down a rabbit hole. I've lost more than one hour hunting a bug that turned out to be a maintenance notice sitting unread in my inbox.
The time it was me
I'll admit this one, because I think it's more common than anyone says.
Everything went silent on a server one morning. Ping was blocked anyway, so no clues there. SSH timed out. I started checking memory graphs, disk graphs, all of it. Everything looked normal.
Then I tried the same IP from my phone's mobile data, and it worked.
My home internet provider had quietly given me a new IP address overnight, and the firewall rule on the server only allowed my old one. The server was perfectly healthy. I had locked myself out.
Now, before I look at a single graph, I test from a second network. It takes 30 seconds and rules out an entire category of embarrassment.
What to do when it's happening right now
Here's the order I follow. It's changed a lot since my early days of frantic refreshing.
- Probe from outside. Run the three checks above and note what still answers. Also test from a second network, like your phone's data.
- Look at the provider's graphs. DigitalOcean, AWS CloudWatch, Google Cloud Monitoring and Azure Monitor all show CPU, disk and network history. The shape matters. A line stuck flat at 100% means saturation. A sudden cliff means a restart or a crash.
- Open the provider's web console. Most providers have one (DigitalOcean calls it the Droplet Console, AWS has EC2 Serial Console and Instance Connect). It doesn't depend on SSH, so it often works when SSH doesn't. Even a slow, laggy login is worth having. It lets you run
uptime,topandfree -mand see what's actually eating the machine. - Wait a few minutes before rebooting. If the cause is swap thrash or a pile of overlapping jobs, the machine sometimes recovers on its own once something finishes or gets killed.
- Use the gentlest reboot first. A soft restart from the dashboard is better than a hard power cycle, because the operating system gets a chance to flush its writes to disk. A hard power-off is the last resort, especially with a database running.
- Collect evidence right after it comes back. Run
uptime, thendmesg -T | tail -50, thenjournalctl -b -1 -e, which shows the tail end of the previous boot's logs, exactly the moments before the freeze.
That last command only works if your system keeps logs across reboots. On many setups it doesn't by default. You can turn it on with:
sudo mkdir -p /var/log/journal
sudo systemctl restart systemd-journaldI did this on every server I manage after one incident where the logs from the freeze simply weren't there.
Setting things up so the next freeze explains itself
The most frustrating part of an unresponsive server is that once it reboots, the evidence disappears. A little setup beforehand changes that.
- Install
sysstatand enable its data collection (on Debian and Ubuntu that's a one-line change in/etc/default/sysstat). Afterwards,sar -rshows yesterday's memory usage andsar -ushows CPU, minute by minute. You can look back at the exact moment things went bad. - Monitor memory, not only CPU. On AWS, the default EC2 metrics don't include memory or disk usage. You need the CloudWatch agent for that, and a lot of people never set it up. Netdata is a good free option if you want real-time graphs without much configuration.
- Add a small swap file as a cushion. It won't fix a memory problem, but it gives the machine room to slow down instead of dying instantly.
- Set an outside uptime check. A tool like UptimeRobot can alert you when the site stops responding. The important detail is to check the page, not just ping. A server can answer ping all day while the site is dead.
What I'd tell past me
Don't resize the server before you know why it froze.
I did exactly that once. The server froze, I upgraded to double the RAM, and it worked fine for about two weeks. Then it froze again, because the real problem was that overlapping cron job. The extra memory just gave it more room to pile up before the same collapse.
A bigger server isn't a diagnosis. It's a delay.
The other lesson is the one I keep coming back to. When a server goes quiet, you don't need to guess. Work out what still answers, and that will usually point you to the cause.

comments