Why a VM Can Freeze Without Crashing


I hit Reset. About ten minutes later, I realized it was the dumbest thing I could have done.

It was a Friday evening, and I had an Ubuntu VM in my home lab running a WordPress staging site and MariaDB for a client's test. The console froze. SSH timed out. I gave it maybe ninety seconds, got impatient, and clicked Reset.

The VM came back, but MariaDB spent a while recovering from the crash, and the last hour of the client's test orders was gone.

Later I opened the host's graphs and saw the truth. The VM hadn't been dead. It had been busy, swapping hard, and it would probably have recovered by itself in a few minutes.

That evening changed how I treat frozen VMs. Here's what I learned.


A Crash Is a Decision. A Freeze Is a Wait.

This is the idea that made everything click for me.

When a VM crashes, the guest operating system notices something fatal and stops on purpose. That's a blue screen on Windows or a kernel panic on Linux. There's a message, a log entry, and a stop code.

When a VM freezes, nothing notices anything. The guest is just waiting for something: a disk to answer, memory to free up, a CPU core to become available. Waiting isn't an error, so nothing gets logged.

That's why frozen VMs feel so mysterious. There's no confession in the logs because, from the guest's point of view, nothing went wrong. It's just very, very patient.

Linux even makes this explicit. By default, the kernel will log a warning that a task has been stuck for a long time and then carry on. It doesn't panic unless you've configured it to. Windows behaves similarly. So a VM can sit in a stuck state for minutes and never produce a crash.


Freeze #1: The One That Was Only a Screen

A few months before the Reset incident, I had a Kali VM in VirtualBox that froze solid. The mouse was dead, the desktop didn't move, and the window looked like a photograph.

I was ready to power it off, but I tried one small thing first: I opened a terminal on my host and ran ssh into the VM.

It answered instantly.

uptime showed six days of runtime and a normal load. The guest was perfectly fine. Only the graphical display had hung.

I restarted the display manager from that SSH session:

sudo systemctl restart display-manager

The desktop came back. (This closes your graphical session, so unsaved GUI work is lost. But everything running in the background survives.)

The root cause was two things stacked together. I had 3D acceleration switched on for the VM, and my Guest Additions were older than the VirtualBox version on the host after an update. Turning off 3D acceleration and updating Guest Additions ended the problem.

The lesson: a frozen window is not the same as a frozen VM. Before you do anything drastic, try a second way in.


Freeze #2: The One Waiting on a Disk That Wasn't Answering

This one was on a client's small office server. It was a Proxmox box running a few VMs, with the VM disks stored on a NAS over the network.

Every so often, one VM would freeze for several minutes, then wake up like nothing had happened. Users said the file share "hung." No crash, no restart.

After recovery, I ran this inside the Linux guest:

dmesg -T | grep -Ei "blocked for more than|soft lockup"

It showed a pile of lines saying tasks had been blocked for more than 120 seconds. That's the kernel's way of saying "I've been waiting on the disk for two minutes and got nothing back."

The cause turned out to be the NAS. It ran a scheduled disk check in the early hours, right when the backup job also started. Storage response times went through the roof, and every VM waiting on a disk read or write simply stopped.

The VM wasn't broken. Its disk had gone slow, and everything else waited politely behind it.

We moved the disk check and the backup to different windows, and shifted that one VM's disk onto local SSD storage. The freezes stopped.

On Windows guests, the equivalent clue is usually Event ID 129 in the System log, a warning that the storage device didn't respond and a reset was issued. If you see that around the time of a freeze, look at your storage first.


Freeze #3: The One That Was Stuck in Line

This one wasn't a real freeze, and that's what makes it sneaky.

In my lab, I gave every VM four virtual CPUs because more felt better. The host had four physical cores and four VMs. When two of them ran heavy scans at the same time, everything stuttered. Keystrokes lagged and windows froze for a few seconds at a time.

Inside the guest, top showed a high number in the st column. That's steal time, meaning the guest wanted to run but the host hadn't given it a turn yet. On VMware, the equivalent is CPU Ready in the performance charts. If it sits above roughly 5–10%, the VM is spending real time waiting in line.

The fix felt backwards. I cut each VM down to two vCPUs, and everything got faster. Fewer virtual CPUs are easier for the host to schedule, so each VM waits less.

If a VM feels frozen in short bursts rather than for long stretches, suspect this one.


Busy Freeze or Dead Freeze? Ask the Host.

After the Reset mistake, I started sorting every frozen VM into one of two groups:

  • A busy freeze: the VM is stuck, but something is still happening. It might be swapping, waiting on a slow disk, or finishing a snapshot job.
  • A dead freeze: nothing is happening at all. The guest is truly stuck, in a deadlock, a spinning kernel loop, or something similar.

You can't tell which one it is from inside the VM, because the VM isn't talking. You tell from the host.

Here's what I look at:

  • Disk activity for that VM. On Proxmox, the node summary shows IO delay. On ESXi, check datastore latency. On a Windows host, open Task Manager or Resource Monitor. If disk activity is high, the VM is probably working through something.
  • Host memory. If the host is swapping, every VM on it will feel it.
  • CPU behavior. If the VM's process is pinned at exactly one core at 100% with no disk traffic, the guest is often spinning in a loop. That's more likely a dead freeze.
  • Free space on the datastore. A full datastore can pause VMs. Hyper-V shows Paused-Critical, and ESXi pops up a question asking you to free space and retry.
  • Snapshots. Deleting a big snapshot can stun a VMware VM while it consolidates, and with a huge delta file that pause can last far longer than "brief." I wrote about why in How VM Snapshots Work.

If everything is busy, wait. If everything is flat and silent, you're probably dealing with the real thing.


What I Do Now, in Order

This is the routine I follow every time before I go near Reset:

  1. Stop mashing keys. Every keypress just queues up. I take a screenshot of the console and note the time.
  2. Try a different door. Ping the VM, try SSH or RDP, and check the hypervisor's own summary page. If the guest agent or VMware Tools status is still reporting an IP address, the guest is at least partly alive. (Test ping when the VM is healthy first, because many Windows setups block it by default.)
  3. Check the host. Look at disk latency, memory, CPU, free space, and whether the VM is paused.
  4. Give a busy freeze real time. If the host shows heavy disk activity, I wait fifteen to twenty minutes. That's my personal rule, not a law, but it would have saved my client's database.
  5. Ask nicely first. Use the hypervisor's shutdown guest option, which sends a normal power-button signal. On a frozen VM it often won't work, but it costs nothing to try. If your keyboard combo for Ctrl+Alt+Del goes to the host, use the "send Ctrl+Alt+Del" option in the hypervisor menu instead.
  6. Power off before Reset. Both are rough, but I treat Reset as the very last option.
  7. After forcing it, check for damage. Let the filesystem check run, then look at your database logs. Any data that wasn't written to disk is gone.

When It Really Is Dead

Sometimes it is a dead freeze. In that case, the last thing worth doing before you force it off is grab evidence, because a restart wipes the state that would explain what happened.

At minimum, take that console screenshot and save the host graphs. If you want to go further, most hypervisors can send an NMI (non-maskable interrupt) to the guest. On Hyper-V that's Debug-VM with -InjectNonMaskableInterrupt, and on KVM it's virsh inject-nmi. The catch is that the guest has to be set up in advance to produce a crash dump when it receives one. So try this in a lab first, not for the first time during an emergency.


Mistakes I Keep Seeing (Including Mine)

  • Resetting too early, which is what I did, and it cost a client an hour of data.
  • Assuming a freeze means faulty hardware. In my experience, it's far more often storage, memory, or CPU scheduling.
  • Adding more vCPUs or RAM to "fix" it. More vCPUs often make scheduling worse, and more RAM on an overcommitted host makes it worse still.
  • Never learning what normal looks like. If you've never seen your host's usual disk latency, you can't spot a bad one.
  • Leaving old snapshots lying around. They slow down disk access and can turn a small hiccup into a long stall.
  • Running real-time antivirus over VM disk files on a Windows host. Scanning huge .vmdk, .vhdx, or .vdi files can make a VM crawl. Most hypervisor vendors recommend adding exclusions.

The Note Taped to My Monitor

After the Friday incident, I wrote one line on a sticky note and stuck it to my monitor:

"Is it dead, or is it busy? Check the host first."

Every frozen VM since then has been a little less scary. Sometimes it was a display bug, sometimes a slow disk, and sometimes a crowded CPU. Twice it really was dead. But I only know that now because I looked before I clicked.

Next time a VM stops responding, take a breath and open the host's graphs. The answer is usually there.


Also Read: Why VMs Suddenly Restart

Hashir
Author At TopicGems • Published Tuesday, September 22, 2026
Hashir is a freelance cybersecurity professional and web developer, working with clients since 2022. He writes about virtualization, networking, and cloud infrastructure based on hands-on client work.

comments