A guy from the cleaning crew backed a floor buffer into a patch panel at a client site I was managing remotely. One second the branch office was fine. The next second, half their Slack channels lit up with "is anyone else's internet down."
I wasn't in the building. All I had was a monitoring dashboard, and I watched the whole recovery happen in almost real time — like watching a body react to a cut before you even feel the pain.
That's genuinely the best way I can describe it. A network doesn't "notice" a failure and calmly reroute like a GPS recalculating. It goes through a very specific, very mechanical sequence of panic, detection, and correction — and depending on what's actually running on that network, the whole thing takes anywhere from under a second to well over a minute.
Here's what that sequence actually looked like, broken down by roughly how long each part took.
Second 0 — The Link Just Dies
The moment that cable got yanked, the physical layer noticed instantly. Both ends of that connection lost their electrical or optical signal, and the interface flipped from up to down almost immediately — this part is nearly instant because it's just hardware detecting the absence of a carrier signal, nothing "smart" involved yet.
But here's the part people don't expect: nothing downstream knows yet. The switch knows its own port died. Nobody else does. For a brief window, every other device on the network is still confidently sending traffic toward a path that no longer exists.
Seconds 1 to 3 — Somebody Has to Notice
This is where it actually gets interesting, because "somebody noticing" depends entirely on what protocol is running.
If the link was part of an EtherChannel or LACP bundle (multiple physical links treated as one logical connection), recovery is almost boring — the bundle just quietly drops down to the remaining member links and traffic keeps flowing, often in well under a second. I've watched this happen and genuinely not noticed anything from a ping test running in the background.
If it's a routed link running OSPF, the neighboring router has to figure out the neighbor is gone. Left alone, OSPF would wait for its dead timer to expire — 40 seconds by default, which is an eternity in networking terms. But most routers today also react to the interface state change directly, which is much faster, because they don't have to wait for a timeout when the physical link itself already reported the failure.
This is exactly what happened at that branch office. The primary router's OSPF neighbor relationship over that link dropped almost immediately once the interface went down, without waiting for any dead timer.
Seconds 3 to 8 — The Network Argues With Itself
Once a device knows a path is gone, it has to tell everyone else, and this is where a slow network and a fast network really separate.
OSPF flooded a Link State Advertisement to every other router in the area, essentially saying "this path no longer exists, update your maps." Every router that received it had to rerun its shortest path calculation and rebuild its routing table.
On a small network — a handful of routers, a simple topology — this is fast. On a large, messy network with dozens of routers and inconsistent area design, this step is where things drag. I've seen convergence take five seconds on a clean setup and genuinely thirty-plus seconds on a network where nobody had touched the OSPF area design in years.
Also Read: How a Router Chooses Between Multiple Routes
If Spanning Tree Protocol is involved on the switching side instead — meaning this was a switch-to-switch link rather than a routed one — the story is much worse without help. Classic STP can take up to 50 seconds to settle after a topology change, because it's built to be cautious about accidentally creating a loop. This is exactly why almost nobody should still be running plain STP in 2026 — Rapid STP exists specifically to fix this, cutting recovery down to a few seconds by changing how switches confirm a port is safe to use.
That branch office, thankfully, was on RSTP. If it hadn't been, that outage would have lasted a full minute instead of a handful of seconds, purely because of a protocol setting nobody thinks about until the day it matters.
Seconds 8 to 15 — Traffic Actually Starts Moving Again
Once the new routing or switching table is in place, traffic starts flowing over the backup path. But — and this tripped me up the first time I really paid attention to it — things don't instantly feel normal.
Every device that was mid-conversation over the old path had already sent packets into a dead end. TCP connections that didn't get acknowledgments start retransmitting, which adds a small but real delay on top of the network-level recovery. This is why a video call might freeze for a couple of seconds even after the network has technically already fixed itself — the applications are still catching up to a network state that changed underneath them.
Voice and video traffic feel this the worst, because they're the least tolerant of even a two or three second gap. A file transfer barely notices. A Zoom call absolutely does.
Seconds 15 to 40 — Everything Quietly Settles
This is the boring part, and boring is good here. ARP tables that pointed toward the dead path age out and refresh. DNS caches that had nothing to do with the failure keep working exactly as before. Monitoring systems that had queued up alerts finally catch up and confirm the new path is stable.
At the branch office, full normal latency and packet loss numbers on the dashboard didn't return to baseline until close to the 25-second mark, even though the actual routing decision had been made in the first ten.
The Mistake I Made Watching This Happen
The first time I watched a failover like this in a live environment, I panicked at the wrong moment. I saw the OSPF neighbor drop, saw a small spike in packet loss, and immediately assumed the backup path had failed too — because the graph didn't snap back to normal instantly.
It hadn't failed. It was just retransmissions and ARP cleanup finishing their job after the actual reroute had already succeeded. I nearly called the ISP over something that resolved itself twenty seconds later on its own.
The lesson stuck with me: the moment a routing table updates is not the moment a network "feels" recovered. There's always a tail of cleanup after the real fix, and mistaking that tail for a second failure is one of the most common false alarms in network troubleshooting.
Watching Recovery Happen on Your Own Network
You don't need enterprise monitoring to see most of this yourself:
ping -t(Windows) or a continuouspingon Linux run during a planned failover shows you the exact packet loss window and how long it lastsshow ip ospf neighboron Cisco gear tells you instantly whether a neighbor relationship has reformed after a failureshow spanning-treeshows you which ports are blocking, forwarding, or still in a transitional state after a topology change- mtr is genuinely the most useful tool for this specific scenario, because it keeps sampling continuously and shows you exactly when loss starts and exactly when it stops, hop by hop
If you want to actually practice this instead of just reading about it, pulling a cable on a lab setup in GNS3 or EVE-NG while running continuous pings is one of the more oddly satisfying things you can do — watching the failure window shrink as you switch from plain STP to RSTP, or from static routing to OSPF, makes the theory click in a way no diagram really can.
Why This Actually Matters Day to Day
Most people only ever experience a link failure as "the internet went out for a bit." What's genuinely happening underneath is a layered sequence — physical detection, protocol notification, table recalculation, application-level catch-up — and each layer has its own speed, its own failure mode, and its own fix.
Also Read: Why Your WiFi Works but Websites Do Not
Understanding which stage you're actually watching, when something goes wrong, is the difference between fixing the real problem in ten seconds and chasing a phantom issue for twenty minutes — which, if I'm honest, used to be me on a fairly regular basis before I started paying attention to the timestamps instead of just the symptom.

comments