Critical Infrastructure Redundancy: Deploying Automated Sat-to-Fiber Failover RoutersConceptual schematic · not to scale. Labels and connections illustrate the concepts explained in this guide.Site LANOT / IT / guestPolicy routerprobes + hysteresisFibre primaryindependent pathSatellite backupindependent pathShape traffic for degraded capacity
FIG. 14 / Conceptual engineering diagram. Not to scale. On small screens, swipe to inspect.

Key takeaways

  • Probe beyond the carrier gateway through each specific WAN.
  • Use hysteresis to prevent repeated switching.
  • Stable sessions require more than a changed default route.

Detect the failure users experience

An Ethernet port can remain up when a carrier has lost upstream connectivity. Probe more than the local gateway: use several controlled destinations and, where practical, an application-level check. Choose a decision rule that tolerates one unreachable probe target. Route the probes explicitly through the WAN being assessed so that a healthy backup cannot conceal a failed primary.

Separate hard loss from degradation. Packet loss or queueing may make a link unusable before it is completely down. Thresholds should reflect the application’s tolerance and observed normal behaviour, especially on a satellite path with more variable delay.

Add hysteresis and hold-down

A sample starting policy might probe every five seconds, fail after three consecutive failed assessments and require 60 seconds of healthy results before restoration. These are illustrative values, not universal recommendations: they create detection delay and should be tested against the outage budget. Hysteresis prevents repeated switching around a borderline threshold.

Prefer a controlled failback during stable conditions. A route that immediately returns to an unstable primary can interrupt users more often than remaining on the backup for a few minutes.

Session continuity is separate

Ordinary multi-WAN failover changes the public source address and usually breaks existing sessions. An overlay can preserve a stable endpoint if it is designed to do so, but its anchor, encryption performance and routing become additional dependencies. Validate MTU, DNS, asymmetric paths and inbound services. Do not promise seamless calls based only on a router status change.

Acceptance runbook

Test upstream blackholing as well as unplugging the modem. Confirm that essential traffic is prioritised and large backups are paused on satellite. Record time to detect, time to reconnect and time to restore full service separately. Store the router configuration offline and include a manual fallback procedure for failed automation or expired subscriptions.

From concept to acceptance

Worked design exercise

With five-second probes and a threshold of three consecutive failed rounds, failure detection is typically on the order of 10–15 seconds after an outage, depending on its position in the probe cycle, before probe timeouts and routing convergence are counted. An application may then need additional time to reconnect. A promise of sub-second continuity cannot be met by this sample policy.

Measure every phase separately and choose thresholds that balance detection speed against false positives. Repeat a partial-failure test in which the local carrier gateway answers but the application path is unreachable.

Regional compliance & specification note

Australia / ACMA. verify backup-terminal authorisation and provider usage terms before treating it as a continuity resource.

Germany / BNetzA. authorised radio equipment is one requirement; regulated organisations must also assess their sector-specific continuity and security duties.

These are procurement checks, not a determination that a particular installation is authorised. Confirm the current equipment, service, site and operating mode with the relevant authority and provider.

Sources & further reading

Official references for standards, programme context and service terms. Worked examples and design checklists are editorial analysis; verify project-specific inputs before implementation.

References reviewed 03 October 2026. Operator terms and regulatory documents may change.

Ask an engineering questionBack to top ↑