ADAM PULSE Knowledge Base
Resiliency · Failover · Dual WAN · Monitoring

Primary Internet Down, Backup Internet Up: How to Know Whether Failover Actually Worked

A business can have two internet connections and still suffer a real outage. That happens because backup availability is not the same thing as successful failover. A secondary circuit can show: Up while users still experience: No internet access Broken VPN tunnels Failed cloud applications Poor voice quality Slow performance DNS problems Authentication failures Routing issues Blocked services The right question is not: “Is WAN2 up?” It is: “Did the business actually continue operating when WAN1 failed?”

Short answer

How Do You Know Whether Internet Failover Worked?

A successful failover should prove all of the following:

If any of those steps fail, you may have partial failover, not successful business continuity.

What Is Internet Failover?

Internet failover is the process of automatically moving network traffic from a failed primary internet connection to a secondary connection. Common backup technologies include: Cable broadband Secondary fiber 5G LTE Starlink Fixed wireless Another carrier SD WAN paths The purpose of failover is not simply to keep a WAN interface green. The purpose is to maintain the applications and communications the business depends on.

What Is the Difference Between Redundancy and Resilience?

These are related but different. Redundancy You have multiple components. Example: Two internet circuits. Resilience The system continues operating when one component fails. Example: The primary fiber fails, traffic shifts to the secondary circuit, VPNs recover, phones continue working, cloud applications remain usable, and users barely notice. That is resilience. A company can therefore have redundancy without resilience.

Why Backup Internet Often Fails When You Need It

Common reasons include:

Backup circuit was already degraded Failover policy was misconfigured Routing did not update properly VPN tunnel did not reestablish DNS stopped resolving Backup bandwidth was insufficient Voice traffic was not prioritized Applications rejected the new public IP Firewall rules differed by WAN Cloud services depended on source IP allowlists SD WAN health checks were poorly configured The backup shared the same physical path as the primary The backup carrier had the same upstream dependency The failover timer was too slow The firewall itself was the single point of failure This is why multi-WAN monitoring needs to answer whether failover actually preserved business operation, not merely whether WAN2 was reachable.

Step 1: Verify the Primary Actually Failed

Before evaluating failover, confirm the failure event itself. Check: WAN1 interface state Gateway reachability Packet loss Latency Routing Public IP reachability Carrier handoff Firewall logs Health check status Exact timestamp Example: 10:42:17 AM — WAN1 packet loss exceeds threshold 10:42:21 AM — ISP gateway unreachable 10:42:25 AM — WAN1 marked failed That gives you a clean starting point.

Step 2: Verify the Firewall Detected the Failure

A backup circuit cannot take over if the firewall or SD WAN controller does not recognize the primary path as unhealthy. Check: Health monitor state Link monitor SLA monitor SD WAN member state Dead gateway detection Routing table Policy state Failover logs Common mistake: The primary circuit is unusable for applications but still responds enough that the firewall considers it healthy. Then traffic never moves. This is sometimes called a brownout rather than a hard outage.

What Is a Brownout?

A brownout occurs when an internet connection is technically up but performs so poorly that applications are unusable. Examples: 40% packet loss Extreme latency High jitter Severe congestion Intermittent routing The interface remains: UP but the business experiences: DOWN A good failover design must detect service quality, not only physical link state.

Step 3: Confirm Traffic Actually Moved to the Backup

After WAN1 fails, verify that production traffic is using WAN2. Check: Active routes Firewall sessions SD WAN member NAT source address Public IP Traffic counters Application flow Do not assume: WAN2 active = user traffic moved. Verify it.

Step 4: Test Basic Internet Reachability

Once traffic moves, test: Public IP reachability Web browsing HTTPS External DNS Common internet destinations If users can reach multiple independent external services, basic failover is functioning. But that is only the beginning.

Step 5: Verify DNS Still Works

Failover can break DNS in several ways. Possible causes: Primary ISP DNS becomes unreachable Firewall continues forwarding to unavailable resolver Internal DNS depends on primary WAN VPN based DNS fails DNS security service allows only primary public IP Routing changes incorrectly Test: Public hostname resolution Internal hostname resolution Configured DNS servers Alternative resolvers where appropriate A backup circuit that passes IP traffic but breaks DNS is not a successful failover.

Step 6: Verify VPN Tunnels Reestablished

This is a major source of partial failover. The internet may work, but corporate connectivity may not. Check: Site to site VPNs Remote access VPNs Azure tunnels AWS tunnels Cloud firewall tunnels Partner tunnels SD WAN overlays Common problems: Tunnel peer expects primary public IP Backup IP is not configured Routing does not move IKE policies differ Firewall policy does not permit backup path NAT traversal changes The result: Internet works. Business network does not.

Step 7: Verify Cloud Applications

Test the actual services employees use. Examples: Microsoft 365 Google Workspace Zoom Teams Salesforce ERP CRM Hosted desktops Contact center Azure AWS Payment processing Cloud file systems Do not stop at: google.com loads That does not prove business continuity.

Step 8: Check Source IP Restrictions

Many enterprise services restrict access based on known public IP addresses. When failover occurs, the source IP may change. That can break: Cloud applications Partner portals Banking systems Security platforms SFTP APIs Hosted databases Remote management Allowlisted SaaS Ask:

Does any critical application expect traffic from the primary WAN public IP?

If yes, the backup IP must be accounted for.

Step 9: Test Voice and Video

A backup connection may keep web browsing alive while real time communications become unusable. Test: Zoom calls Teams calls VoIP Contact center SIP Video meetings Monitor: Latency Jitter Packet loss MOS where available A 50 Mbps backup may be enough for email but not: 100 phones Video meetings Cloud desktops File synchronization at the same time.

Step 10: Verify Backup Bandwidth Is Sufficient

Ask:

What is the backup designed to support?

Example: Primary: 1 Gbps fiber Backup: 100 Mbps broadband When failover occurs, the organization may suddenly have only 10% of normal capacity. That may be acceptable if traffic is prioritized. It may be disastrous if everything continues normally. Check: Backup utilization Top applications QoS Traffic shaping Guest WiFi Backups Cloud sync Video Large downloads Noncritical traffic

Should You Limit Traffic During Failover?

Often, yes. A useful failover policy may temporarily deprioritize or block: Guest WiFi Large backups Software updates Cloud synchronization Nonessential video Bulk transfers while preserving: Voice Critical SaaS VPN Payment systems Business applications This is business aware failover.

Step 11: Verify Firewall Policies on the Backup Path

A secondary WAN can fail because firewall rules do not apply correctly. Check: Outbound policies Inbound policies NAT VIPs Port forwards SD WAN rules Security profiles DNS policies Web filtering Application control Route policies The backup path should be tested before an outage, not discovered during one.

Step 12: Verify Monitoring Still Works

This one is often overlooked. When failover occurs:

Can your monitoring platform still reach the site?

Do remote agents still report?

Does SNMP monitoring continue?

Does external monitoring see the new public IP?

Do dashboards show the actual current path?

A monitoring platform that goes blind during failover can misclassify a successful recovery as a total site outage.

Step 13: Verify Inbound Services

Outbound internet may work while inbound services fail. Examples: Hosted VPN Remote access VoIP trunks Web servers Remote support SFTP Security cameras Remote desktops Inbound NAT If these depend on the primary public IP, they may remain unavailable during failover. This may be acceptable if documented. But it should not be mistaken for full continuity.

Step 14: Check Authentication and Identity

Some cloud services may behave differently after a public IP change. Check: Microsoft Entra conditional access Identity provider policies Geo restrictions Risk based login Security controls VPN authentication Partner allowlists A user might suddenly receive: Blocked sign in MFA challenge Suspicious login alert because traffic now originates from a new public IP or carrier.

Step 15: Test Critical Business Workflows

This is where technical failover becomes business continuity. Do not ask only:

Can users browse?

Test the workflows that generate revenue or keep operations running. Examples: Process a credit card Place a phone call Access CRM Open ERP Connect to hosted desktop Reach warehouse application Send/receive email Use contact center Connect to VPN Print shipping label Use cloud POS Failover is successful only if the required business workflows remain functional.

Step 16: Determine Business Impact

Classify failover outcomes. Successful Failover Users remain operational with little or no impact. Successful With Degradation Business continues but: Bandwidth reduced Voice quality degraded Some applications slower Noncritical services disabled Partial Failover Some services recover while others remain unavailable. Failed Failover Backup exists but users remain substantially offline. This classification is more useful than: WAN2 = UP.

Step 17: Measure Failover Time

Track: Primary failure detected Backup selected Routing changed VPN restored Applications healthy Business confirmed operational Example: 10:42:17 — Primary degradation begins 10:42:25 — WAN1 declared failed 10:42:29 — WAN2 selected 10:42:37 — Internet restored 10:42:51 — VPN restored 10:43:08 — Critical applications healthy Total business interruption: 51 seconds That is a meaningful resilience metric.

What Is RTO for Internet Failover?

Recovery Time Objective, or RTO, is the target time within which service should be restored after disruption. For internet failover, you can define something practical like: Critical internet services should recover within 60 seconds of primary circuit failure. Different businesses require different targets.

Step 18: Test Failback

When the primary returns:

Should traffic immediately move back?

Maybe not. A circuit can flap. Example: WAN1 returns. Traffic moves back. WAN1 fails again 30 seconds later. Traffic moves to WAN2. Repeat. That creates instability. Use: Hold down timers Stability timers Performance thresholds Failback delay Manual confirmation where appropriate A good policy may require the primary to remain healthy for several minutes before restoring it as preferred.

What Is Failback?

Failback is the process of returning traffic from the backup connection to the preferred primary connection after the primary recovers. Failback should be tested just as carefully as failover. Poor failback can create: Session drops VPN resets Voice interruptions Routing instability Repeated path changes

Step 19: Watch for Flapping

A flapping primary circuit may repeatedly trigger: WAN1 WAN2 WAN1 WAN2 That can be worse than staying on the backup circuit. Monitor: Number of transitions Duration between transitions Packet loss before transition Recovery stability Repeated path transitions should become one ongoing incident, not dozens of new alerts.

Step 20: Verify the Backup Is Truly Diverse

Two providers do not automatically mean two independent paths. They may share: Conduit Building entrance Local fiber Pole line Carrier hotel Last mile provider Transport provider Central office Utility power A backhoe can therefore cut both. Ask both carriers about: Physical path diversity Building entrance diversity Last mile ownership Central office Upstream network Power independence True resilience sometimes requires: Different carrier Different medium Different path Example: Fiber + 5G or: Fiber + Starlink may provide more physical diversity than two circuits delivered through the same underground conduit.

Step 21: Check Backup Health Before the Primary Fails

Do not wait for an outage to discover: Backup circuit has 30% packet loss Public IP changed VPN configuration is stale Billing problem suspended the line SIM expired Starlink terminal offline Router lost configuration Bandwidth was downgraded A useful monitoring platform should continuously answer: “Will my backup actually work if the primary fails right now?”

What Is Failover Readiness?

Failover readiness is an assessment of whether backup connectivity is currently capable of supporting the intended business services. It should consider: Backup reachability Latency Packet loss Bandwidth DNS VPN readiness Routing Application reachability Public IP requirements Firewall rules Physical diversity Recent test Rather than: Backup WAN: Up a more meaningful status might be: FAILOVER READINESS 92 / 100 Any production score should use a published and defensible methodology.

How Often Should You Test Internet Failover?

Do not make the first test a real outage. A reasonable program may include: Continuous passive monitoring Monthly validation Quarterly controlled failover Post configuration change testing Post carrier change testing Post firewall upgrade testing Frequency should reflect business criticality. A hospital, contact center, or high volume retail environment may require more frequent testing than a small administrative office.

What Should an Internet Failover Test Include?

A proper test should verify: Primary path failure detection Backup activation Internet access DNS VPN Cloud apps Voice Critical inbound services Public IP behavior Bandwidth Monitoring Security controls Authentication Business workflows Failback Then document the results. Example Internet Failover Test Environment Primary: 1 Gbps AT&T DIA Backup: 300 Mbps Comcast Business Test Start 9:00 AM Action Primary WAN administratively disconnected. Result 8 sec — primary marked failed 11 sec — backup selected 19 sec — public internet restored 31 sec — site to site VPN restored 36 sec — Microsoft 365 healthy 42 sec — Zoom test call successful 44 sec — CRM healthy Issues Found Partner SFTP failed due to IP allowlist. Guest WiFi consumed 65 Mbps during backup. Overall Status PARTIAL PASS Corrective Action Add backup public IP to partner allowlist. Disable guest WiFi automatically during failover. That is a useful resilience test. Example Failed Failover Primary Failure WAN1 fiber lost. Backup WAN2 interface shows UP. User Experience No internet. Findings Backup gateway reachable. Firewall routing still preferred WAN1. SD WAN health rule tested only physical interface status. WAN1 remained physically up despite 100% upstream packet loss. Root Cause Health monitoring design failed to detect upstream carrier outage. Notice: The backup circuit itself was fine. The failover logic failed.

Why Ping Alone Is Not Enough for Failover Health Checks

Suppose the carrier gateway responds to ping but: DNS fails Internet routing is broken Cloud apps cannot connect Traffic experiences 60% packet loss The circuit should probably be considered unhealthy. Use multiple signals such as: Gateway Public DNS HTTP/HTTPS Packet loss Latency Application probe VPN health A mature failover decision can combine several signals.

Should Failover Be Based on Packet Loss and Latency?

Often, yes. Physical link state alone may detect: Hard outages but miss: Brownouts Congestion Route failures Severe packet loss Extreme latency A policy might define failure as: Gateway unreachable OR Packet loss above threshold for defined duration OR Latency above threshold OR Critical application probes failing The exact thresholds should match the business and network.

Why Does VPN Break After Internet Failover?

Common causes include:

Peer restricted to primary IP Routing not updated IKE does not renegotiate Backup NAT different Tunnel bound to WAN1 Cloud gateway not configured for backup peer Firewall rule missing Authentication policy tied to primary path This is why VPN recovery should be an explicit failover test.

Why Does VoIP Break After Failover?

Possible reasons: High jitter Insufficient bandwidth SIP provider IP restrictions NAT change Firewall session state QoS missing on backup Backup latency higher SIP registration delay SD WAN policy Voice is often the first service that reveals poor backup quality.

Why Do Some Applications Work and Others Fail After Failover?

Different applications depend on different things. Examples: Website only needs HTTPS. CRM may require SSO. Partner portal may require IP allowlisting. VPN requires peer configuration. VoIP needs low jitter. Remote desktop may need private routing. Therefore: Some applications working does not prove complete failover success.

Can a Backup Internet Connection Be Too Slow?

Absolutely. Consider: Primary: 1 Gbps Backup: 50 Mbps If the organization normally consumes 300 Mbps, the backup will saturate immediately unless traffic is prioritized. Failover design should include: Capacity planning QoS Critical application priority Noncritical traffic restriction

What Should You Monitor on the Backup Connection?

Continuously monitor: Availability Gateway Packet loss Latency Jitter Bandwidth Public IP DNS VPN readiness Interface errors Carrier state Modem/ONT status Signal quality for cellular Starlink health where applicable The worst time to discover the backup is broken is after the primary fails.

What Is a Backup Circuit Confidence Test?

A backup circuit confidence test verifies whether the secondary path can carry production traffic before a real outage. This might involve: Periodic application probes through backup WAN DNS resolution VPN health checks HTTP tests Latency measurement Speed/capacity test Route validation Public IP check Small controlled production tests The purpose is: Prove readiness continuously rather than assume it.

How Do You Document a Failover Incident?

Capture: Primary carrier Backup carrier Circuit IDs Failure timestamp Detection timestamp Failover timestamp Recovery timestamp Failback timestamp Packet loss Latency DNS status VPN state Application tests Voice test Bandwidth Public IP change User impact Problems found Corrective actions This should become part of your RCA. Failover Severity Should Be Based on Business Impact Compare: P1 Primary and backup failed. Site offline. P2 Primary failed. Backup active but critical applications broken. P3 Primary failed. Backup active. Business operational but degraded. P4 Primary failed. Backup active. No measurable user impact. That is far more useful than giving every primary circuit failure the same severity.

Should Customers Be Notified When Failover Works?

Often, yes, but the communication should reflect actual impact. Example: The Tampa location's primary internet circuit failed at 8:42 AM. Traffic automatically transitioned to the backup circuit and users remain operational. The primary carrier has been engaged. We are monitoring backup performance while the carrier restores service. That demonstrates value without creating unnecessary panic. The ADAM PULSE Failover Model The monitoring objective should move from: WAN1: DOWN WAN2: UP to: Primary WAN Failure Detection: Confirmed Backup Path: Active Internet: Healthy DNS: Healthy VPN: Healthy Cloud Applications: Healthy Voice: Healthy Backup Utilization: 62% Customer Impact: None Failover Time: 28 seconds Current Risk: Backup is now single path Next Action: Carrier escalation That tells the operator what actually matters. ADAM Pulse is designed to help answer whether the backup will actually work if the primary fails — by evaluating primary and secondary ISP health, gateway status, routing, public IP, DNS, VPN, bandwidth, and failover readiness rather than simply reporting link state. Internet Failover Checklist PRIMARY ☐ WAN1 failed ☐ Gateway failed ☐ Failure timestamp captured ☐ Cause investigated DETECTION ☐ Firewall detected failure ☐ Health checks worked ☐ Brownout detection tested BACKUP ☐ WAN2 active ☐ Gateway healthy ☐ Packet loss acceptable ☐ Latency acceptable ☐ Bandwidth acceptable INTERNET ☐ Public IP connectivity ☐ HTTPS ☐ Multiple external targets DNS ☐ Public DNS ☐ Internal DNS ☐ Resolver path VPN ☐ Site to site tunnels ☐ Cloud tunnels ☐ Private routes ☐ Remote access if required CLOUD ☐ Microsoft 365 ☐ Google Workspace ☐ CRM ☐ ERP ☐ Critical SaaS ☐ Cloud infrastructure VOICE ☐ VoIP ☐ Zoom/Teams ☐ Contact center ☐ Jitter ☐ Packet loss SECURITY ☐ NAT ☐ Firewall policies ☐ IP allowlists ☐ Identity controls CAPACITY ☐ Utilization ☐ QoS ☐ Critical traffic ☐ Noncritical traffic OPERATIONS ☐ Monitoring active ☐ Users operational ☐ Business workflows tested ☐ Customer impact classified FAILBACK ☐ Primary stable ☐ Traffic returned cleanly ☐ Sessions stable ☐ VPN stable ☐ No flapping

Frequently asked questions

How do I know if my backup internet is actually working?

Do not check only whether the interface is up. Verify real internet traffic, DNS, VPNs, cloud applications, voice, bandwidth, monitoring, and critical business workflows over the backup path.

What is internet failover?

Internet failover automatically moves network traffic from a failed primary WAN connection to a secondary connection so business services can continue.

What is the difference between WAN redundancy and failover?

Redundancy means multiple connections exist. Failover is the process of actually moving traffic to the alternate path when the primary fails.

Why is my backup WAN up but I still have no internet?

Possible causes include incorrect routing, DNS failure, firewall policy, NAT, failed health checks, or a backup connection that is technically linked but not providing usable internet service.

Why does internet work after failover but VPN does not?

VPN configuration may depend on the primary public IP, WAN interface, route, NAT behavior, or remote peer configuration.

Why do some applications fail after internet failover?

Applications may depend on IP allowlists, private routes, authentication policies, DNS, VPNs, or other path specific settings.

How do I know whether failover is fast enough?

Measure the time between primary service degradation and the point when critical business applications are restored. Compare that with your recovery objective.

What is a failover brownout?

A brownout occurs when the primary circuit remains technically up but performs poorly enough to disrupt applications. Failover systems that monitor only link state can miss these conditions.

Should backup internet be the same speed as primary?

Not necessarily. It should provide enough capacity for the critical business services that must continue during an outage.

How often should I test backup internet?

Continuously monitor the backup and perform controlled failover tests periodically, especially after carrier, firewall, SD WAN, routing, or VPN changes.

What is failback?

Failback is the process of returning traffic to the preferred primary connection after it recovers.

Why does my connection keep switching between primary and backup?

The primary circuit may be flapping, or failover and failback thresholds may be too aggressive. Stability timers and better health criteria can help.

What should network monitoring show during failover?

Ideally, it should show the primary failure, backup status, actual application health, VPN state, DNS, business impact, failover time, capacity, and recommended next action.

How do I know if my backup internet is actually working?

Do not check only whether the interface is up. Verify real internet traffic, DNS, VPNs, cloud applications, voice, bandwidth, monitoring, and critical business workflows over the backup path.

What is internet failover?

Internet failover automatically moves network traffic from a failed primary WAN connection to a secondary connection so business services can continue.

What is the difference between WAN redundancy and failover?

Redundancy means multiple connections exist. Failover is the process of actually moving traffic to the alternate path when the primary fails.

Why is my backup WAN up but I still have no internet?

Possible causes include incorrect routing, DNS failure, firewall policy, NAT, failed health checks, or a backup connection that is technically linked but not providing usable internet service.

Why does internet work after failover but VPN does not?

VPN configuration may depend on the primary public IP, WAN interface, route, NAT behavior, or remote peer configuration.

Why do some applications fail after internet failover?

Applications may depend on IP allowlists, private routes, authentication policies, DNS, VPNs, or other path specific settings.

How do I know whether failover is fast enough?

Measure the time between primary service degradation and the point when critical business applications are restored. Compare that with your recovery objective.

What is a failover brownout?

A brownout occurs when the primary circuit remains technically up but performs poorly enough to disrupt applications. Failover systems that monitor only link state can miss these conditions.

Should backup internet be the same speed as primary?

Not necessarily. It should provide enough capacity for the critical business services that must continue during an outage.

How often should I test backup internet?

Continuously monitor the backup and perform controlled failover tests periodically, especially after carrier, firewall, SD WAN, routing, or VPN changes.

What is failback?

Failback is the process of returning traffic to the preferred primary connection after it recovers.

Why does my connection keep switching between primary and backup?

The primary circuit may be flapping, or failover and failback thresholds may be too aggressive. Stability timers and better health criteria can help.

What should network monitoring show during failover?

Ideally, it should show the primary failure, backup status, actual application health, VPN state, DNS, business impact, failover time, capacity, and recommended next action.

Bottom line

A green backup WAN interface does not prove your business is protected. Successful failover means: The primary failed. The system detected it. Traffic moved. DNS continued working. VPNs recovered. Cloud applications remained available. Voice remained usable. The backup had enough capacity. Security policies still worked. Users remained productive. And finally: Traffic returned cleanly when the primary recovered. The most useful question is therefore not: “Is my backup connection up?” It is: “If my primary internet failed right now, would my business actually keep working?” That is the difference between having backup internet and having real network resilience.

← More from the ADAM Pulse Knowledge Base