Network troubleshooting checklist: a step-by-step guide to LAN, firewall, ISP and application faults
When someone says:
"The network is down."
that statement tells IT almost nothing.
Is the problem:
- One computer?
- WiFi?
- LAN?
- DNS?
- Firewall?
- ISP?
- VPN?
- Cloud application?
- SD WAN?
- Remote service?
Good network troubleshooting is not random.
It is a process of progressively narrowing the failure domain.
For ADAM Pulse and USA Telecom, a practical troubleshooting model is:
DEFINE → SCOPE → VERIFY → ISOLATE → TEST → ESCALATE → RESTORE → DOCUMENT
This checklist provides a repeatable process IT teams can use to investigate network incidents without immediately blaming the ISP, rebooting everything, or jumping randomly between tools.
What Is Network Troubleshooting?
Network troubleshooting is the systematic process of identifying, isolating and resolving connectivity or performance problems.
The goal is to move from:
"Something is broken."
to:
"The likely problem exists in this part of the network path."
That process should be evidence based.
What Is the First Step in Network Troubleshooting?
Start by defining the problem.
Ask:
What exactly is not working?
Avoid vague descriptions.
Instead of:
"The internet is slow."
determine:
- Which user?
- Which location?
- Which application?
- Wired or wireless?
- When did it begin?
- Is it constant or intermittent?
- Are other users affected?
Good troubleshooting begins with good intake.
Network Troubleshooting Checklist
Use this sequence:
- Define the symptom.
- Record the exact incident time.
- Determine scope.
- Check the endpoint.
- Check local connectivity.
- Check WiFi or LAN.
- Check DNS.
- Check the gateway.
- Check the firewall.
- Check the primary WAN.
- Check the backup WAN.
- Test multiple external destinations.
- Review latency.
- Review packet loss.
- Review jitter when relevant.
- Review routing or path.
- Check failover.
- Check the application.
- Review historical monitoring.
- Identify the probable failure domain.
- Escalate with evidence.
- Verify restoration.
- Document the incident.
- Look for recurrence.
Step 1: Define the Symptom
Determine what the user actually experiences.
Examples:
No internet
Website will not load
Zoom freezes
VoIP audio breaks up
VPN disconnects
Application is slow
These symptoms may have different causes.
Step 2: Record the Exact Time
The timestamp is critical.
Ask:
When did it start?
When did it stop?
Is it happening now?
Historical network data becomes much more useful when correlated to an exact time.
Step 3: Determine the Scope
Ask:
One user?
One department?
One location?
Multiple locations?
Everyone?
Scope immediately narrows the possibilities.
One User
Consider:
- Endpoint
- WiFi
- Local configuration
- Application
One Location
Consider:
- LAN
- Firewall
- WAN
- ISP
Multiple Locations
Consider:
- Shared application
- Cloud provider
- DNS
- Central network
- Common carrier
- SaaS platform
Step 4: Check the Endpoint
Before investigating the carrier, verify the user's device.
Check:
- Network connection
- IP configuration
- WiFi association
- VPN
- Local firewall
- Application state
- Other applications
If one device is affected while everyone else works, begin locally.
Step 5: Check Physical Connectivity
For wired devices:
- Cable connected?
- Link light?
- Switch port active?
- Docking station healthy?
For network infrastructure:
- Power?
- Cabling?
- Interface state?
Physical problems remain common.
Step 6: Check WiFi
If the user is wireless, evaluate:
- Signal
- Access point
- Interference
- Roaming
- Congestion
- Client behavior
Compare with a wired device when practical.
If wired works and WiFi does not, the ISP becomes less likely.
Step 7: Check IP Configuration
Validate:
- IP address
- Subnet
- Gateway
- DNS
An incorrect configuration can look like an internet outage.
Step 8: Test the Local Gateway
Can the endpoint reach its gateway?
If not, investigate:
- LAN
- VLAN
- WiFi
- Switching
- Local routing
Do not immediately escalate to the ISP.
Step 9: Test DNS Separately
Can the device reach an external IP address but not resolve a hostname?
If yes, investigate DNS.
This distinction is important:
Connectivity works
but:
Name resolution fails.
Users may still describe that as:
"The internet is down."
Step 10: Check the Firewall
Determine:
- Is it reachable?
- Is it powered?
- Are WAN interfaces up?
- Are resources normal?
- Did it reboot?
- Were changes made?
- Are VPNs affected?
The firewall is a major troubleshooting boundary.
Step 11: Check the Primary WAN
Evaluate:
- Link state
- Gateway reachability
- External reachability
- Latency
- Packet loss
- Historical events
A circuit can be online but degraded.
Step 12: Check the Backup WAN
If the site has redundant internet, verify:
- Backup is healthy
- Failover occurred if required
- Backup performance is acceptable
- Critical applications work
Do not assume WAN2 is healthy because WAN1 normally carries traffic.
Step 13: Test Multiple External Destinations
Avoid drawing conclusions from one target.
If:
Destination A fails
but:
B and C work,
the problem may be with A or its path.
If:
All external destinations fail simultaneously,
the issue may be closer to the site or carrier.
Step 14: Review Latency
Compare current latency against:
- Normal baseline
- Other locations
- Previous performance
A major increase can indicate congestion, routing changes or carrier problems.
Step 15: Review Packet Loss
Determine:
- Is loss local?
- Does it begin at the WAN?
- Is it persistent?
- Is it intermittent?
- Does it affect multiple targets?
Packet loss is particularly important for real time applications.
Step 16: Review Jitter
For:
- Zoom
- VoIP
- Video
- Contact centers
review jitter where available.
A network can have acceptable average latency but inconsistent packet delivery.
Step 17: Review the Network Path
Use path testing to understand how traffic reaches the destination.
Look for:
- Route changes
- Increased delay
- Failure boundaries
- Unexpected paths
Remember that individual network hops may treat diagnostic traffic differently, so interpret traceroute carefully and in context.
Step 18: Check for WAN Failover
Ask:
Did traffic move from primary to backup?
If yes:
- Why did it fail over?
- How long did it take?
- Was the backup healthy?
- Did applications survive?
- Did traffic fail back?
Failover itself can explain short application interruptions.
Step 19: Check the Application
If infrastructure appears healthy, test the actual service.
Examples:
- Zoom
- Microsoft 365
- CRM
- ERP
- VoIP platform
- SaaS application
Infrastructure health does not guarantee application health.
Step 20: Review Vendor or Cloud Service Status
When appropriate, check whether the application provider reports a service issue.
Do not spend hours changing local infrastructure when the service itself is experiencing an incident.
Step 21: Review Historical Monitoring
If the problem is no longer occurring, history becomes critical.
Review the incident window for:
- Availability
- Latency
- Packet loss
- Jitter
- Firewall status
- WAN state
- Failover
- Previous similar incidents
This is where continuous monitoring dramatically improves troubleshooting.
Step 22: Check What Changed
Ask:
What Changed Before the Problem?
Review:
- Firewall changes
- Firmware
- ISP changes
- Routing
- DNS
- VPN
- SD WAN policy
- Application updates
- New equipment
A recent change is not automatically the cause, but it is an important clue.
Step 23: Identify the Failure Domain
Before escalating, classify the probable problem.
Endpoint
One device.
LAN or WiFi
Local connectivity.
Firewall
Network edge.
WAN or ISP
External connectivity.
DNS
Name resolution.
Application
Service or SaaS platform.
Cloud or Remote Network
External infrastructure.
The goal is not always to know the final root cause immediately.
The goal is to narrow the investigation.
Step 24: Escalate to the Correct Party
Once the failure domain is understood, involve:
- Internal IT
- Network engineer
- Firewall provider
- ISP
- Cloud provider
- Application vendor
- Managed service provider
Escalating too early to the wrong party wastes time.
What Information Should Be Included in an ISP Escalation?
Provide:
- Location
- Circuit ID
- Incident start
- Duration
- Current status
- Firewall status
- Gateway status
- Packet loss
- Latency
- Relevant path information
- Previous similar incidents
Avoid:
"Internet slow. Please investigate."
What Information Should Be Included in an Application Escalation?
Provide:
- Application
- Users affected
- Locations affected
- Exact incident time
- Error messages
- Network health
- Reproduction steps
- Screenshots or logs where appropriate
This helps distinguish application issues from network problems.
Step 25: Verify Restoration
Do not close an incident only because:
The ISP says it is fixed.
or:
The application vendor says resolved.
Independently validate:
- Gateway
- Firewall
- WAN
- External connectivity
- Latency
- Packet loss
- Application access
Then confirm user experience where appropriate.
Step 26: Monitor Stability
A service that comes back for two minutes may not be truly restored.
Continue monitoring for recurrence.
Watch for:
- Flapping
- Packet loss
- Latency
- Repeat failover
- Application errors
Step 27: Document the Incident
Record:
- What happened
- Start time
- End time
- Impact
- Failure domain
- Root cause if known
- Actions taken
- Provider ticket
- Resolution
- Follow up
Documentation turns individual troubleshooting into organizational knowledge.
Step 28: Look for Recurrence
One incident may be random.
Ten similar incidents are a pattern.
Historical monitoring should help answer:
Has this happened before?
Repeated incidents deserve deeper investigation.
What Is Root Cause Analysis?
Root cause analysis attempts to identify the underlying reason an incident occurred.
Examples:
Symptom:
Zoom froze.
Failure domain:
WAN quality.
Root cause:
Carrier packet loss caused by provider issue.
Do not confuse the symptom with the root cause.
What Is Fault Isolation?
Fault isolation narrows the problem to a specific network area.
For example:
User
→ LAN
→ Firewall
→ WAN
→ ISP
→ Internet
→ Application
Testing each boundary helps determine where normal behavior stops.
Why Is Fault Isolation Better Than Random Troubleshooting?
Random troubleshooting often looks like:
Reboot modem.
Change DNS.
Restart firewall.
Call ISP.
Reinstall application.
This may eventually work, but it can destroy evidence and consume time.
Fault isolation follows the network logically.
Should You Start with Ping?
Ping is useful, but it is only one tool.
It can help test:
- Reachability
- Latency
- Packet loss
But it does not fully evaluate:
- DNS
- Application health
- TCP behavior
- WiFi
- Firewall policy
- User experience
Use it as part of a broader process.
When Should You Use Traceroute?
Traceroute can help understand network path and potential changes.
Use it when:
- Latency changes
- Routes appear abnormal
- External destinations fail
- Carrier path questions arise
Interpret intermediate hop behavior cautiously because devices may deprioritize or block diagnostic traffic.
When Should You Use a Speed Test?
A speed test can help measure throughput.
It is useful when investigating:
- Bandwidth
- Congestion
- Capacity
But it should not be treated as the only network health test.
A high speed result does not eliminate:
- Packet loss
- Jitter
- Short outages
- Routing problems
When Should You Use Packet Capture?
Packet capture is a deeper troubleshooting tool useful when basic monitoring cannot explain the problem.
It can help investigate:
- Protocol behavior
- Retransmissions
- Session problems
- Application communication
Packet capture generally requires more technical expertise and should be used purposefully.
What Network Troubleshooting Tools Should IT Teams Have?
A practical toolkit may include:
- Ping
- Traceroute
- DNS tools
- Speed testing
- SNMP monitoring
- Syslog
- Flow monitoring
- Packet capture
- Firewall diagnostics
- Application monitoring
- Continuous WAN monitoring
- Historical performance data
The best tool depends on the question.
What Is the ADAM Pulse Network Troubleshooting Model?
ADAM Pulse can organize troubleshooting around:
USER → LAN → GATEWAY → FIREWALL → WAN → CARRIER → INTERNET → APPLICATION
At every layer, ask:
Is it reachable?
Is performance normal?
What changed?
What does history show?
This creates a repeatable troubleshooting method.
The ADAM Pulse Troubleshooting Framework
DEFINE → SCOPE → VERIFY → ISOLATE → TEST → ESCALATE → RESTORE → DOCUMENT → LEARN
Define
Describe the actual symptom.
Scope
Determine who and what is affected.
Verify
Confirm the condition.
Isolate
Find the likely failure domain.
Test
Collect evidence.
Escalate
Engage the correct party.
Restore
Return service.
Document
Preserve the incident record.
Learn
Identify patterns and improvements.
A Network Troubleshooting Checklist Should Be Used Before the Outage
Do not invent the process while users are waiting.
Create:
- Standard intake questions
- Monitoring points
- Carrier information
- Escalation contacts
- Runbooks
- Restoration procedures
before incidents occur.
ADAM Pulse Turns Troubleshooting into a Repeatable Process
The purpose of network monitoring is not simply to generate graphs.
It should help answer:
What failed?
Where did it fail?
When did it happen?
What remained healthy?
Has it happened before?
Who should act next?
ADAM Pulse provides managed network monitoring designed to help USA Telecom customers detect problems, preserve historical evidence, isolate network failure domains, and improve operational response.
Define the problem.
Follow the path.
Collect the evidence.
Escalate intelligently.
Verify restoration.
Learn from the incident.
Talk with USA Telecom about using ADAM Pulse to create a more systematic network monitoring and troubleshooting process.
Frequently asked questions
What are the basic steps of network troubleshooting?
Define the problem, determine scope, verify local connectivity, test the gateway, firewall and WAN, review DNS and application health, isolate the failure domain, escalate with evidence and verify restoration.
What should I check first when the internet is down?
First determine whether the problem affects one user or an entire location. Then check local connectivity and the gateway before assuming the ISP is responsible.
How do I tell whether a network problem is the firewall or ISP?
Monitor and test both boundaries independently. If the firewall remains reachable while external connectivity fails, the investigation may move toward the WAN or carrier path.
Is ping enough to troubleshoot a network?
No. Ping is useful for reachability, latency and loss testing, but complete troubleshooting may also require DNS testing, path analysis, firewall diagnostics, application testing and historical monitoring.
Why is historical monitoring important for troubleshooting?
Many intermittent problems disappear before technicians investigate. Historical monitoring preserves evidence from the actual incident period.
Sources
- IETF — RFC 792: Internet Control Message Protocol (September 1981, Internet Standard, STD 5). Why devices may rate-limit or deprioritise the ICMP that ping and traceroute depend on.
- Cisco — Troubleshoot Packet Drops. Congestion, buffer exhaustion and interface errors as drop causes.
- Cisco — What Is Network Latency?
- FCC — Measuring Broadband America. Methodology for measuring latency and packet loss alongside throughput.
- NIST — The NIST Cybersecurity Framework (CSF) 2.0 (NIST CSWP 29, 26 February 2024). Continuous monitoring (DE.CM) and the logging that supports it (PR.PS-04).
A single test from a single location at a single moment rarely proves where a fault sits. Correlate against history, test from more than one point, and preserve timestamped evidence before changing configuration or rebooting equipment.
USA Telecom Consulting LLC is a Service-Disabled Veteran-Owned Small Business running a 24/7 NOC. We monitor networks, circuits and firewalls for regulated and defense-supply-chain organizations.
Ping vs traceroute vs MTR · Is it the ISP or the firewall? · What is network fault isolation? · How to troubleshoot packet loss · i · i