Capture the last known good state. Follow the evidence when it changes.
Network outages are easy to notice and harder to reconstruct. By the time troubleshooting starts, the route, resolver state, interface condition, or latency change that mattered may already be gone.
LastKnownGood is a local first Linux troubleshooting and observability tool built to preserve that evidence. The name comes from its central idea: keep a trusted baseline of the network's last known good state, then compare against it when something changes. It detects meaningful changes, correlates related symptoms, tracks incidents through recovery, and produces evidence backed reports.
Healthy state → change → failure → evidence → diagnosis → recovery
Core diagnosis stays local. AWS adds durable storage and operational visibility, but LastKnownGood does not depend on cloud connectivity to investigate an outage that may have broken that connectivity in the first place.
New to LastKnownGood? Read the User Guide for installation, first run steps, command examples, monitoring, Docker demonstrations, AWS integration, and troubleshooting.
A snapshot can include:
- Routes and default gateway
- Network interfaces
- IP addresses
- DNS resolvers
- Failed systemd services
- Reachability probes
- DNS lookup results
- Latency
- Packet loss
- MTU information
The system can then compare a known good baseline against the current state.
A basic monitoring tool might report:
Default route missing
External probe failed
DNS lookup failed
LastKnownGood attempts to connect those symptoms.
For example:
Default route disappeared
|
External reachability failed
|
Multiple network probes failed
|
Likely diagnosis:
Default gateway or routing failure
Confidence: HIGH
The diagnosis is accompanied by the evidence that produced it.
This makes the result easier to investigate and verify instead of simply producing more alerts.
Create an isolated Python environment:
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -e ".[dev]"Capture a known good baseline:
nfr snapshot \
--probe 1.1.1.1 \
--dns example.com \
--output snapshots/baseline.jsonAfter an authorized network change or outage, capture the current state:
nfr snapshot \
--probe 1.1.1.1 \
--dns example.com \
--output snapshots/current.jsonCompare the two states:
nfr compare \
snapshots/baseline.json \
snapshots/current.json \
--output reports/incident.mdNo Linux lab available?
Run the included fixture demonstration:
nfr compare tests/fixtures/healthy.json tests/fixtures/broken.jsonThe project includes a Docker based failure injection lab that creates a real networking failure inside an isolated container.
Run:
./lab/run_demo.shThe demonstration:
- Creates a disposable Docker environment.
- Records the healthy network state.
- Removes the container's default route.
- Captures the resulting outage.
- Compares the healthy and broken states.
- Diagnoses the routing failure.
- Generates an incident report.
- Restores the Docker managed network.
The failure occurs only inside the disposable container.
This initial diagnostic demonstration is preserved as a historical validation record. The current validation record covers the expanded test suite, guarded recovery, and CI checks.
For the approval gated recovery demonstration, run:
./lab/run_guarded_recovery.shThat workflow restores the known good default route inside the container,
verifies the result, and writes reports/remediation.md with before and after
evidence.
LastKnownGood can identify changes such as:
- Default route disappearance
- Gateway or routing changes
- DNS configuration changes
- DNS lookup failures
- Resolver response changes
- Interface state changes
- Address changes
- MTU changes
- Increased latency
- Packet loss
- Reachability failures
- Newly failed systemd services
When multiple symptoms point to the same underlying problem, the report can lead with a correlated diagnosis such as:
Default gateway or routing failure: HIGH confidence
or:
Local interface or link failure: HIGH confidence
Supporting evidence is included with the diagnosis.
Network diagnostic data can expose sensitive infrastructure information.
LastKnownGood supports deterministic pseudonymization of:
- Hostnames
- IP addresses
- Probe targets
Set a private redaction key:
export NFR_REDACTION_KEY="use-a-long-private-value"Then capture a protected snapshot:
nfr snapshot \
--redact \
--probe 1.1.1.1 \
--dns example.com \
--output snapshots/shared.jsonThe same values remain comparable across snapshots without exposing the original network information.
Retention cleanup also defaults to a dry run:
nfr prune --directory snapshots --keep 100 --max-age-days 30Changes are only applied when explicitly requested:
nfr prune \
--directory snapshots \
--keep 100 \
--max-age-days 30 \
--applyLastKnownGood also includes a Terraform managed AWS operations layer.
Current capabilities include:
- Private S3 evidence storage
- S3 versioning
- SSE S3 encryption
- Encryption in transit
- Public access blocking
- Redacted evidence upload
- CloudWatch health metrics
- Incident summary storage
- Minimal operational dashboarding
Infrastructure is defined under:
infra/aws/
This keeps the AWS environment reproducible and managed as code.
Watch mode turns one time snapshot comparison into an incident timeline. It can detect failure and recovery transitions, maintain incident state, preserve evidence, recover an open incident after restart, and write a final lifecycle summary.
Start a bounded watch session with preserved snapshot history:
nfr watch \
--interval 30 \
--cycles 10 \
--probe 1.1.1.1 \
--dns example.com \
--snapshot-dir snapshots/watch \
--snapshot-limit 100Watch mode detects failure and recovery transitions, maintains incident state, protects incident evidence from rolling snapshot cleanup, and writes a final lifecycle summary after recovery.
The watch demonstration captures an outage, identifies the likely routing failure, detects recovery, closes the incident, and records the total outage duration.
A continuous watch session was also validated across two separate outages. Each incident was independently opened, recovered, and closed.
See the CLI reference for all commands.
LastKnownGood is intentionally conservative about remediation.
Detection and diagnosis can be automated.
Making changes to a network requires stronger safeguards.
The recovery layer therefore uses:
- Explicitly allowed remediation actions
- Approval required remediation plans
- Human review before changes
- Guardrails around supported actions
- Automatic post change verification
- Before and after remediation reports
Execution is currently limited to the isolated Docker lab and one tightly scoped route restoration action. The lab also demonstrates automatic rollback when post remediation verification fails. Generic or production remediation is not implemented.
flowchart LR
A[Healthy Network] --> B[Baseline Snapshot]
C[Current Network] --> D[Current Snapshot]
B --> E[Change Analyzer]
D --> E
E --> F[Evidence]
F --> G[Root Cause Correlation]
G --> H[Incident Report]
H --> I[Redaction]
I --> J[AWS Evidence Storage]
J --> K[CloudWatch and Dashboard]
G --> L[Guarded Remediation Plan]
See the architecture document for more detail.
Every push and pull request is checked automatically with GitHub Actions.
The CI pipeline currently includes:
Python
├── Ruff linting
└── pytest
Security
└── Python dependency audit
Infrastructure
├── terraform fmt
├── terraform init
└── terraform validate
This continuously checks the application, dependencies, and infrastructure configuration.
Completed:
- Capture Linux network state
- Compare healthy and current snapshots
- Generate Markdown and JSON incident reports
- Test common outage signatures
Completed:
- Detect address, MTU, latency, and packet loss changes
- Detect DNS resolver behavior changes
- Correlate related symptoms
- Add snapshot redaction
- Add retention controls
- Build a reproducible Docker outage lab
Completed:
- Terraform managed AWS infrastructure
- Private and encrypted S3 evidence storage
- Redacted evidence upload
- CloudWatch health metrics
- Incident summaries
- Minimal operational dashboard
- CI security and infrastructure checks
Completed:
- Approval required remediation plans
- Explicit action allowlist
- Safety guardrails
- Approved recovery execution in the isolated Docker lab
- Automatic recovery verification
- Before and after remediation reporting
- Rollback after failed recovery verification in the isolated Docker lab
Completed:
- Bounded or continuous
nfr watchmonitoring - Failure and recovery transition detection
- Rolling snapshot retention
- Protected incident evidence
Completed:
- Open and closed incident tracking
- Restart recovery for persisted open incidents
- Outage duration calculation
- Final failure to recovery summaries
Multiple sequential incidents have also been validated in one continuous watch session, with each outage independently opened, recovered, closed, and summarized.
See ROADMAP.md for the full roadmap.
- TCP/IP
- Routing
- DNS
- Network interfaces
- Default gateways
- Reachability testing
- Latency
- Packet loss
- Python
- pytest
- Ruff
- Command line application design
- JSON
- Git
- Ubuntu
- systemd
- Network troubleshooting
- Permissions
- Safe subprocess execution
- AWS
- Amazon S3
- Amazon CloudWatch
- IAM
- Terraform
- GitHub Actions
- Docker
- Dependency auditing
- Encryption
- Evidence redaction
- Infrastructure validation
A few rules keep LastKnownGood from becoming an unsafe "self healing network" demo:
- Diagnosis must survive the outage. Core analysis stays local.
- Evidence comes before action. Findings and remediation decisions are tied to observable state.
- A diagnosis must be inspectable. Correlation should show why it reached a conclusion.
- Sensitive evidence is protected before central storage.
- Remediation stays narrow, approved, and verifiable.
- Failure testing stays isolated. The Docker lab changes container networking, not the host.
The project is meant to make troubleshooting clearer without giving automation unrestricted authority over the network.
I document the design lessons behind this project in my Cloud Network Architecture Journal, including local first observability and guarded remediation.
Ashley “Patience” Hopkins
WGU B.S. Cloud and Network Engineering, AWS Track
CompTIA A+ · CompTIA Network+ · LPI Linux Essentials · ITIL 4 Foundation



