Context
URL ingestion is a boundary of the scanner, but its normalization and deduplication behavior is currently exercised only indirectly. A regression here can silently change which targets are scanned.
Scope
Add focused tests for src/headerproof/input.py. Keep production behavior unchanged unless a test exposes a clearly incorrect edge case; discuss that before changing implementation.
Cover bare hostnames and scheme-relative URLs; JSONL extraction from supported target keys; invalid input; fragment removal while preserving path/query; stable duplicate suppression and max_urls; and SQLite-backed deduplication across repeated input.
Acceptance criteria
- Tests are deterministic and use only local temporary files.
- Duplicate targets are emitted once without changing first-seen order.
- Invalid records are skipped without becoming scan targets.
- In-memory and SQLite-backed deduplication have equivalent observable target output for the same input.
- Existing tests remain green.
Comment before starting so ownership can be confirmed.
Context
URL ingestion is a boundary of the scanner, but its normalization and deduplication behavior is currently exercised only indirectly. A regression here can silently change which targets are scanned.
Scope
Add focused tests for src/headerproof/input.py. Keep production behavior unchanged unless a test exposes a clearly incorrect edge case; discuss that before changing implementation.
Cover bare hostnames and scheme-relative URLs; JSONL extraction from supported target keys; invalid input; fragment removal while preserving path/query; stable duplicate suppression and max_urls; and SQLite-backed deduplication across repeated input.
Acceptance criteria
Comment before starting so ownership can be confirmed.