Skip to content

Add URL input normalization and deduplication regression tests #24

Description

@tayfuryldz

Context

URL ingestion is a boundary of the scanner, but its normalization and deduplication behavior is currently exercised only indirectly. A regression here can silently change which targets are scanned.

Scope

Add focused tests for src/headerproof/input.py. Keep production behavior unchanged unless a test exposes a clearly incorrect edge case; discuss that before changing implementation.

Cover bare hostnames and scheme-relative URLs; JSONL extraction from supported target keys; invalid input; fragment removal while preserving path/query; stable duplicate suppression and max_urls; and SQLite-backed deduplication across repeated input.

Acceptance criteria

  • Tests are deterministic and use only local temporary files.
  • Duplicate targets are emitted once without changing first-seen order.
  • Invalid records are skipped without becoming scan targets.
  • In-memory and SQLite-backed deduplication have equivalent observable target output for the same input.
  • Existing tests remain green.

Comment before starting so ownership can be confirmed.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions