Skip to content

Stage: Research #1

Description

@craig-dt

Stage 1 of 7 · Command: /project:research · Artifacts: docs/research-brief.md + docs/research.md

Ground the idea in options, prior art, and risks. Runs in three parts: grill for a research brief (gated on Craig's sign-off) → execute the research → publish to Notion.

What this stage produces

docs/research-brief.md — Objective · Key questions to answer · Constraints · Exit criteria · Out of scope. Craig must approve this before any research begins.

docs/research.md — anchored to the brief's key questions:

  1. Two or three viable technical approaches with tradeoffs (complexity, maintenance, performance, stack fit).
  2. Existing libraries/services that already solve part of this, each with last-release/maintenance status.
  3. Top 5 risks or unknowns, ranked, and what would de-risk each.
  4. What's needed from Craig to make a recommendation.
  5. 3–5 open questions, also asked in chat.

{RESEARCH} items carried over from docs/prep-n-research.md

These are the reason this stage exists. The labels produced by flabel become ML training ground truth, so verdict trustworthiness is the dominant quality requirement. Every ruleset and feed selected needs written justification recorded in docs/research.md.

Content/inline detection sources

  • Snort/Suricata rulesets: which are HIGH confidence? Industry-standard, detection-focused not threat-hunting, strong emphasis on low false positives. Compare candidates (e.g. ET Pro vs ET Open vs Snort Talos registered/subscriber) on FP posture, licensing, and update cadence.
  • NGFW option: Palo Alto Networks VM-Series vs FortiGate — availability, licensing, and how detections are queried back programmatically.
  • Stretch: is there a free/open-source L7 equivalent to PANW/FortiGate app-level detection?

Encrypted traffic (JA3/JA4)

  • Identify JA3/JA4 fingerprint feeds from HIGHLY trusted sources — multiple, industry-standard / government / research-grade. Document provenance and update cadence for each.
  • How to collate and deconflict the feeds into a single consolidated list (what happens when two feeds disagree on the same fingerprint?).
  • Can a threat name be derived from a JA3/JA4 match, or is the verdict only "known-bad fingerprint"? This directly affects the labels.json schema and how usable the labels are for training.
  • What is the realistic false-positive rate of JA3/JA4 matching (fingerprint collisions across benign and malicious clients)? This bounds achievable label quality.

Cross-cutting

  • Are there existing labeled-pcap corpora or prior-art labeling pipelines to compare against or validate with?
  • How would we measure label quality at all — is there a validation strategy?

Exit criteria

  • Craig signed off on docs/research-brief.md.
  • Every {RESEARCH} checkbox above is answered or explicitly deferred with a reason.
  • Every selected ruleset/feed has a cited justification.
  • Sources cited; uncertainty flagged rather than guessed.
  • Findings published to the Notion tracker row; current_stage advanced to prd.

Metadata

Metadata

Assignees

No one assigned

    Labels

    stageOne of the 7 pipeline stages

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions