A browser extension that detects phishing pages entirely client-side: it looks at a page the way PhishIntention and similar academic systems do — find a brand logo on the page, check whether the page is asking for credentials (a "CRP", credential-requesting page), and flag a mismatch between the claimed brand and the actual domain. All inference (logo matching, OCR, layout detection, CRP classification) runs locally via ONNX Runtime; no page content is sent to a server. This repository contains the extension itself plus everything used to train, export, and evaluate the models it runs.
Project state: This is a prototype. The extension is functional and installable in both Firefox and Chrome but there are open issues.
| Directory | What it is |
|---|---|
extension/ |
The browser extension itself (WXT + TypeScript + React) — the only piece meant to ship. Everything else here supports building, training, or evaluating it. |
training_yolo/ |
Trains and exports the YOLO layout detector (finds logos, input fields, buttons, etc. on a page) to ONNX. |
export_crp_classifier/ |
Trains/downloads and exports the CRP (credential-requesting-page) classifier to ONNX. |
export_logo_matcher/ |
Trains and exports the logo-matching embedding models (and the OCR text-recognition model) to ONNX; see MODELS.md for benchmarks. |
evaluation/ |
Automated pipeline that runs the built extension (via Playwright + a local pywb replay server) against archived phishing/benign samples and reports accuracy. Includes sub-tools for CRP-classifier evaluation, logo-DB deduplication/extension, and whitelist-coverage analysis. |
phish_archiver/ |
CLI that crawls and archives phishing (PhishTank) and benign (Tranco-derived) login pages into WARC files, for use as evaluation's input dataset. |
.github/ |
CI: unit tests (Vitest) and end-to-end tests (Playwright) for the extension. |
Each directory has its own README with setup and usage instructions.