Note
TLDR This project is a system designed to compare various phishing detection methods on visual data. It was developed as part of a bachelor's thesis at the Warsaw University of Technology. The implemented methods are:
- Phishpedia - a CNN model for phishing detection based on logos
- VisualPhish - a deep learning model for phishing detection based on visual features
- Baseline: perceptual hashing and similarity search, custom implementation
A summary of all evaluation runs can be found in EVALUATION_RESULTS.md.
Article published in MDPI Applied Sciences. See Citations for more details.
- Phishing Target Recognition
- Table of Contents
- Technologies Used
- Prerequisites
- Quickstart
- Unit Tests
- Citations
This project is built with a modern Python stack, leveraging the following key technologies:
-
Programming Language:
- Python 3.9
-
Machine Learning & Data Processing:
- PyTorch
- scikit-learn
- NumPy & Pandas
- Pillow (PIL)
- FAISS
-
API & Web Interface:
- FastAPI
- Streamlit
- Docker
-
Development & Tooling:
- uv
- just
- WandB
Before starting, install these tools:
- Just: A command runner. Installation instructions can be found here.
- uv: An advanced Python package and environment manager. Install using the following commands:
curl -LsSf https://astral.sh/uv/0.5.18/install.sh | sh source $HOME/.local/bin/env
- unzip: A tool for decompressing ZIP files.
Initial setup (required for all paths):
- Install development tools:
just tools
- Set the
PROJECT_ROOT_DIRenvironment variable to point to the main project directory:You can also add this command to your shell configuration file (e.g.,export PROJECT_ROOT_DIR=$(pwd)
~/.zshrcor~/.bashrc) to make it available in every new terminal session:echo "export PROJECT_ROOT_DIR=$(pwd)" >> ~/.zshrc # For Zsh # or echo "export PROJECT_ROOT_DIR=$(pwd)" >> ~/.bashrc # For Bash source ~/.zshrc # or source ~/.bashrc
This project uses the src/data_splitter tool to create 60:20:20 train/val/test splits and automatically organize the
data
for Phishpedia and VisualPhish.
The CSV must include:
"url",
"fqdn",
"screenshot_object",
"affected_entity",
"is_phishing"Add a data_split section. Choose a label_strategy that matches your dataset layout (subfolders, labels_file, or
directory).
{
"data_split": {
"random_state": 42,
"output_directory": "data_splits",
"create_symlinks": true,
"datasets": {
"my_dataset": {
"path": "data/interim/my_dataset",
"label_strategy": "directory",
"target_mapping": {
"phishing": "phishing",
"benign": "trusted_list"
}
}
}
}
}Reference for label strategies:
subfolders: Each sample has its own folder containingshot.pngandinfo.txt.labels_file: A flat directory of images with alabels.txtfile.directory: Images grouped by target as subdirectories, e.g., the VisualPhish dataset.
just setup-data-splitter
uv run src/data_splitter/split_data.py config.jsonOutputs under $PROJECT_ROOT_DIR/data_splits/my_dataset:
train.csv,val.csv,test.csvvisualphish/data/{train,val,test}/{trusted_list,phishing}/...phishpedia/data/{train,val,test}/{trusted_list,phishing}/...withshot.png/info.txt
Set up (via just; uses src/models/phishpedia/justfile):
# Install dependencies and download models
just run-pp setup
# Or only download models (if dependencies are already installed)
just run-pp setup-modelsmkdir -p $PROJECT_ROOT_DIR/logs/phishpedia
uv run src/models/phishpedia/phishpedia.py \
--folder $PROJECT_ROOT_DIR/data_splits/my_dataset/phishpedia/data/test/phishing \
--output_txt $PROJECT_ROOT_DIR/logs/phishpedia/test_results.txt \
--logSet up once:
cd src/models/visualphishnet
uv sync
uv run wandb login YOUR_API_KEYTrain (uses trusted_list and phishing under .../visualphish/data/train):
uv run src/models/visualphishnet/trainer.py \
--dataset-path $PROJECT_ROOT_DIR/data_splits/my_dataset/visualphish/data/train \
--logdir $PROJECT_ROOT_DIR/logs/visualphish/my_dataset \
--output-dir $PROJECT_ROOT_DIR/data/processed/VisualPhish/my_datasetEvaluate on test set:
uv run src/models/visualphishnet/eval_new.py \
--emb-dir $PROJECT_ROOT_DIR/data/processed/VisualPhish/my_dataset \
--data-dir $PROJECT_ROOT_DIR/data_splits/my_dataset/visualphish/data/test \
--phish-folder phishing \
--benign-folder trusted_list \
--threshold 8.0 \
--result-path $PROJECT_ROOT_DIR/logs/visualphish/my_dataset \
--save-folder $PROJECT_ROOT_DIR/logs/visualphish/my_dataset_resultsNote (macOS TMP space):
export TMPDIR=$HOME/tmpUse the EER-based optimizer to tune the decision threshold using validation splits created by the splitter.
uv run src/models/visualphishnet/threshold_optimizer.py \
--emb-dir $PROJECT_ROOT_DIR/data/processed/VisualPhish/my_dataset \
--val-phish-dir $PROJECT_ROOT_DIR/data_splits/my_dataset/visualphish/data/val/phishing \
--val-benign-dir $PROJECT_ROOT_DIR/data_splits/my_dataset/visualphish/data/val/trusted_list \
--mean 8 --std 22 --max 100 \
--output-dir $PROJECT_ROOT_DIR/logs/visualphish/threshold_opt \
--plotOutputs:
$PROJECT_ROOT_DIR/logs/visualphish/threshold_opt/optimal_threshold.json$PROJECT_ROOT_DIR/logs/visualphish/threshold_opt/threshold_sweep.csvUse the reported threshold witheval_new.pyvia--threshold.
Set up the environment:
cd src/models/baseline
uv syncBuild the FAISS index from the training data (phishing, then benign):
mkdir -p $PROJECT_ROOT_DIR/logs/baseline
cd $PROJECT_ROOT_DIR/src/models/baseline
uv run load.py \
--images $PROJECT_ROOT_DIR/data_splits/my_dataset/visualphish/data/train/phishing \
--index $PROJECT_ROOT_DIR/data/processed/baseline/my_dataset.faiss \
--is-phish \
--batch-size 256 \
--log
uv run load.py \
--images $PROJECT_ROOT_DIR/data_splits/my_dataset/visualphish/data/train/trusted_list \
--index $PROJECT_ROOT_DIR/data/processed/baseline/my_dataset.faiss \
--batch-size 256 \
--log \
--appendQuery the test data (run once per class):
uv run query.py \
--images $PROJECT_ROOT_DIR/data_splits/my_dataset/visualphish/data/test/phishing \
--index $PROJECT_ROOT_DIR/data/processed/baseline/my_dataset.faiss \
--output $PROJECT_ROOT_DIR/logs/baseline/test_phishing.csv \
--threshold 0.5 \
--batch-size 256 \
--log \
--is-phish
uv run query.py \
--images $PROJECT_ROOT_DIR/data_splits/my_dataset/visualphish/data/test/trusted_list \
--index $PROJECT_ROOT_DIR/data/processed/baseline/my_dataset.faiss \
--output $PROJECT_ROOT_DIR/logs/baseline/test_benign.csv \
--threshold 0.5 \
--batch-size 256 \
--logRun grid search on the validation set to find a good threshold. Ensure you point to the validation CSV and images produced by the splitter.
cd $PROJECT_ROOT_DIR/src/models/baseline
uv run threshold_optimization.py \
--val-csv $PROJECT_ROOT_DIR/data_splits/my_dataset/visualphish/val.csv \
--data-base $PROJECT_ROOT_DIR/data_splits/my_dataset/visualphish/data/val \
--index-path $PROJECT_ROOT_DIR/data/processed/baseline/my_dataset.faissOutputs include results_summary.csv and best-threshold files:
best_threshold_identification_rate.txtbest_threshold_mcc.txtbest_threshold_target_mcc.txtUse the chosen threshold with subsequentquery.pyruns on test sets.
docker-compose up -d
uv run streamlit run src/website.pyPaths are relative to the project root unless stated otherwise.
-
api: No host volumes are mounted. Configuration is via environment variables (
MODELS,PORT,VP_PORT,PP_PORT,BS_PORT). -
visualphish:
./data/processed/VisualPhish/model2.h5→/code/model/model2.h5(trained model)./data/processed/VisualPhish/whitelist_emb.npy→/code/model/whitelist_emb.npy(embeddings)./data/processed/VisualPhish/whitelist_file_names.npy→/code/model/whitelist_file_names.npy(file name list)./data/processed/VisualPhish/whitelist_labels.npy→/code/model/whitelist_labels.npy(labels)
-
phishpedia:
./src/models/phishpedia/models→/code/models(model directory)./src/models/phishpedia/LOGO_FEATS.npy→/code/LOGO_FEATS.npy(logo features)./src/models/phishpedia/LOGO_FILES.npy→/code/LOGO_FILES.npy(logo file mapping)
-
baseline:
./src/models/baseline/index.faiss→/code/index/index.faiss(FAISS index)./src/models/baseline/index.csv→/code/index/index.csv(index metadata)
- Ensure
config.jsonpaths are relative toPROJECT_ROOT_DIRunless absolute.
Located in scripts/ and src/models/*/:
-
scripts/augment_data.py: Augment benign samples per target to reach a minimum count; outputs a mirrored structure with a log file.uv run scripts/augment_data.py $PROJECT_ROOT_DIR/data_splits/my_dataset/visualphish/data/train \ --output $PROJECT_ROOT_DIR/data_splits/my_dataset_aug/visualphish/data/train \ --threshold 20
-
scripts/create_dataset.py: Build a VisualPhish-formatted dataset from separate benign/phishing folders using symlinks and mappings.uv run scripts/create_dataset.py \ --benign-dir $PROJECT_ROOT_DIR/path/to/benign \ --phishing-dir $PROJECT_ROOT_DIR/path/to/phishing \ --output-dir $PROJECT_ROOT_DIR/data/interim/VisualPhish
-
scripts/combine_images.py: Combine three example images into a single figure (for reporting/presentations).uv run scripts/combine_images.py --images img1.png img2.png img3.png --layout horizontal -o combined.png
-
scripts/setup.sh: A convenience bootstrap script (installs uv/just, sets PROJECT_ROOT_DIR, performs a wandb login). Review before running. -
VisualPhish utilities:
src/models/visualphishnet/generate_whitelist_filenames.py: Export whitelist filenames fromtrusted_list.src/models/visualphishnet/evaluate_visualphishnet.py: Evaluate a results CSV into metrics and an optional ROC curve.
-
Baseline utilities:
src/models/baseline/load.py,src/models/baseline/query.py: index and query helpers used above.src/models/baseline/threshold_optimization.py: threshold grid search (see section above).
Unit tests are located under src/api/tests and are configured via src/api/pyproject.toml.
Run them with uv:
cd src/api
uv sync --frozen --group dev
uv run pytestExamples:
# Run a single test
uv run pytest tests/test_routes.py::TestPredictEndpoint::test_predict_endpoint_success@Article{app16020640,
AUTHOR = {Jarczewski, Marcin and Białczak, Piotr and Mazurczyk, Wojciech},
TITLE = {Phishing Website Impersonation: Comparative Analysis of Detection and Target Recognition Methods},
JOURNAL = {Applied Sciences},
VOLUME = {16},
YEAR = {2026},
NUMBER = {2},
ARTICLE-NUMBER = {640},
URL = {https://www.mdpi.com/2076-3417/16/2/640},
ISSN = {2076-3417},
ABSTRACT = {With the rapid advancements in technology, there has been a noticeable increase in phishing attacks that exploit users by impersonating trusted entities. The primary attack vectors include fraudulent websites and carefully crafted emails. Early detection of such threats enables the more effective blocking of malicious sites and timely user warnings. One of the key elements in phishing detection is identifying the entity being impersonated. In this article, we conduct a comparative analysis of methods for detecting phishing websites that rely on website screenshots and recognizing their impersonation targets. The two main research objectives include binary phishing detection to identify malicious intent and multiclass classification of impersonated targets to enable specific incident response and brand protection. Three approaches are compared: two state-of-the-art methods, Phishpedia and VisualPhishNet, and a third, proposed in this work, which uses perceptual hash similarity as a baseline. To ensure consistent evaluation conditions, a dedicated framework was developed for the study and shared with the community via GitHub. The obtained results indicate that Phishpedia and the Baseline method were the most effective in terms of detection performance, outperforming VisualPhishNet. Specifically, the proposed Baseline method achieved an F1 score of 0.95 on the Phishpedia dataset for binary classification, while Phishpedia maintained a high Identification Rate (>0.9) across all tested datasets. In contrast, VisualPhishNet struggled with dataset variability, achieving an F1 score of only 0.17 on the same benchmark. Moreover, as our proposed Baseline method demonstrated superior stability and binary classification performance, it should be considered as a robust candidate for preliminary filtering in hybrid systems.},
DOI = {10.3390/app16020640}
}