(ACL 2026) Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring
git clone https://github.com/haotian-liu/LLaVA.git
conda create -n llava python=3.10 -y
conda activate llava
cd LLaVA && pip install -e . && cd ..
pip install -r requirements_llava.txt # additional packages and dependencies required to run our experiments (please do not skip)Do not switch the order of the last two commands above. After running the last, it might tell you about a version mismatch that llava requires an older PyTorch, it is fine.
We also provide environment requirement files for different models. Our scripts kcd.py and mcd.py choose the model from the command line (--model / -m, or a positional llava, qwen, or internvl argument), then map that choice to a fixed local checkpoint path under model/. If no model is specified, they default to llava.
For Qwen2.5-VL:
conda create -n qwen25vl python=3.10 -y
conda activate qwen25vl
pip install -r requirements_qwen25vl.txtFor InternVL3-8B:
conda create -n internvl3 python=3.10 -y
conda activate internvl3
pip install -r requirements_internvl3.txtThe project supports multiple vision-language models such as LLaVA, Qwen, and InternVL. Download all required models:
python download_models.pyThis will download the following models to the model/ directory:
- LLaVA-v1.6-Vicuna-7B: Default for most experiments
- FLAVA: Facebook's multimodal model for baseline comparisons (
code/baseline_flava.py) - Qwen2.5-VL-3B-Instruct: Qwen vision-language model
- Qwen2.5-VL-7B-Instruct: Larger Qwen model (~13GB, default in our Qwen experiments)
- InternVL3-8B: OpenGVLab InternVL3 (~15GB)
Note that you can customize the models to download by editing MODELS_TO_DOWNLOAD in download_models.py by considering the experiments you want to run and available disk space. If you only need to run the main experiment, you only need to download the specific target model (FLAVA only for the baseline).
The model downloading script will download the model to ./model locally for faster testing and development. If you wish not to do so, you can skip this step and the model will download when you first run the script.
python download_datasets.pyAfter this, download the rest of the datasets with this link (recommended) or manually following the instructions in the terminal.
python verify_setup.py
mkdir resultsRun experiments from the project root directory (not inside the code directory).
python code/kcd.py # by default llava is used
python code/mcd.py --model qwen # or -m for controlling which model to run
python code/run_multiple_experiments.py --script kcd --model qwen --runs 5 # repeat with different seeds- Scripts
kcd.py,mcd.pyare the main scripts of our methods. - Scripts starting with
hidden_detect_are our best-effort replication of HiddenDetect (ACL 2025) in our scenario, including its proposed layer selection heuristics and detection. - Use
run_multiple_experiments.pyto run an experiment multiple times and aggregate the results. feature_cache,load_datasets,profiling_utils,feature_extractor*are helper scripts- Code in
analysiscan be used to replicate several visualizations such as PCA analysis and visualization of our layer selection heuristics.
Please contact Peichun Hua at peichunhua04@gmail.com for any question about the code or paper.
If you use this code or find our work helpful, please cite:
@inproceedings{hua2026rethinking,
author = {Hua, Peichun and Li, Hao and Shi, Shanghao and Yu, Zhiyuan and Zhang, Ning},
editor = {Liakata, Maria and Moreira, Viviane P. and Zhang, Jiajun and Jurgens, David},
booktitle = "Proceedings of the 64th Annual Meeting of the {A}ssociation for {C}omputational {L}inguistics (Volume 1: Long Papers)",
month = jul,
year = "2026",
address = "San Diego, California, United States",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2026.acl-long.992/",
pages = "21748--21785",
ISBN = "979-8-89176-390-6",
}