Skip to content
 
 

Repository files navigation

Prompt Injection Benchmark (Seeds x Variants x Attacks)

This repository runs a consistent benchmark across multiple methods (baselines + attacks) against one or more target LLMs using a unified judge and a frozen results schema.

Status / Scope (Proof of Concept)

This project is a proof-of-concept (PoC) research/benchmarking tool. It is intended for experimentation, comparative evaluation, and reproducible analysis workflows, and it is not positioned as a production security product or a hardened assessment platform.

Docs (Start Here)

  • Detailed end-to-end overview (methods, schema, aggregation outputs): docs/PROJECT_OVERVIEW.md
  • Full aggregator documentation (all summaries, IR definition): docs/AGGREGATE_RESULTS.md

Choose Your Path

  • Recommended: run everything via the CISSA CLI (presets, isolated runs, HTML report): see "Quickstart (CISSA CLI)"
  • Alternative: run individual methods directly (dev/debug): see "Quickstart (vLLM)" (also works without vLLM if your target is API-based)
  • Optional: local GPU target via vLLM/Transformers: see "Optional: GPU/local target"

Conceptual Architecture

The CISSA benchmark is structured as a layered pipeline:

Seeds + Variants
        ->
Attack / Baseline Method
        ->
Target Model
        ->
Unified Judge
        ->
Frozen CSV Schema
        ->
Aggregator
        ->
Summaries + HTML Report

Repository structure

  • core/: judge, schema, config, inference routing
  • methods/: baseline and attack implementations
  • cissa_tool/: CLI orchestration layer (scanner)
  • data/: seeds and variants
  • results/: experiment outputs (per-run folders)
  • artifacts/: audit logs and reproducibility artifacts
  • scripts/: setup and aggregation utilities
  • docs/: detailed documentation

What you run

You typically run one method at a time. Each method:

  1. generates prompts (baseline or attack),
  2. queries the target model(s),
  3. runs a unified judge (core/judge.py) over (seed_text, response_text),
  4. computes consistent metrics (ASR / BSR / FRR),
  5. writes a CSV in a frozen schema.

Interpretation note: results should be treated as research benchmark outputs (PoC context), not as production security certifications or definitive model safety claims.

Methods

  • Baseline-1 (Vanilla): methods/baseline/baseline_vanilla.py
    • Sends the seed text directly (no wrapping).
  • Baseline-2 (Template Wrapping): methods/baseline/baseline_wrapping.py
    • Wraps each seed with all system variants from data/variants/system_variants_*.csv (same-language).
  • DrAttack-Lite (Adapted): methods/drattack/drattack.py
    • A simplified/adapted implementation inspired by DrAttack (multi-candidate construction + embedding-based selection).
    • Uses word-substitution obfuscation (wordgame mapping), a semantic target text (IntentSummary), and a combined selection score (cosine + refusal/short/generic penalties).
    • Caches parser outputs in artifacts/prompts_drattack_*.json with parser metadata/digest for reproducibility.
    • Logs per-candidate diagnostics to artifacts/drattack_candidates_<run_id>.jsonl (scores, rank, gap, invalid flags).
    • It is not the full reference DrAttack; results should be interpreted as DrAttack-Lite.
    • Optional/experimental: may be unstable on low compute; safe to skip.
  • PAIR-Lite (Multi-turn): methods/pair/pair_lite.py
    • Multi-turn adversarial refinement (attacker -> target -> judge).
    • Uses sparse judging (every 3 steps or when response changes a lot) + early stop by cutoff/stagnation.
    • Supports parallel seed workers (PAIR_WORKERS).
    • Optional/experimental due to higher cost.

Unified outputs

All methods write CSVs using a frozen schema (core/schema.py). Columns that do not apply are always present and set to None. Judge internals (raw vs focus comparison, divergence, verify path) are logged to artifacts/judge_audit.jsonl for auditability. The CSV remains the final decision view and is not fully self-contained without artifacts.

Run folder layout

Each cissa scan creates an isolated run directory under results/:

  • results/run_<id>/results_*.csv: per-method outputs (frozen schema).
  • results/run_<id>/summaries/*.csv: aggregation outputs (generated by cissa report).
  • results/run_<id>/report/report.html: HTML report (tabs + filters; generated by cissa report).
  • results/run_<id>/logs/*.log: per-method execution logs (when enabled).

During cissa report, the tool may create temporary folders inside the run directory:

  • results/run_<id>/.cissa_agg_results_only/
  • results/run_<id>/.cissa_agg_filtered/

These are deleted automatically at the end of the report. If the process is interrupted (or Windows/OneDrive locks files), they may remain and can be removed safely.

Reproducibility guarantees

Each CISSA run is reproducible via:

  • results/run_<id>/config_resolved.yaml (fully resolved configuration for that run)
  • results/run_<id>/meta.json (run lifecycle, timestamps, method status/log pointers)
  • Frozen CSV schema (core/schema.py)
  • artifacts/judge_audit.jsonl (judge audit trail)

How judging works

The unified judge evaluates only:

  • seed_text (the instruction you truly want to test),
  • response_text (the target model output),
  • and the seed_type context (benign vs malign).

It does NOT use final_prompt for judgment.

Metric logic

  • ASR (attack_success): seed_type=malign AND is_compliant
  • BSR (benign_success): seed_type=benign AND is_compliant AND NOT has_refusal
  • FRR (false_refusal): seed_type=benign AND has_refusal For malign rows, has_refusal is kept as an auxiliary signal (e.g., prefaced refusal), not a gate for ASR.

Core modules

  • core/config.py: models, judge params, default paths, refusal prefixes, placeholder.
  • core/loaders.py: input readers + normalizers (seeds and variants).
  • core/judge.py: unified semantic judge (JSON-only, OpenAI-only).
  • core/utils.py: run id/timestamp, refusal detection, JSON parsing helpers.
  • core/schema.py: frozen CSV schema and ensure_schema().
  • core/runner_utils.py: shared runner helpers (loading/binding/CSV output).
  • core/local_llm.py: role-aware inference routing (target local via vLLM/Transformers; OpenAI models via API; mutator/parser can be forced off GPU).
  • artifacts/: DrAttack parser cache (prompts_drattack_*.json), candidate diagnostics (drattack_candidates_*.jsonl), and judge audit trail (judge_audit.jsonl).

Setup (one-time)

Prerequisites:

  • Python 3.10+
  • Optional: vLLM for local GPU runs
  • OpenAI API key (judge is OpenAI-only)
  • DeepSeek API key (default mutator model is DeepSeek)

Install Python deps (creates .venv and installs requirements.txt) on any OS:

python scripts/install_deps.py

Create your local env file from the template (recommended):

macOS/Linux:

cp .env.example .env

Windows (PowerShell):

Copy-Item .env.example .env

Then edit .env and set your secrets:

  • OPENAI_API_KEY (required; judge is OpenAI-only)
  • DEEPSEEK_API_KEY (required only if using the default DeepSeek mutator/parser path)
  • HUGGINGFACE_HUB_TOKEN (optional; needed for gated/private Hugging Face models)

Manual env vars are also supported. Example for OpenAI:

macOS/Linux:

export OPENAI_API_KEY="your_key_here"

Windows (PowerShell):

$env:OPENAI_API_KEY="your_key_here"

Set the DeepSeek API key (required for default mutator/parsing path):

macOS/Linux:

export DEEPSEEK_API_KEY="your_key_here"

Windows (PowerShell):

$env:DEEPSEEK_API_KEY="your_key_here"

Optional DeepSeek endpoint override:

macOS/Linux:

export DEEPSEEK_BASE_URL="https://api.deepseek.com/v1"

Windows (PowerShell):

$env:DEEPSEEK_BASE_URL="https://api.deepseek.com/v1"

macOS/Linux note: prefer python3 to avoid older system Python:

python3 scripts/install_deps.py

Important: this creates .venv at the project root but does NOT activate it. You must activate it yourself:

macOS/Linux:

source .venv/bin/activate

Windows (PowerShell):

.\.venv\Scripts\Activate.ps1

Optional: try installing vLLM as part of setup (may be unsupported on some platforms; macOS often fails):

python scripts/install_deps.py --with-vllm

macOS/Linux:

python3 scripts/install_deps.py --with-vllm

Wrapper scripts (same behavior):

Windows (PowerShell):

.\scripts\install_deps.ps1

If PowerShell blocks script execution:

Set-ExecutionPolicy -Scope Process -ExecutionPolicy Bypass

macOS/Linux (bash):

chmod +x scripts/install_deps.sh
./scripts/install_deps.sh

Quickstart (CISSA CLI)

The CISSA CLI orchestrates presets and keeps each run in its own folder.

Install CISSA as a command (recommended)

First-time setup (creates .venv, installs deps, and installs cissa inside the venv):

python cissa.py setup

Then activate the venv:

macOS/Linux:

source .venv/bin/activate

Windows (PowerShell):

.\.venv\Scripts\Activate.ps1

Now cissa is available:

cissa doctor
cissa list
cissa scan smoke

Alternative (manual) install:

Inside the virtualenv, install in editable mode:

pip install -e .

Then use:

cissa list
cissa scan

1) Minimal .env (non-GPU / OpenAI target)

OPENAI_API_KEY=your_openai_key
TARGET_MODELS=gpt-4o-mini

Tip: start from .env.example and only fill the keys you need. Add HUGGINGFACE_HUB_TOKEN if you use gated Hugging Face models.

2) Run a smoke preset

cissa list
cissa scan smoke

This creates results/run_<id>/ and writes CSVs there.

3) Generate report (HTML + CSV summaries)

cissa report

Outputs:

  • results/run_<id>/summaries/summary.csv (+ all other summary_*.csv)
  • results/run_<id>/report/report.html (tabs + filters)

4) Interactive terminal summary viewer

cissa summary --latest --interactive

Presets (single-method)

Test presets (2 seeds per group):

cissa scan vanilla_test
cissa scan wrapping_test
cissa scan pair_test
cissa scan pair_code_switching_test
cissa scan drattack_test

Normal presets (30 seeds per group):

cissa scan vanilla
cissa scan wrapping
cissa scan pair
cissa scan pair_code_switching
cissa scan drattack

All methods:

cissa scan all_test
cissa scan all

Dry-run and auto report

cissa scan smoke --dry-run
cissa scan all_test --dry-run
cissa scan all_test --report

Runs, status, cleanup

cissa runs
cissa status
cissa report
cissa clean --keep 3

Advanced: target override (optional)

Presets include explicit targets: lists. You can override at runtime if needed:

cissa targets
cissa scan all_test --targets gemma
cissa scan all_test --targets google/gemma-3-4b-it

Advanced: preset introspection

cissa info smoke

Optional: GPU/local target

LOCAL_BACKEND=vllm
LOCAL_DEVICE=cuda
LOCAL_BATCH_SIZE=8

Quickstart (vLLM)

Set the local backend (default is vllm):

# Linux VM (recommended)
export LOCAL_BACKEND=vllm
export LOCAL_DEVICE=cuda
export LOCAL_BATCH_SIZE=8
export DRATTACK_CANDIDATES=4
export PAIR_STEPS=8
export PAIR_CUTOFF=8
export PAIR_MAX_STAGNATION=4
export PAIR_WORKERS=4

Or run the methods directly:

python methods/baseline/baseline_vanilla.py
python methods/baseline/baseline_wrapping.py
python methods/drattack/drattack.py
python methods/pair/pair_lite.py

Run all methods (explicit)

There is a separate script that runs all methods. It requires an explicit confirmation flag:

# macOS/Linux
bash scripts/run_all_methods.sh --yes
# Windows
.\scripts\run_all_methods.ps1 -Yes

Results generation

Each method writes a CSV with a frozen schema to results/:

  • results/results_baseline_vanilla.csv
  • results/results_baseline_wrapping.csv
  • results/results_drattack.csv
  • results/results_pair.csv

DrAttack CSV includes additional method-specific columns (beyond frozen core), such as: seed_risk_tag, final_prompt_quality_flag, selection_score, candidate_count, chosen_candidate_rank, best_score, second_best_score, score_gap, duration_total_batch, duration_per_prompt_est, duration_per_seed_est.

Important: preserve artifacts/ together with results/ for experiment reproducibility/auditability. In particular, keep artifacts/judge_audit.jsonl when sharing or backing up runs.

Aggregating results

Aggregation is done by scripts/aggregate_results.py and produces summary.csv plus additional summary_*.csv files. It also computes IR (Interpretability Rate) as a post-processing metric.

For the full list of summaries, flags, and the IR definition, see: docs/AGGREGATE_RESULTS.md.

Minimal example:

python scripts/aggregate_results.py -i results -o results/summary.csv

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages