DPO/SafeDPO/OPAD training + eval for teaching tool-using LLMs to refuse falsely-benign MCP exploits
-
Updated
Jul 8, 2026 - Python
DPO/SafeDPO/OPAD training + eval for teaching tool-using LLMs to refuse falsely-benign MCP exploits
The language-level attack defense skill every agent should keep. Covers 12 attack categories. Works with Claude, GPT, Gemini, Copilot, and any LLM.
Risk-adaptive prompt-injection defense layer for commercial APIs and local LLMs.
Pre-registered adversarial robustness study testing whether model-initiated session termination provides defensive coverage beyond refusal training against multi-turn attacks. Minimal Python harness with scorer and analysis pipeline. Pilot on Gemma 4 26b. Preliminary findings in FINDINGS.md.
[ICLR 2025] Reinforced Blue Teaming for VLMs Against Jailbreak Attacks
Tactical AI security posture and prompt injection vulnerability scanner for AI system instructions.
Paired safety-boundary evaluation for LLM safety reasoning methods.
Runtime defense toolkit against prompt injection for LLM APIs — intercepts, analyzes, and protects prompts in real-time
A reproducible safety framework and defense pipeline designed to protect large language models from multi turn jailbreak attacks
Semantic-layer prompt injection defence that separates untrusted instructions from authority while preserving the legitimate task.
Add a description, image, and links to the jailbreak-defense topic page so that developers can more easily learn about it.
To associate your repository with the jailbreak-defense topic, visit your repo's landing page and select "manage topics."