A CLI that statically analyses prompt template files and reports token budget, variable injection risk, instruction conflicts, and structural metadata. Targets prompt engineers and developers maintaining large prompt libraries.
- Token counting across multiple model families (gpt-3.5-turbo, gpt-4, gpt-4o, text-davinci-003) using tiktoken — fully offline, no API keys required
- Variable extraction for
{{double_curly}},{single_curly}, and Jinja2{% block %}styles - Injection-risk scoring: heuristic rules flag unescaped, user-controlled variable slots that sit inside system/instruction blocks, raw format-string (
{var}) interpolation, and slots placed next to privileged/override language — each flag includes a line number, severity (warn/error), and a suggested fix - Instruction-conflict detection: parses imperative sentences (regex + an offline spaCy model), clusters them by topic keyword (language, verbosity, tone, format, emoji use, disclaimers, etc.), and flags pairs that contradict each other — either by opposite polarity ("never use emojis" vs. "feel free to use emojis") or by pinning incompatible values for the same topic ("respond in English" vs. "reply in the user's language")
- Readability scoring: Flesch Reading Ease, Flesch-Kincaid Grade Level, and average sentence length computed over the template's prose (variable slots and Jinja2 tags are stripped first so they don't skew word/sentence stats)
- Composite prompt-health report: a single
prompt-lens reportcommand combines token count, injection risk, instruction conflicts, and readability into a 0-100 "prompt health" score, emits JSON or a human-readable summary, and exits with a small CI-friendly status code (0healthy,1warnings,2critical) - Accepts
.txt,.md,.jinja2,.j2prompt files - Clean Python package with
src/layout andpyproject.toml
Requires Python 3.9+.
# Clone the repository
git clone https://github.com/ahnaf-lab/prompt-lens.git
cd prompt-lens
# Install in editable mode with dev dependencies
pip install -e ".[dev]"Instruction-conflict detection depends on spaCy's small English pipeline (en_core_web_sm), which installs automatically as a pinned wheel dependency — no extra download step or network access needed at analysis time.
# Default: count tokens for gpt-4 and list variables
prompt-lens analyse path/to/prompt.txt
# Count tokens for a specific model
prompt-lens analyse path/to/prompt.txt --model gpt-4o
# Show token counts for all supported models
prompt-lens analyse path/to/prompt.txt --all-models
# Skip variable extraction
prompt-lens analyse path/to/prompt.txt --no-variables
# Skip injection-risk scoring
prompt-lens analyse path/to/prompt.txt --no-injection
# Skip instruction-conflict detection
prompt-lens analyse path/to/prompt.txt --no-conflicts
# Skip readability scoring
prompt-lens analyse path/to/prompt.txt --no-readability# Human-readable summary, exits 0/1/2 for CI gating
prompt-lens report path/to/prompt.txt
# Machine-readable JSON report (same schema, same exit codes)
prompt-lens report path/to/prompt.txt --jsonprompt-lens report runs every check (tokens, variables, injection risk,
instruction conflicts, readability) and reduces the findings to a single
0-100 prompt health score, plus an exit code meant for CI pipelines:
| Exit code | Meaning |
|---|---|
0 |
Healthy — no injection errors, health score ≥ 90 |
1 |
Warnings — no injection errors, but health score < 90 (conflicts and/or minor injection warnings present) |
2 |
Critical — at least one injection error, or health score < 50 |
Example CI usage (fails the build on critical findings, but not on warnings):
prompt-lens report prompts/system_prompt.txt --json > report.json
if [ $? -eq 2 ]; then
echo "prompt-lens: critical issues found"; exit 1
fiExample output:
File: examples/system_prompt.txt
Tokens (gpt-4): 312
Variables found: ['assistant_name', 'context', 'domain', 'user_name']
Jinja2 blocks: 3
Injection risk: 1 error(s), 0 warning(s)
[ERROR] line 6 (PL-INJ001) variable 'context': User-controlled variable 'context' is interpolated directly inside a system/instruction block without any delimiter, so injected text could be read as an instruction.
fix: Wrap {{ context }} in an explicit delimiter (e.g. <user_input>{{ context }}</user_input> or a fenced code block) and/or move it under a clearly labeled 'User:' section.
prompt-lens list-modelsfrom prompt_lens import (
count_tokens, count_tokens_all_models, extract_variables,
scan_injection_risks, detect_conflicts, analyze_readability, build_report,
)
text = open("my_prompt.txt").read()
# Token counts
print(count_tokens(text, model="gpt-4")) # single model
print(count_tokens_all_models(text)) # all models dict
# Variable extraction
result = extract_variables(text)
print(result.all_variables) # ['context', 'name', 'role']
print(result.double_curly) # variables from {{ }}
print(result.single_curly) # variables from { }
print(result.jinja2_blocks) # raw content of {% %} tags
# Injection-risk scoring
for flag in scan_injection_risks(text):
print(flag.rule_id, flag.severity, flag.line, flag.variable, flag.suggestion)
# Instruction-conflict detection
for flag in detect_conflicts(text):
print(flag.rule_id, flag.topic, flag.line_a, flag.line_b, flag.reason)
# Readability scoring
readability = analyze_readability(text)
print(readability.flesch_reading_ease, readability.flesch_kincaid_grade)
# Composite health report (same dict the CLI --json output emits)
report = build_report(text, file="my_prompt.txt")
print(report["health_score"], report["exit_code"])| Rule | Severity | Trigger |
|---|---|---|
PL-INJ001 |
error | User-controlled variable interpolated inside a system/instruction block without any delimiter (fenced code block, triple-quoted block, or matching XML tag pair) |
PL-INJ002 |
error (system block) / warn | User-controlled variable uses raw format-string {var} syntax instead of {{ var }}, which has no escaping |
PL-INJ003 |
warn | User-controlled variable sits on the same line as privileged/override language (e.g. "ignore previous instructions", "system prompt", "secret") |
A template's whole content is treated as an implicit system block unless it declares explicit role headers (System:, Instructions:, User:, etc.) on their own line.
Prompts that grow organically often accumulate contradictory rules (e.g. one section says "always respond in English", another says "reply in the user's language"). prompt-lens parses imperative sentences out of the template using spaCy sentence segmentation plus a regex/POS-based imperative filter, clusters them by topic keyword (language, verbosity, tone, format, emoji, disclaimer, code, greeting, citation), and flags pairs that share a topic but contradict:
| Rule | Severity | Trigger |
|---|---|---|
PL-CONF001 |
warn | Two imperative sentences share a topic keyword cluster but assert contradictory constraints — either opposite polarity ("never use emojis" vs. "feel free to use emojis") or incompatible fixed values for a topic known to take mutually-exclusive values, i.e. language, format, verbosity, tone ("respond in JSON" vs. "write plain prose") |
This is heuristic, not a proof of contradiction: it can miss paraphrased conflicts outside its keyword coverage, and can occasionally flag related-but-compatible instructions. Treat findings as a prompt for human review.
python -m pytest tests/ -v| Model | Encoding |
|---|---|
| gpt-3.5-turbo | cl100k_base |
| gpt-4 | cl100k_base |
| gpt-4o | o200k_base |
| text-davinci-003 | p50k_base |
prompt-lens/
├── pyproject.toml
├── README.md
├── src/
│ └── prompt_lens/
│ ├── __init__.py
│ ├── tokenizer.py # tiktoken-based token counting
│ ├── extractor.py # variable slot extraction
│ ├── injection.py # heuristic injection-risk scoring
│ ├── conflicts.py # instruction-conflict detection (spaCy + regex)
│ ├── readability.py # Flesch readability scoring
│ ├── report.py # composite health score + CI report/exit codes
│ └── cli.py # click CLI entry point
└── tests/
├── fixtures/ # prompt fixture files covering edge cases
│ ├── conflicts/ # known-conflicting prompt pairs
│ └── library/ # small prompt-library fixture for the report e2e test
├── test_tokenizer.py
├── test_extractor.py
├── test_injection.py
├── test_conflicts.py
├── test_readability.py
└── test_report.py
This project is built autonomously and gated on passing tests. Token counting, variable extraction, heuristic injection-risk scoring, instruction-conflict detection, readability scoring, and the composite CI health report are implemented and verified by a pytest suite covering tokenisation, extraction, injection rules (PL-INJ001–003), instruction-conflict rules (PL-CONF001), Flesch readability metrics, and the report command's JSON schema and exit codes — including an end-to-end test that runs prompt-lens report --json over every file in a small prompt-library fixture directory (tests/fixtures/library/) and validates the resulting report against its schema.