Skip to content

Latest commit

 

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

prompt-lens

A CLI that statically analyses prompt template files and reports token budget, variable injection risk, instruction conflicts, and structural metadata. Targets prompt engineers and developers maintaining large prompt libraries.

Features

  • Token counting across multiple model families (gpt-3.5-turbo, gpt-4, gpt-4o, text-davinci-003) using tiktoken — fully offline, no API keys required
  • Variable extraction for {{double_curly}}, {single_curly}, and Jinja2 {% block %} styles
  • Injection-risk scoring: heuristic rules flag unescaped, user-controlled variable slots that sit inside system/instruction blocks, raw format-string ({var}) interpolation, and slots placed next to privileged/override language — each flag includes a line number, severity (warn/error), and a suggested fix
  • Instruction-conflict detection: parses imperative sentences (regex + an offline spaCy model), clusters them by topic keyword (language, verbosity, tone, format, emoji use, disclaimers, etc.), and flags pairs that contradict each other — either by opposite polarity ("never use emojis" vs. "feel free to use emojis") or by pinning incompatible values for the same topic ("respond in English" vs. "reply in the user's language")
  • Readability scoring: Flesch Reading Ease, Flesch-Kincaid Grade Level, and average sentence length computed over the template's prose (variable slots and Jinja2 tags are stripped first so they don't skew word/sentence stats)
  • Composite prompt-health report: a single prompt-lens report command combines token count, injection risk, instruction conflicts, and readability into a 0-100 "prompt health" score, emits JSON or a human-readable summary, and exits with a small CI-friendly status code (0 healthy, 1 warnings, 2 critical)
  • Accepts .txt, .md, .jinja2, .j2 prompt files
  • Clean Python package with src/ layout and pyproject.toml

Install

Requires Python 3.9+.

# Clone the repository
git clone https://github.com/ahnaf-lab/prompt-lens.git
cd prompt-lens

# Install in editable mode with dev dependencies
pip install -e ".[dev]"

Instruction-conflict detection depends on spaCy's small English pipeline (en_core_web_sm), which installs automatically as a pinned wheel dependency — no extra download step or network access needed at analysis time.

Usage

Analyse a prompt file

# Default: count tokens for gpt-4 and list variables
prompt-lens analyse path/to/prompt.txt

# Count tokens for a specific model
prompt-lens analyse path/to/prompt.txt --model gpt-4o

# Show token counts for all supported models
prompt-lens analyse path/to/prompt.txt --all-models

# Skip variable extraction
prompt-lens analyse path/to/prompt.txt --no-variables

# Skip injection-risk scoring
prompt-lens analyse path/to/prompt.txt --no-injection

# Skip instruction-conflict detection
prompt-lens analyse path/to/prompt.txt --no-conflicts

# Skip readability scoring
prompt-lens analyse path/to/prompt.txt --no-readability

Composite health report (CI)

# Human-readable summary, exits 0/1/2 for CI gating
prompt-lens report path/to/prompt.txt

# Machine-readable JSON report (same schema, same exit codes)
prompt-lens report path/to/prompt.txt --json

prompt-lens report runs every check (tokens, variables, injection risk, instruction conflicts, readability) and reduces the findings to a single 0-100 prompt health score, plus an exit code meant for CI pipelines:

Exit code Meaning
0 Healthy — no injection errors, health score ≥ 90
1 Warnings — no injection errors, but health score < 90 (conflicts and/or minor injection warnings present)
2 Critical — at least one injection error, or health score < 50

Example CI usage (fails the build on critical findings, but not on warnings):

prompt-lens report prompts/system_prompt.txt --json > report.json
if [ $? -eq 2 ]; then
  echo "prompt-lens: critical issues found"; exit 1
fi

Example output:

File: examples/system_prompt.txt
Tokens (gpt-4): 312
Variables found: ['assistant_name', 'context', 'domain', 'user_name']
Jinja2 blocks: 3
Injection risk: 1 error(s), 0 warning(s)
  [ERROR] line 6 (PL-INJ001) variable 'context': User-controlled variable 'context' is interpolated directly inside a system/instruction block without any delimiter, so injected text could be read as an instruction.
    fix: Wrap {{ context }} in an explicit delimiter (e.g. <user_input>{{ context }}</user_input> or a fenced code block) and/or move it under a clearly labeled 'User:' section.

List supported models

prompt-lens list-models

Python API

from prompt_lens import (
    count_tokens, count_tokens_all_models, extract_variables,
    scan_injection_risks, detect_conflicts, analyze_readability, build_report,
)

text = open("my_prompt.txt").read()

# Token counts
print(count_tokens(text, model="gpt-4"))       # single model
print(count_tokens_all_models(text))            # all models dict

# Variable extraction
result = extract_variables(text)
print(result.all_variables)      # ['context', 'name', 'role']
print(result.double_curly)       # variables from {{ }}
print(result.single_curly)       # variables from { }
print(result.jinja2_blocks)      # raw content of {% %} tags

# Injection-risk scoring
for flag in scan_injection_risks(text):
    print(flag.rule_id, flag.severity, flag.line, flag.variable, flag.suggestion)

# Instruction-conflict detection
for flag in detect_conflicts(text):
    print(flag.rule_id, flag.topic, flag.line_a, flag.line_b, flag.reason)

# Readability scoring
readability = analyze_readability(text)
print(readability.flesch_reading_ease, readability.flesch_kincaid_grade)

# Composite health report (same dict the CLI --json output emits)
report = build_report(text, file="my_prompt.txt")
print(report["health_score"], report["exit_code"])

Injection-risk rules

Rule Severity Trigger
PL-INJ001 error User-controlled variable interpolated inside a system/instruction block without any delimiter (fenced code block, triple-quoted block, or matching XML tag pair)
PL-INJ002 error (system block) / warn User-controlled variable uses raw format-string {var} syntax instead of {{ var }}, which has no escaping
PL-INJ003 warn User-controlled variable sits on the same line as privileged/override language (e.g. "ignore previous instructions", "system prompt", "secret")

A template's whole content is treated as an implicit system block unless it declares explicit role headers (System:, Instructions:, User:, etc.) on their own line.

Instruction-conflict detection

Prompts that grow organically often accumulate contradictory rules (e.g. one section says "always respond in English", another says "reply in the user's language"). prompt-lens parses imperative sentences out of the template using spaCy sentence segmentation plus a regex/POS-based imperative filter, clusters them by topic keyword (language, verbosity, tone, format, emoji, disclaimer, code, greeting, citation), and flags pairs that share a topic but contradict:

Rule Severity Trigger
PL-CONF001 warn Two imperative sentences share a topic keyword cluster but assert contradictory constraints — either opposite polarity ("never use emojis" vs. "feel free to use emojis") or incompatible fixed values for a topic known to take mutually-exclusive values, i.e. language, format, verbosity, tone ("respond in JSON" vs. "write plain prose")

This is heuristic, not a proof of contradiction: it can miss paraphrased conflicts outside its keyword coverage, and can occasionally flag related-but-compatible instructions. Treat findings as a prompt for human review.

Run tests

python -m pytest tests/ -v

Supported Models

Model Encoding
gpt-3.5-turbo cl100k_base
gpt-4 cl100k_base
gpt-4o o200k_base
text-davinci-003 p50k_base

Project Structure

prompt-lens/
├── pyproject.toml
├── README.md
├── src/
│   └── prompt_lens/
│       ├── __init__.py
│       ├── tokenizer.py      # tiktoken-based token counting
│       ├── extractor.py      # variable slot extraction
│       ├── injection.py      # heuristic injection-risk scoring
│       ├── conflicts.py      # instruction-conflict detection (spaCy + regex)
│       ├── readability.py    # Flesch readability scoring
│       ├── report.py         # composite health score + CI report/exit codes
│       └── cli.py            # click CLI entry point
└── tests/
    ├── fixtures/             # prompt fixture files covering edge cases
    │   ├── conflicts/        # known-conflicting prompt pairs
    │   └── library/          # small prompt-library fixture for the report e2e test
    ├── test_tokenizer.py
    ├── test_extractor.py
    ├── test_injection.py
    ├── test_conflicts.py
    ├── test_readability.py
    └── test_report.py

Status

This project is built autonomously and gated on passing tests. Token counting, variable extraction, heuristic injection-risk scoring, instruction-conflict detection, readability scoring, and the composite CI health report are implemented and verified by a pytest suite covering tokenisation, extraction, injection rules (PL-INJ001–003), instruction-conflict rules (PL-CONF001), Flesch readability metrics, and the report command's JSON schema and exit codes — including an end-to-end test that runs prompt-lens report --json over every file in a small prompt-library fixture directory (tests/fixtures/library/) and validates the resulting report against its schema.

About

A CLI that statically analyses a prompt template file and reports token budget, variable injection risk (e.g…

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages