Cut AI Agent Response Time by 70% & Save 80% on Tokens
The transparent proxy that makes your AI coding agents faster, cheaper, and smarter
Quick Start • Features • How It Works • Stats • Docs
jevXagent is a transparent proxy that supercharges your AI coding agents (Claude Code, Codex, Kilo, Cline) by routing prompts through Jev's blazing-fast decision engine before the main LLM generates a response.
Your AI agent wastes time and tokens doing the same classification work over and over:
- "Is this a DevOps question or a bug fix?"
- "Should I search the codebase or generate code?"
- "What's the user actually asking for?"
Every prompt burns 200-300 expensive output tokens just thinking about what to do.
jevXagent implements "Line J" - injecting Jev's 120ms decisions directly into the agent's context, so it sees what was already decided before generating:
Without jevXagent: 2,800ms, 250 tokens
User → Agent → [classify + think + generate] → Response
With jevXagent: 950ms, 50 tokens
User → Jev (120ms) → Line J → Agent → [generate only] → Response
Result: 3x faster, 80% fewer output tokens, same quality.
Updated weekly - see statistics/ folder
| Metric | Without jevXagent | With jevXagent | Improvement |
|---|---|---|---|
| Response Time | 2,400ms | 890ms | 2.7x faster |
| Output Tokens | 240 tokens | 48 tokens | 80% reduction |
| Cost per Request | $0.0072 | $0.0018 | 75% cheaper |
| Time to First Token | 850ms | 310ms | 63% faster |
- Zero-Config Proxy - One command to start, works with all your agents
- 4 Agent Support - Claude Code, Codex, Kilo Code, Cline (auto-detected)
- Line J Injection - Jev decisions as system context (the secret sauce)
- Fail-Open Design - Jev timeout? Agent still works perfectly
- Real-Time Metrics - See latency and token savings per request
- Live Console Display - Track every request's performance
- JSONL Statistics - Machine-readable logs for analysis
- Ctrl+S Snapshots - Instant summaries of cumulative stats
- Weekly Reports - Automated statistics visualization
- Streaming Support - SSE with token extraction
- Multi-Turn Aware - Conversation context preserved
- Media Bypass - Images/PDFs skip Jev automatically
- Secure by Default - No credential logging, localhost only
# Clone and install
git clone https://github.com/j1s4nn/jevXagent.git
cd jevXagent
pip install -e .python -m jevxagentAnswer a few prompts:
- Select your installed agent (Claude Code, Codex, etc.)
- Enter Jev API key
- Configure stats path
- Done! Config saved to
~/.jevxagent/config.json
Terminal 1 - Start the proxy:
python -m jevxagentYou'll see:
==================================================
jevXagent v0.2
==================================================
Mode: Project
Agent: claude-code
Jev: ON
Proxy: http://127.0.0.1:9099
Ctrl+S Save statistics snapshot
Ctrl+C Stop jevXagent
==================================================
Proxy running. Waiting for requests...
Terminal 2 - Use your agent normally:
# For Claude Code / Kilo
export ANTHROPIC_BASE_URL=http://127.0.0.1:9099
claude
# For Codex
export OPENAI_BASE_URL=http://127.0.0.1:9099
codexThat's it! Your agent now routes through jevXagent automatically.
Traditional Agent (Slow):
User: "restart the pod"
↓
Agent LLM: "Let me think... this is DevOps... kubernetes...
I should generate a kubectl command... checking best
practices... here's the command with explanation..."
↓ 2,400ms, 240 tokens
Response: [Long explanation + code]
With jevXagent (Fast):
User: "restart the pod"
↓
Jev: Classifies → "DevOps, kubectl, urgent" [120ms]
↓
Line J: Injects context into agent request
↓
Agent LLM: *Sees decisions already made*
"kubectl rollout restart..."
↓ 890ms, 48 tokens
Response: [Concise code only]
Line J is the architectural innovation that makes this possible. Instead of making Jev's decisions available to the agent, we inject them directly into the agent's system context:
{
"messages": [
{
"role": "system",
"content": "SYSTEM EXECUTION CONTEXT:\nCategory: DevOps\nUrgency: High\nAction: Generate kubectl script only"
},
{
"role": "user",
"content": "restart the pod"
}
]
}Now the agent sees what Jev decided before generating, eliminating redundant classification and reasoning.
Every request shows detailed metrics:
──────────────────────────────────────────────────
Request #42 | Trace: a3f8b2c1
Jev: success | 142ms
Tokens: in=31 out=0
Agent: 1120ms
TTFT: 310ms
Tokens: in=1842 out=126 total=1968
Total: 1289ms
──────────────────────────────────────────────────
Press Ctrl+S anytime for cumulative statistics:
==================================================
STATISTICS SNAPSHOT
==================================================
Total Requests: 347
Jev Enabled: 347
Bypassed: 23
Avg Jev Latency: 128ms
Avg Total Latency: 1,056ms
Saved to: ./jevxagent_stats.jsonl
==================================================
Generate visual statistics reports saved to statistics/:
# Generate weekly report
python scripts/generate_stats_chart.pySee statistics/README.md for details.
- High-Volume Development - Save 80% on token costs across your team
- Performance-Critical Workflows - Cut response times from 3s to 1s
- Token Budget Management - Track and optimize AI spending
- Production Monitoring - Real-time visibility into agent performance
- Coding tasks with clear classification (bugs, features, refactors)
- DevOps automation prompts
- Repetitive agent workflows
- Multi-agent orchestration systems
Compare performance by toggling Jev:
# Edit ~/.jevxagent/config.json
{
"jev_enabled": false # Set to false to disable
}Restart the proxy. Now you can measure the exact improvement Jev provides.
Per-project statistics:
cd my-project
python -m jevxagent
# Stats saved to: my-project/jevxagent_stats.jsonl# Override config location
export JEVXAGENT_CONFIG=/path/to/config.json
# Change proxy port
export JEVXAGENT_PORT=9100- Quick Reference - Cheat sheet for daily use
- Usage Guide - Comprehensive walkthrough
- Implementation Report - Technical deep dive
- Architecture Diagrams - Visual architecture reference
Run the test suite:
pytest tests/ -vResult: 19/19 tests passing
Coverage includes:
- Agent adapter detection
- Line J context injection
- End-to-end request flows
- Fail-open error handling
- Statistics persistence
- Multi-turn conversations
- Core proxy architecture
- 4 agent adapters
- Line J context injection
- Real-time metrics
- Statistics persistence
- Automated weekly reports
- Web dashboard (live metrics)
- Model-specific pricing
- Docker containerization
- Multi-agent switching
- Prometheus metrics export
Contributions are welcome! Please see CONTRIBUTING.md for guidelines.
# Install dev dependencies
pip install -e ".[dev]"
# Run tests
pytest tests/ -v
# Run linting
black src/ tests/MIT License - see LICENSE for details.
- Jev by TypeSafe.ai - The fast decision engine powering jevXagent
- Anthropic, OpenAI - For the amazing agent APIs
- The open-source community - For the incredible tools we build upon
- Issues: GitHub Issues
- Discussions: GitHub Discussions
- Email: myprojectjisan@gmail.com
Speed up your AI agents today
Get Started • View Stats • Read Docs
Made with care by j1s4nn
Star us on GitHub if jevXagent helps you!
