Skip to content

Repository files navigation

deepagents-jev

Choose relevant original context instead of generating a summary. A small, reversible context-selection middleware for LangChain Deep Agents, powered by TypeSafe JEV.

All 3 turns, including initial summary generation — versus summarization.

Final-output benchmark: JEV cut costs by 83.9% on GLM-5.3 and 91.8% on Opus 4.8. Both methods produced 12/12 correct artifacts per model, each passing 32 functional tests.

Measured on a constrained Python-generation task; includes cache and JEV fees. The percentages above include the initial summary cost. They are not savings against an already-cached full history.

Continuation only (turns 2–3) — versus cached full context.

With about 99.5% full-context cache hits, JEV reduced continuation costs by 68.0% for GLM and 70.9% for Opus. Reused summaries were cheaper than both. All three methods passed 8/8 artifacts per model.

These are the same recorded runs, with the first turn excluded for every method. Reused summaries were cheapest on these follow-up turns. JEV beat cached full context, but did not beat an already-created summary. A longer session with growing history and repeated summarization could have a different cost balance.

Inspect the results · Reproduce the experiment

JEV scores older conversation fragments against the current task. The middleware sends selected original messages to your chat model while keeping the complete history in the checkpoint. A fragment omitted now can return for a later question. It replaces Deep Agents' built-in summarization slot; no fork or monkey patch is required.

Read the experiment report · Configuration · Reproduce the experiments

Install

Python 3.11+ is required. Compatibility is currently pinned and tested against Deep Agents 0.7.16. The package has not been published to PyPI; install from Git:

pip install 'git+https://github.com/hiroshi75/deepagents-jev.git'

For development and the included experiment tools:

git clone git@github.com:hiroshi75/deepagents-jev.git
cd deepagents-jev
uv sync --locked --group dev

The lockfile uses public PyPI packages. No local Deep Agents checkout is needed.

Add it to an agent

Set JEV_API_KEY in your environment, or copy .jev.example to .jev and fill in your key. Configure your chat model's provider credentials separately. JEV selects context; your ordinary chat model still generates answers and calls tools.

from deepagents import create_deep_agent
from deepagents_jev import JevContextMiddleware, SelectionConfig
from langgraph.checkpoint.memory import InMemorySaver

# Use the LangChain chat model already configured in your application.
def create_agent(model):
    context = JevContextMiddleware(
        config=SelectionConfig(
            max_tokens=16_000,
            trigger_tokens=4_000,
            keep_recent=6,
            relevance_threshold=0.2,
        ),
    )
    return create_deep_agent(
        model=model,
        middleware=[context],
        checkpointer=InMemorySaver(),
    )

# agent = create_agent(your_chat_model)
# config = {"configurable": {"thread_id": "my-project"}}
# agent.invoke({"messages": [("user", "Investigate the authentication code.")]}, config)
# Reuse the thread_id for subsequent questions; ainvoke is also supported.

A complete model-selectable example is included:

uv run --no-sync python examples/agent.py \
  --model anthropic:claude-opus-4-8 --task 'Explain the package structure.'

This example makes paid JEV and chat-model calls and runs the normal Deep Agents tool loop. Install any additional LangChain provider integration your model needs. To exercise selection and fragment recovery with JEV alone:

uv run --no-sync python examples/live_selection.py

What stays intact

  • Full checkpoint history; only the next model request is filtered.
  • All user/system messages and the recent-message tail.
  • Tool calls paired with their results, in original order.
  • Opaque media, signed reasoning, incomplete tool transactions, and oversized groups.

The budget covers history messages, not system prompts, tool schemas, or output. If protected history alone is too large, the middleware raises ContextBudgetExceeded. Scoring failures raise JevError by default. The optional fail_open setting allows full-history fallback, which can exceed your budget. See configuration.

Eligible context text is sent to the TypeSafe API. The standard Deep Agents summarizer also archives originals; this project's difference is selection of original text for the next request, rather than generation of a replacement summary.

Measured cost and correctness

Four cases: two long-history sizes × two seeds, with three questions per case. Each answer must recover four exact values from synthetic release-ledger records buried in historical source-file reads. Costs include observed caching and JEV fees.

Model Method Total USD Correct fields Fully correct answers
GLM-5.3 Summarization $0.484313 48/48 12/12
GLM-5.3 JEV selection $0.059991 48/48 12/12
GLM-5.3 Full context $0.590291 48/48 12/12
Claude Opus 4.8 Summarization $2.907931 48/48 12/12
Claude Opus 4.8 JEV selection $0.130132 48/48 12/12
Claude Opus 4.8 Full context $3.878537 48/48 12/12

JEV cost 87.6% less for GLM and 95.5% less for Opus 4.8 than summarization across this three-turn workload, at the same measured exact-answer accuracy. Summarization was cheaper on the subsequent two turns after its initial cost. This is a sparse fact-retrieval benchmark, not a general coding-quality evaluation. The 16k summary trigger was deliberately earlier than the default context-window threshold, and tools could not reopen archived history. No invoice comparison, long-session savings guarantee, or statistical equivalence claim is implied.

Original datasets, questions, answers, raw API usage, and checksum manifests are included. Regenerate costs and independently grade answers without API calls:

uv run --no-sync python -m benchmarks.report --verify
uv run --no-sync python -m benchmarks.quality

For live reruns, see experiments/README.md.

Final executable output

A separate follow-up asks each model to generate a Python request-admission function from the same historical ledgers. We execute the final code against 32 independent-oracle checks per artifact, including boundary values, invalid inputs, rejection priority, exact metadata, and output types. Every test must pass for an artifact to count as correct. These are fresh API calls with their own costs, not a relabeling of the JSON-answer experiment above.

Model Method Total USD Passing artifacts Functional tests USD / passing artifact
GLM-5.3 Summarization $0.508680 12/12 384/384 $0.042390
GLM-5.3 JEV selection $0.082033 12/12 384/384 $0.006836
GLM-5.3 Full context $0.603380 12/12 384/384 $0.050282
Claude Opus 4.8 Summarization $3.037862 12/12 384/384 $0.253155
Claude Opus 4.8 JEV selection $0.250455 12/12 384/384 $0.020871
Claude Opus 4.8 Full context $3.975434 12/12 384/384 $0.331286

JEV cost 83.9% less for GLM and 91.8% less for Opus 4.8 than summarization on this executable-policy task. All three methods passed all artifacts: there was no observed loss of final-output correctness on these cases. This is a small, constrained programming task, not evidence of equal quality on arbitrary tasks. Tests within an artifact are correlated, and the four histories share a task schema; they are not independent real-world software projects.

Replay all 72 generated artifacts and 2,304 functional checks without API calls:

uv run --no-sync python -m benchmarks.end_to_end_report --verify

The saved code, expected/actual test outputs, raw API responses, and costs are included in experiments/recorded/end-to-end/ and final-artifacts.json.

Repository layout

src/deepagents_jev/       Installable middleware and JEV client
examples/                Minimal agent and live selection examples
tests/                   Offline client, middleware, graph, and evaluation tests
benchmarks/              Live runners, cost aggregation, final-output evaluators
docs/                    Configuration and design constraints
experiments/recorded/    Immutable GLM and Opus 4.8 datasets and API records
experiments/summary/     Published cost and correctness results
experiments/SHA256SUMS   Checksums of recorded evidence
.github/workflows/      Offline CI (no API keys)

Development

uv run --no-sync pytest -q
uv run --no-sync ruff check src tests examples benchmarks
uv run --no-sync ty check src
uv build

The distribution contains the middleware, not the large recorded experiment data. Clone this repository for the full experiment bundle. See CONTRIBUTING.md.

Inspiration and license

Inspired by @CompleteSkeptic's post and the memo Why yet another agent?. This repository tests an independent implementation of query-aware selection.

MIT. Dataset source excerpts retain LangChain's MIT notice; see NOTICE and LICENSES/DeepAgents-MIT.txt.

About

No description, website, or topics provided.

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages