Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
19 changes: 19 additions & 0 deletions .changeset/jev-evidence-coverage.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
---
"@sensegrep/core": minor
"@sensegrep/cli": minor
"@sensegrep/mcp": minor
---

Improve Jev search with query-sensitive scoring, helper-reserved candidate pools up to 80, per-aspect evidence coverage, bounded helper recovery, and source-free diagnostic traces. Preserve local fallback and separate local scores from model judgments. Add explicit query aspects and experimental AST excerpt/constant context options, plus offline grouped calibration and four-arm evaluation tools. Search/context use Jev automatically when a dedicated Jev credential is configured; explicit off and SENSEGREP_JEV_MODE take precedence, and a generic OpenRouter key alone does not enable it.

Remove concentration-based relevance discounting, add isolated Noul/Score panels and useful-mass ranking, stop redundant zero-gain context tails, and recover source-backed missing helpers/configuration. Add frozen candidate ablations, exact-contract probability calibration, offline question discovery with a configurable proposer (validated with GPT-6 Luna), and read-only end-to-end agent evaluations. No learned ranker or calibrated confidence is activated automatically.

Add independent rerank/evidence/recovery stages, advisory evidence categories, optional three-way literal-requirement verification and bounded multi-hop resolved-call recovery. Permit relevant delegating wrappers to trigger recovery without counting as answer evidence. Screen research questions across development groups, lint unsupported question types, retain score distributions, prune correlated features with CV checks, and feed the proposer out-of-fold errors and correct examples. Add real-source reviewed fixtures and stable-build stage ablations; keep learned artifacts experimental.

Require absolute eligible evidence for coverage selection. Retry isolated candidate state within the shared deadline when no eligible packet fits or helper recovery reaches its candidate cap without additions. Preserve the actual local selection when no usable packet is available; incomplete source cannot close an evidence gap.

Preserve global ranking when deduplicating overlapping chunks, preventing low-ranked file siblings from exhausting retrieval limits. Retain up to two already-retrieved, source-verified constant dependencies of leading implementations, including arithmetic definitions without evaluating them. Add context-outcome AutoResearch with grouped development selection, exact question-batch replay, revalidated seed features, local-score blend controls, resumable frozen datasets and separate held-out evaluation. Learned weights remain experimental.

Repair nonempty Jev context packets, retain uncertain resolved dependencies without claiming completeness, evaluate candidates in isolated states by default, and extend offline AutoResearch with contextual states and full Score distributions.

Replace runtime local-anchor repair with directed evidence packets, caller-conditioned dependency judgments and one bounded packet recovery cycle. Use a smaller staged panel and absolute eligibility buckets with stable local ties by default. Reserve final verification time, invalidate verdicts when packet membership changes, and recover omitted independent implementations through conditional contribution. Expose selection reasons and separate source-reviewed fact coverage, alternative implementations and exact symbol linkage in CLI evaluation. Historical panels/rankers remain explicit ablations; no learned weights or Luna calls are used in this validation.
7 changes: 5 additions & 2 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -38,8 +38,11 @@ session-*.md
# Local external repositories
bench/

# Temporary files
temp_audio.*
# Temporary files
temp_audio.*
__pycache__/
.jev-local/
.test-indexer/

# Demo generated artifacts
demo/out/
Expand Down
79 changes: 79 additions & 0 deletions docs/evaluations/jev-v12-acceptance-2026-09-25/REPORT.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,79 @@
# Jev v12: implementation and final acceptance

Date: 2026-09-25. Local implementation; no release or global installation performed.

## Changes

- Interleave local and vector candidates so truncation cannot erase the local reservation.
- Add bounded file discovery before full-source evaluation: up to 240 retrieved rows, 160 file catalogs, 12 excerpts per catalog, and six seconds within the overall deadline. Catalogs nominate sources; they never prove an answer.
- Separate domain compatibility from useful evidence, including useful negative answers. Unrelated service implementations cannot qualify solely through behavioral similarity.
- Select implementation packages within the real token budget. Preserve a strong alternative implementation and, when independently justified, a directly linked caller matching an explicit query term.
- Assess additional packages against already selected code rather than filling the package limit automatically.
- Preserve the actual local packet on fallback. Mark unresolved known dependencies as incomplete while excluding built-ins and callback parameters from false dependency obligations.
- Expose discovery and selection diagnostics; use a 30-second default deadline for package context. Explicit overrides remain supported.
- Extend the research budget only through recorded authorization, preserving previous spending and uncertain reservations.

These are bounded application rules and Jev judgments, not newly trained or calibrated weights.

## Final frozen run

Configuration: original 30-case manifest and reference bindings, 4,000 context tokens, maximum five packages, 32 full-package candidates, 30-second deadline. The discovery catalog limit is separate from the full-package candidate limit.

| Measure | Final result |
| --- | --- |
| Evaluated queries | 30 |
| Positive queries with all reference sources | 27 / 29 |
| Negative query | Correctly `no-evidence-found` |
| Mean query latency | 18,182.8 ms |
| Direct-evidence verdicts | 16 |
| Direct verdict missing reference sources (proxy) | 0 |

On the same 26 positive IDs completed by v11, complete source coverage increased from 21/26 to 25/26. The final version's 27/29 must not be compared directly with 21/26 without this denominator adjustment.

Remaining misses:

- `working-break-premise`: expected `isWithinWorkingHours` absent; other real working-hours implementations selected. Different implementations handle breaks differently and the query does not specify one. Final verdict remained `partial-evidence`.
- `sms-nonjson`: only one of three required reference sources retained; final verdict `partial-evidence`. A targeted run had recovered all three, but that success did not repeat in final acceptance.

Both cases have oscillated across evaluations. No expected labels were changed to improve the score. The earlier v12-confirmed result of 28/29 and the targeted SMS success are development evidence, not substitutes for the final 27/29 result.

Artifacts: [summary](summary.json), [package audit](package-audit.json), and split outputs `dev.json`, `validation.json`, `reserved.json`. Previous comparison: [v11](../jev-v11-final-2026-09-25/).

## Two fresh questions

Two questions were source-reviewed and registered before their inference, in [cases.json](../jev-v12-fresh-2026-09-25/cases.json). They were not used to tune this implementation.

- Asaas HTTP 200 with invalid JSON: reference source recovered, but verdict `partial-evidence` and Jev diagnostics `request-failed`.
- Integra ICP invalid JSON versus HTTP failure: reference source not recovered; verdict `not-assessed`, diagnostics `request-failed`.

The harness completed both rows, but model assessment did not complete cleanly. These are not two successful acceptance tests and do not establish generalization. See [raw results](../jev-v12-fresh-2026-09-25/dev.json). The available diagnostics do not establish the cause of `request-failed`; it must not automatically be attributed to the spending limit.

## Validation and limits

- Build and workspace TypeScript checks passed after the final runtime changes.
- Vitest: 391 tests in 55 files passed.
- Budget/evaluator Node tests: seven passed.
- Whitespace diff check passed.

The 30 main queries are reused regression cases, including those named validation/reserved. They are not an untouched statistical holdout. Reference-source coverage is not answer accuracy, and zero unsupported-direct proxies does not prove zero semantic false positives.

The fixes improve this regression set, but do not eliminate retrieval misses or inference variability. No claim of perfection, 29/29 final coverage, or statistically demonstrated generalization is warranted.

## Budget

Shared ledger: `../jev-v9-reviewed-2026-09-25/budget.json`.

- Conservative charge before this request: US$0.988151390568.
- Authorized cap: US$1.988151390568.
- Conservative charge after all evaluations: US$1.984672371332.
- Additional charge: **US$0.996520980764**, below the US$1 authorization.
- Remaining allowance: US$0.003479019236. No further paid calls planned.

The ledger includes uncertain reservations; these figures are conservative accounting, not a reconciled provider invoice. Spending was never reset.

## Documentation basis

- [TypeSafe state](https://docs.typesafe.ai/concepts/state): explicit, bounded judgment context.
- [Skill suggestion](https://docs.typesafe.ai/cookbooks/skill_suggestion): separate relative preference and absolute usefulness.
- [Classifying RAG passages](https://docs.typesafe.ai/cookbooks/classifying_rag_passages): focused evidence classification.
- [Jev limitations](https://docs.typesafe.ai/model-jaggedness/jev-1.13): typed output does not guarantee semantic correctness.
145 changes: 145 additions & 0 deletions docs/jev-context-repair-2026-09-24.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,145 @@
# Jev context repair — 2026-09-24

## Implemented behavior

- Nonempty remote packets can recover bounded local anchors and verified complements.
Token limits, result limits and explicit file/symbol diversity remain enforced.
Retained uncertainty is not counted as established aspect support.
- Contribution and packet completeness are separate questions. A high score for one
aspect or one implementation alone cannot satisfy the direct-evidence gate.
Completeness is still a fallible model judgment, not an exhaustive code proof.
- Bounded one-hop discovery starts from local anchors before semantic admission.
Existing candidates acquire call relations without replacing source or local rank.
Caller source is supplied separately with a 4,000-character bound/truncation flag.
- Recovery distinguishes already retrieved from already evaluated candidates, merges
repeat identities, and retains uncertain resolved helpers without calling them direct.
- The judging shortlist reserves more verified dependencies. Default flexible file
diversity no longer eliminates complements before coverage selection. Explicit
caps remain strict, including after packet repair.
- One candidate per state is the new default; up to four requests run concurrently.
Multiple questions about that candidate share its state. Explicit larger batches
remain available. Partial aspect coverage can trigger isolated re-evaluation when
batching was explicitly requested.
- Referenced constants retain dependency identities even when their scores need no
promotion. Missing constants can therefore participate in context repair.

## AutoResearch and calibration

`--contextual` measures a candidate against up to three short complete symbols from
its frozen local packet, excluding itself. This is a reproducible contextual
baseline, not a dynamically re-evaluated greedy selection at every insertion.
`--encoding distribution` supports all Score probabilities; the accepted feature in
this run was a Noul, so this run does not demonstrate a benefit from Score encoding.
Exact question batches, selected state, model and encoding remain part of replay.
A changed seed state/encoding is rejected. These selectors remain offline.

The first contextual run rejected oversized source: the old dataset included a
105,463-character symbol marked complete. It did not silently accept truncated
feature measurements. The next run excluded 36 candidates longer than 12,000
characters (380 -> 344). Required labels were not used for filtering. Original
baselines are retained, and sameSelectorLocal is the comparable filtered-pool control.
`prepare-jev-context-measurement.py` makes this preprocessing reproducible.
The failed attempt incurred some calls; its cost is not included in the successful
run's usage total. No successful artifact was created for that failed run.

Two discovery rounds with GPT-6 Luna accepted `direct_requested_behavior`.
Grouped development CV completed 19/19 packets versus 15/19 for the same structural
selector with local scores. The blend is 50% predictor / 50% local. This is not a
like-for-like increase from the prior 18/19 because the candidate pool and state
changed. Approximately 68.5% of selected candidates remain unreviewed.
Successful discovery recorded 696 fresh requests and $0.055842 provider cost.

Frozen evaluation on two analytics-privacy questions (one new family) completed
1/2 with learned weights and 1/2 with the local structural control. Learned nDCG
was 0.5000 versus 0.6186 for that control. The family was excluded from discovery;
this tiny check cannot establish generalization. The initial CLI check for this
family used the intermediate v2 build and is retained as `reserved.json`, not
presented as a final-build validation. All previously used families are regression
data, not new holdout data.

An experimental isotonic calibration of the accepted Noul used 24 reviewed DEV
examples for fitting and eight SMS-family examples for validation. Brier improved
from 0.017225 to 0.002551. This calibrates that particular feature/state contract,
not every runtime evidence score. The sample is too small for production threshold
promotion; the artifact remains `deploy:false`. Runtime 0.2/0.65/0.8 boundaries are
still heuristic policy thresholds. No claim of domain-wide calibrated probabilities.

## Runtime and agent validation

See `evaluations/jev-context-repair-2026-09-24/runtime-verified.json` for the final
build comparison: 12 queries across local and Jev (24 commands), frozen index and
compiled-JavaScript fingerprint checked before/after. Earlier runs are retained.
Some earlier runs overlapped remote research/agent calls; their latencies are not
controlled measurements. The final CLI run is executed without other remote work.

The agent benchmark (`agent.json`) used GPT-6 Luna for six tasks, each with and
without Jev, one repetition. Exact JSON comparison yielded 5/6 in both arms.
The local miss was capitalization of AES-256-GCM; the Jev miss added a correct
32-byte explanation to the expected `base64` string. Manual inspection finds no
factual error in those two answers. Do not report either as a retrieval failure or
claim a Jev accuracy win. This benchmark preceded the final diversity/parent-retention
hardening; its exact artifacts are preserved rather than relabeled as another build.

No learned weights, calibrated thresholds or deep traversal were promoted. No
commit, push, npm publication or global CLI update was performed in this task.

## Validation

- 334 Vitest tests in 52 files passed with four workers.
- 26 Python tests (18 context research, five discovery, three calibration) passed.
- Five Node evaluator tests passed.
- `npm run check` (build plus all workspace typechecks) passed.
- `git diff --check` passed.

An earlier unrestricted parallel test run hit the existing 750ms helper-analysis
wall deadline under load. The bounded-worker full suite passed; the runtime deadline
was not increased merely to hide the test timing issue.

## Reproduction

```sh
python scripts/prepare-jev-context-measurement.py <linked-dev.json> <measurable-dev.json>
python scripts/autoresearch-jev-context.py train <measurable-dev.json> <new-model.json> --live --proposals scripts/fixtures/jev-manual-features.json --contextual --encoding distribution --rounds 1 --max-cost 2 --max-requests 2600
python scripts/prepare-jev-context-measurement.py <linked-test.json> <measurable-test.json>
python scripts/autoresearch-jev-context.py test <measurable-test.json> <new-heldout.json> --live --artifact <new-model.json> --max-cost 1
```

References: [RAG classification](https://docs.typesafe.ai/cookbooks/classifying_rag_passages),
[reranking](https://docs.typesafe.ai/cookbooks/rerank_typesafe),
[AutoResearch](https://docs.typesafe.ai/cookbooks/autoresearch_feature_discovery),
[Jev limitations](https://docs.typesafe.ai/model-jaggedness/jev-1.13).


## Direct engineering review (no Luna)

After the user's workflow change, no additional proposer or agent-Luna calls were
made. Earlier research/agent runs are historical artifacts, not the basis for
claiming the following implementation is correct.

Direct control-flow inspection found a second semantic selection after packet
repair: `selectWithinTokenBudget` applies a relative-score admission gate and
reranks candidates. Calling it after dependency repair discarded low-scoring
helpers that had just been restored. The finalization now applies only explicit
diversity caps and recomputes token usage; it does not make another semantic
admission decision. Repair has already bounded result count and token use.

Recovery also sliced its candidate budget in vector-store enumeration order.
It now orders resolved children by the current caller frontier, then complete
source and bounded implementation length, with deterministic ties. Tests cover
storage-order interference and retrieved-but-unassessed dependencies. A packet
test separately proves that high contribution scores cannot override a low
completeness assessment.

`manual-review.json` is the frozen-build CLI comparison for this revision.
The other runtime reports remain available as earlier revisions. These queries
are regression cases; no fresh generalization claim is made.


The recommended future research command now uses manually authored question sets
(`--proposals`) and only Jev for measurement. It never calls a proposer model in
that mode; proposal JSON is validated and hashed into the output artifact. The
legacy explicit `--proposer-model` route remains available for historical replay,
but is not invoked by the recommended workflow. The manual feature set is not
activated in runtime and has not been represented as a measured quality gain.

Direct-review CLI result: 5/10 complete positive packets with Jev versus 4/10 local; both flagged 2/2 out-of-domain negatives. Median wall time: 6303.0 ms Jev versus 2840.0 ms local. One run per query; this is a regression comparison, not statistical generalization. Crypto helpers, one working-hours query and consent cases remain incomplete.
Loading
Loading