diff --git a/.agents/skills/augustus/SKILL.md b/.agents/skills/augustus/SKILL.md index 66487a6..7b11097 100644 --- a/.agents/skills/augustus/SKILL.md +++ b/.agents/skills/augustus/SKILL.md @@ -1,6 +1,6 @@ --- name: augustus -description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, openjev-lm, Nimble, encoder DeBERTa, LoRA distill), announced open decision-model (Watch), constrained-AR (TypeAR, pcdServer), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"eval path\", \"jevals\", \"Harbor taskset\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge, TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, or \"formally verify with Jev\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." +description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill, open multimodal RLCD / blackwood), announced open decision-model (Watch — still not landed), constrained-AR (TypeAR, pcdServer, decision-token LoRA), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"wait for Archer vs open multimodal RLCD\", \"screenshot/DOM candidates → typed Choice\", \"eval path\", \"jevals\", \"Harbor taskset\", \"shared bake-off ECE/NLL/Brier\", \"LLM-as-judge is not the System One score\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge (allowlist *proves*; fail-open cannot block), \"missing other → confident wrong Choice\", \"lint the Jev request\", \"training confronts Choice other / none-of-the-above\", \"S1 reflex keeps control / optional S2 one-use advice\", \"soft AGENTS.md rules vs linter\" (Abide / jev-pref), \"edit-phase vs turn-phase observation window\", banded confidence fail-open preference lint, \"extractive selection / pointer-not-generator\", \"encoder GLiNER compaction vs Jev Score compaction (same job; pointer not summarizer)\", \"fail-closed keep_full under mutation envelope\", \"CI flaky-vs-real merge gate\", \"fail-open VOI wake/resume (Horvitz)\", \"claim/evidence Stop integrity\", \"S1 extract + escalate-S2 indexer\", \"Harbor on/off routing\", \"policy-as-judgment PR marshal\", \"shadow-mode compaction rollout\", \"Jev Ultrafast vs GLiNER Ultrafast (observe-score-act backend-agnostic)\", \"hybrid local decide + remote fill\", \"DONE ≠ verified success\", \"observed a11y/DOM candidates vs screenshot multimodal\", \"evidence-preserving stdout prune (not summarize)\", \"hard token/format envelope then soft Noul\", \"fail-safe keep original on prune failure\", \"stdout prune vs session compaction\", \"specialist S1 computer-use (Cua-S1 form-v0; not TypeSafe Jev)\", \"plan ≠ execute / dry-run default\", \"observed-element option head (fill/check/click/skip)\", \"local /v1/systemone drop-in (stub until hf scorer)\", \"dataframe-native semantic columns\", \"route≠memory\", \"advisory sidecar receipts\", \"structure induction over bags\", \"AST ∩ semantic lint\", \"extractable-from-state / retrieve first\", \"decision-model vs constrained-LLM bake-off\", \"dual-process S1 decide / S2 generate\", \"combinatorial grid ≠ extractive\", TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, \"uncalibrated local likelihoods vs Noul / CUDA replica\", \"decision-native RAG retrieve wide then decide then evidence set\", \"classify-first MCP / read selectively\", \"living applied-mappings atlas / class patterns not a 342 hit list\", \"draft-gate silence as safer / heartbeat\", \"robotics text-state vs pixels\", \"verbatim session ledger / scored recall\", \"judgment as language primitive / English-as-config\", \"pre-registered AMBIGUOUS eval / cascade sign-flip\", \"healthcare Harbor-shaped S1+S2\", \"pre-exec tool gate allow/block/review\", \"productized public primitive / judgment wall\", \"meaning-search without embeddings\", \"attention≠correctness PR review\", \"skills→oxlint / AST prove ∩ remainder\", \"session-sticky first-prompt routing\", \"measured RAG rerank vs generative rerank\", \"Stagehand extract pick-and-copy / judge\", \"harness observe-score-act productization\", or \"formally verify with Jev\", \"capability kernel / secrets never in the agent\", \"Jev is SENSOR not policy\", \"type-safe ≠ correct\", \"typed control plane around DSPy\", \"native-probability calibration / Brier/ECE arena\", \"fan-out as measurement economics\", \"engine owns truth / Jev owns judgment\", \"human-confirmed kill gate\", \"train specialist when downstream reads p vs few-shot hosted when only argmax\", \"decide→policy→LLM leftover cascade\", \"Noul 0.5 cannot-tell never rounded\", \"calibration ≠ sortable / ORDER BY over Jev probs\", \"pairwise inversion / Score ordinality / two-decimal ties\", \"wire-compat self-hosted /v1/systemone GLiFormer\", \"class-backend economics\", \"loopback gateway hosted + local OpenJev\", \"do not distill Jev as teacher of record\", \"active-learning triage / training-data VOI\", \"index-once ask-many / citable evidence packets\", \"meaning-grep AND/OR/NOT line Nouls\", \"closed-vote-only computer-use / no planner LLM\", \"Jev vs local MLX PCD Harbor\", \"PCD O(1) speed ≠ calibrated Noul\", \"host-owned handlers × System One\", \"OMP/pi fail-open acceptance gate\", \"permission vs probability / operator owns the safety bar\", \"judgment ≠ permission / Jev never grants access\", \"eval integrity / instrument not score / dinostomp jev-as-if\", \"constrained optimizer + S1 features / never sole hot-path gate\", \"privilege ≠ verdict / effect contracts not tokens\", \"attention filter / VOI for human review / never blocks / never green unless sure\", \"measurement owns endorsement / evidence-gated question packs\", \"Jev supplies evidence / code owns authority\", \"ranking ≠ calibration / never hard-threshold raw p as frequency\", \"hot-click CU / indexed element table / S1 on click path\", \"Jev judges relevance / code decides structure / never rewrite\", \"local rules first then remainder / never auto-train on model's own hides\", \"combinators / System One as control plane / not chat turns\", \"receipts not leaderboard / type-safe ≠ correct jaggedness\", \"VOI over skill library / skillranker abstention\", \"OOD calibration / AUC ≠ ECE / sign of miscalibration by type\", \"Jev vs thinking-budget small models / frontier-100\", \"turnstile / replayable evidence≠authority\", \"MLX one-pass schema→JSON / Apple Silicon replica economics\", \"memory leases ended by new evidence\", \"never confidently wrong / TLA+ compose with judgment / escalate instead of hard-gate\", \"no seal no advance / coverage ledger / mint ≠ product brain\", \"skill-broker sibling turnstile/skillranker / judgment ≠ permission\", \"sureness / CERTAIN|CONFIDENT|LEANING|TORN|CLUELESS / max_prob is generous\", \"JevBench Harbor/jevals practice / calibration not in Main Score\", \"CI typed gate before expensive review / ci-gatekeeper\", \"Codex MCP host adapter / jev_select_capability\", \"judgment as attention redirect not merge blocker / jev-preflight\", \"compress-before-first-send / dizk jev-lens vs rashed attention filter\", \"tools≠use / SessionStart over hoping the model recalls\", \"observational memory / keep-kind verbatim / pi-om\", \"open-Jev class / openvons / JevPick menu decode\", \"physical-world System One / HA-Jev / not for locks\", \"judgment outside the store / jevql CLI\", \"landed-script trust / headless≠auto-approve\", \"digital-design combinators / extended Router Loop Retry Fallback Memory\", \"VOI cache admission / same-intent skip LLM\", \"BM25 vs Jev skill routing Harbor harness\", \"zeroshot vs BERT / contamination DiD / label-equivalence\", \"typed escalate continue abort baton / inverted loop\", \"worth-your-attention VOI / ThinkyMiner Winnow vs kevinpita winnow\", \"Jev WHETHER Python HOW LLM WHAT\", \"conflict vs ignorance / named Choice escape\", \"Playwright executes Jev chooses / sample-from-distribution\", \"OpenJev /v1/decide not TypeSafe drop-in\", \"SemIf wire-compat runoff\", \"decision-as-memory flywheel\", \"record/replay CI / jevassert\", \"failure-finding arena / jevarena ≠ jev-arena\", \"BBQ stereotype/uncertainty/cost\", \"decider≠executor / jeffrey\", \"sentence-as-rule lint / jevlint ≠ JevLint\", \"VOI hunk prune / prune-review\", \"whole-repo intent VERIFIED/VIOLATION/UNKNOWN\", \"GLiNER2 System One spec ≠ replica\", \"Rust/WebGPU grande / Clojure Laya byte parity / CPU SemIf\", \"ONNX ModernBERT local-jev measured not equivalent\", \"persist constraints across compaction / pi-heed\", \"calibration+cost as first-class gates\", \"Harbor-shaped Jev vs schema-guided LLM-as-judge / jev-judge-bench ≠ jevarena ≠ jevbench\", \"hand no-text steps to Jev / jev-use / Vercel drops confidence / margin fallback\", \"Pi System-One control plane / pi-jev-control\", \"generation as tree of Choices / never free-generates / jev-gpt\", \"OpenRouter recipe atlas / samples not benches / jev-cookbook\", \"personal history feed / no social graph / jevfeed\", \"competing NAR claims / dual-channel ECE / claim-verification / openJev-verdict ≠ OpenJev\", \"empty compaction-proxy skip / IPECTER\", \"throughput ≠ latency / like-for-like ECE\", \"1-token logprob endpoint ≠ Noul / coverage ≠ correctness / chakuho\", \"open replica engine / jevinf / argmax-parity ≠ ECE\", \"unofficial Elixir SDK ≠ OTP peer / dannote/jev\", \"jevex rename + n=16 SWE VOI / files-to-read\", \"commit pre-review attention≠verdict / middle band never rounded / commitjev\", \"Hermes plugin is Agnes not TypeSafe\", \"pi-jev-compact ≠ pi-jev-compaction / verbatim summarizer replacement\", \"empty Codex-proxy skip / IPECTER runway\", \"decision-native inbox / mailordinal / humans own ambiguity\", \"unofficial jev-cli not ready / ≠ jevql\", \"laya-multilingual / English checkpoint confident-wrong OOD / ships uncalibrated\", \"schema-conditioned DeBERTa scorer / peaked ranking ≠ calibration\", \"HF 401 access / GitHub 404 Hub-only\", \"productized System One HTTP / classifier.dev / label+confidence public contract\", \"escalate-under-threshold / smart tier 0.7 / multi-label ignores tier\", \"silent-fallback FALLBACK marker / granite 0.546 vs advertised 0.800\", \"vs_jev tracked JSON not transcription / read eval/README before quoting\", \"choxos/jev-reviewer ≠ egma-ai / systematic-review pointer-not-generator\", \"two-pass Choice+Noul / relative which-line + absolute does-this-line\", \"not-found is an answer / no paraphrase invent\", \"human check as productized judgment / checked never overwritten\", \"githubnext/localjev ≠ kunchenguid/local-jev / prompted JSON ≠ structured logit read\", \"wire-compat ≠ logit-equiv / self-reported probs / entropy confidence\", \"institutional open-replica / GitHub Next /v1/systemone\", \"Harbor-shaped bake-off AG News BoolQ SST-5 / 1200-request caveats\", \"LM Studio runner gap / structured-read primitives for OpenJev parity\", \"NandhaKishorM/laya packaging ≠ Hub-only / Router script-before-p\", \"post-T ECE ≠ raw ECE / Banking77 token-budget / 0.85 still soft / not TypeSafe drop-in / external census ≠ scored bake-off / GLiNER2+routers class-boundary / incomplete vs watch / Harbor honesty watch / JevBench v1.2 geometric-mean I/C/S/K / cal now ON rank / weight sensitivity / option-order 72→21 / instruction models class-boundary / ×2 latency assumption / est. costs / Laya absent gap / Qwen3.8 27B ≠ Archer\", \"hourly already-folded watch / apply-the-five / skip thin noise\", \"hard-gate Noul as PR gate is soundness theater / totally-tim/jev-gate ≠ jev-gateway\", \"S1 never stalls waiting / S2 one-use advisory\", \"purple telemetry = consumed not arrived\", \"Local controller ≠ githubnext/localjev\", \"seed = geometry not async replay\", \"20% starting gate still soft / schema-safe ≠ correct\", \"no pixels to either provider / confidence ≠ selected probability\", \"experimental viz not a flight controller / S2 never grants\", \"OCR+AX observe-score-act / typesafe-computer-use\", \"never send screenshot to frontier for the decision\", \"overlapping CU options = false low confidence\", \"split kind/item/site / offscreen\", \"writer/decider split + post-type Noul still soft\", \"155× one-screenshot Harbor-shaped ≠ taskset\", \"AX never sole / Spotify 0\", \"decision ≠ answer-reader capture\", \"typesafe-computer-use ≠ jev-ultrafast ≠ cua-s1 ≠ camoufox\", \"ASR observe-score-act / jev-voice-browser\", \"partial-speech VOI / complete Noul / free-text waits\", \"spoken confirm ≠ hard auth\", \"numbered overlay disambiguate without another model\", \"moritzkremb/jev-voice-browser ≠ jev-voice-control ≠ nikolas-j\", \"wrap-as-execution / AgentGhost ALLOW ASK DENY\", \"rules first then Jev remainder / ASK throws / fail-closed\", \"reddpy/AgentGhost ≠ jwen5419807/agentghost ≠ vventirozos\", \"JP genre atlas / studio_yebisu / stars ephemeral ≠ eval\", \"Jev Clearly Explained / akshay_pachaar / LLM hammer\", \"schema-safe ≠ correct / 200× 400× TypeSafe ceiling\", \"questions-as-code / shadow first / not a TypeSafe how-to\", \"proposition ≠ embedding / contrast-set refund\", \"boolean composition of soft Nouls / AND OR NOT after threshold\", \"uehaj/jev-semgrep ≠ semgrep.dev\", \"meaning-grep dedicated fold / not a gate\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." license: MIT metadata: version: 0.3.0 @@ -23,8 +23,8 @@ Noul), not the monopoly. This skill owns **where judgment belongs**; the official `typesafe-ai` skill plus the live docs own Jev integration contracts — read them before writing Jev API code. Neighbor skills `tenbin` (lint/measure) and `decision-first` (try-Jev-first habit) own -their jobs. Do not collapse into a TypeSafe how-to, a Laya install, or -a GLiClass or GLiNER tutorial. +their jobs. Do not collapse into a TypeSafe how-to, a Laya install, a kev +serve, a blackwood vLLM how-to, a jev-local Docker install, or a GLiClass or GLiNER tutorial. Pick the **pillar** from the hole (expected utility, VOI, MCDA, signal detection, search/control, org/safety, formal methods), then the @@ -51,7 +51,10 @@ classical method you already trust, substitute it, classify the win this?", "is Jev the only model?", "is this only for software?", "formally verify with Jev / replace TLA+ / Dafny / DST", "Alloy vs Apalache", "GLiNER vs Jev", "LLM-as-judge", - "paraphrase brittleness", "allowlist then judge", or "TOCTOU-of-Noul": + "paraphrase brittleness", "allowlist then judge", "TOCTOU-of-Noul", + "Jev inside the database / sqlite-jev", "Jev picks bitrate / join + order / the model", "wait for Archer", "lint the request / missing + other", "training confronts Choice other / none-of-the-above", "soft AGENTS.md rules vs the linter", "screenshot Choice / omni System One", "extractive quotes / pointer not generator", "compaction summarize vs pointer", "encoder vs Jev compaction backend", "shadow-mode compaction rollout", "CI flaky-vs-real merge gate", "fail-open VOI wake/resume", "claim vs session evidence", "S1 indexer escalate-S2", "Harbor on/off routing", "fail-open vs fail-closed wake vs CI gate", "encoder vs Jev computer-use backend", "hybrid local decide + remote fill", "DONE vs verified success", "stdout prune vs session compaction", "hard envelope then Noul prune", "Cua-S1 vs TypeSafe Jev", "plan vs execute dry-run", "specialist computer-use vs general agent", "local drop-in vs stub scorer", "route vs memory", "when does it hold / extractable from state", "decision model vs constrained LLM", "dual-process S1/S2", "combinatorial grid vs extractive", "uncalibrated local likelihoods", "decision-native RAG", "classify-first / read selectively", "living applied-mappings atlas / class patterns", "silence as safer / draft-gate heartbeat", "robotics text-state vs pixels", "verbatim ledger vs summary", "judgment as language primitive", "Stagehand extract pick-and-copy", "harness observe-score-act vs demo loop", "public judgment wall / six parallel questions", "meaning-search without embeddings", "attention ≠ correctness", "skills→oxlint / AST prove ∩ remainder", "session-sticky first-prompt routing", "measured RAG rerank vs generative rerank", "capability kernel / secrets never in the agent", "Jev is SENSOR not policy", "type-safe ≠ correct", "typed control plane around DSPy", "native vs verbalized confidence", "engine owns truth / Jev owns judgment", "human-confirmed kill gate", "train specialist vs few-shot hosted", "decide→policy→LLM leftover", "Noul 0.5 cannot-tell never rounded", "calibration ≠ sortable / ORDER BY", "pairwise inversion / Score ordinality / two-decimal ties", "wire-compat GLiFormer /v1/systemone", "class-backend economics", "loopback gateway hosted + local", "do not distill Jev as teacher", "active-learning triage", "evidence-packet explorer", "meaning-grep AND/OR/NOT", "closed-vote-only / no planner LLM", "Jev vs PCD Harbor", "PCD O(1) ≠ Noul", "host-owned handlers × System One", "OMP/pi fail-open gate", "permission vs probability / operator owns thresholds", "judgment ≠ permission / Jev never grants access", "eval integrity / instrument not score", "constrained optimizer + S1 features / never sole hot-path gate", "privilege ≠ verdict / effect contracts not tokens", "attention filter / VOI for human review / never blocks / never green unless sure", "measurement owns endorsement / evidence-gated question packs", "Jev supplies evidence / code owns authority", "ranking ≠ calibration / never hard-threshold raw p as frequency", "hot-click CU / indexed element table", "Jev judges relevance / code decides structure", "local rules first then remainder / never auto-train on own hides", "combinators / System One as control plane", "receipts not leaderboard / type-safe ≠ correct jaggedness", "VOI over skill library / skillranker abstention", "OOD calibration / AUC ≠ ECE", "Jev vs thinking-budget small models", "turnstile / replayable evidence≠authority", "MLX one-pass schema→JSON / Apple Silicon replica economics", "memory leases ended by new evidence", "never confidently wrong / TLA+ compose / escalate instead of hard-gate", "no seal no advance / coverage ledger / mint ≠ product brain", "skill-broker sibling / judgment ≠ permission", "sureness bands / max_prob is generous", "JevBench / calibration not in Main Score", "CI typed gate before expensive review", "Codex MCP host adapter", "judgment as attention redirect / jev-preflight", "compress-before-first-send / dizk jev-lens", "tools≠use / SessionStart over hoping", "observational memory / pi-om keep-kind", "open-Jev class / openvons / JevPick", "physical-world System One / HA-Jev / not for locks", "judgment outside the store / jevql", "landed-script trust / headless≠auto-approve", "digital-design combinators / extended five", "VOI cache admission / same-intent skip LLM", "BM25 vs Jev skill routing Harbor harness", "zeroshot vs BERT / contamination DiD", "typed escalate continue abort baton / inverted loop", "worth-your-attention VOI / ThinkyMiner Winnow", "Jev WHETHER Python HOW LLM WHAT", "conflict vs ignorance / named Choice escape", "Playwright executes Jev chooses", "OpenJev /v1/decide not drop-in", "SemIf wire-compat runoff", "decision-as-memory flywheel", "record/replay CI / jevassert", "failure-finding arena / jevarena ≠ jev-arena", "BBQ not a bias cert", "decider≠executor", "sentence-as-rule lint / jevlint", "VOI hunk prune", "whole-repo intent VERIFIED/VIOLATION/UNKNOWN", "GLiNER2 spec ≠ replica", "open replica substrates / grande / laya-jolt / JEV-CPU", "ONNX local-jev not equivalent", "persist constraints across compaction / pi-heed", "calibration+cost first-class gates", "Harbor-shaped Jev vs SGR LLM-as-judge / jev-judge-bench ≠ jevarena ≠ jevbench", "hand no-text steps / jev-use / Vercel drops confidence", "Pi System-One control plane / pi-jev-control", "never free-generates / jev-gpt tree of Choices", "OpenRouter recipe atlas / samples not benches", "personal history feed / jevfeed / no social graph", "competing NAR claims / dual-channel ECE / openJev-verdict ≠ OpenJev", "empty compaction-proxy skip / IPECTER", "throughput ≠ latency / like-for-like ECE", "1-token logprob endpoint ≠ Noul / coverage ≠ correctness", "open replica engine / jevinf", "unofficial Elixir SDK ≠ OTP peer", "jevex n=16 files-to-read VOI", "commit pre-review attention≠verdict / middle band", "Hermes plugin is Agnes not TypeSafe", "pi-jev-compact ≠ pi-jev-compaction", "empty Codex-proxy skip / IPECTER runway", "decision-native inbox / mailordinal", "unofficial jev-cli not ready / ≠ jevql", "laya-multilingual / English checkpoint confident-wrong OOD", "schema-scorer peaked ranking ≠ calibration", "HF 401 / GitHub 404 Hub-only", "productized System One HTTP / classifier.dev", "escalate-under-threshold / smart tier / multi-label ignores", "silent FALLBACK / granite 0.546 vs advertised 0.800", "vs_jev tracked JSON / read eval/README", "choxos/jev-reviewer ≠ egma-ai / systematic-review pointer", "two-pass Choice+Noul evidence extraction", "not-found is an answer", "human check as productized judgment", "githubnext/localjev ≠ kunchenguid/local-jev", "wire-compat ≠ logit-equiv / prompted JSON ≠ structured read", "self-reported probs / entropy confidence", "GitHub Next local /v1/systemone", "LM Studio runner gap / structured-read primitives", "NandhaKishorM/laya packaging ≠ Hub-only / Router script-before-p", "post-T ECE ≠ raw ECE / Banking77 token-budget", "0.85 still soft / not TypeSafe drop-in", "external census ≠ scored bake-off", "GLiNER2+routers class-boundary", "incomplete openjev census vs watch", "Harbor honesty watch / silent fallback", "JevBench v1.2 geometric mean / cal ON rank / weight sensitivity", "option-order 72→21 / instruction models in the class table", "self-host latency ×2 assumption / est. costs", "Laya absent is a gap not a named exclusion", "Qwen3.8 27B ≠ Archer", "hourly already-folded watch / apply-the-five / skip thin noise", "hard-gate Noul as PR/quality gate is soundness theater", "S1 never stalls waiting / S2 one-use advisory", "Local controller ≠ githubnext/localjev", "purple telemetry = consumed not arrived", "seed = geometry not async replay", "20% starting gate still soft", "no pixels to either provider", "OCR+AX observe-score-act / typesafe-computer-use", "never send screenshot to frontier for the decision", "overlapping CU options = false low confidence", "split kind/item/site", "155× one-screenshot ≠ Harbor taskset", "decision ≠ answer-reader capture", "ASR observe-score-act / jev-voice-browser", "partial-speech VOI / free-text waits", "spoken confirm ≠ hard auth", "numbered overlay without another model", "wrap-as-execution / AgentGhost ALLOW ASK DENY", "rules first then Jev remainder / ASK throws / fail-closed", "reddpy/AgentGhost ≠ jwen5419807/agentghost ≠ vventirozos", "JP genre atlas / studio_yebisu / stars ephemeral ≠ eval", "Jev Clearly Explained / akshay_pachaar / LLM hammer", "schema-safe ≠ correct / 200× 400× TypeSafe ceiling", "questions-as-code / shadow first / not a TypeSafe how-to", "proposition ≠ embedding / contrast-set", "boolean composition of soft Nouls / AND OR NOT", "uehaj/jev-semgrep ≠ semgrep.dev", "meaning-grep dedicated fold / not a gate", or "cascade sign-flip / calibration theater": read `references/faq.md`, then `references/mental-models.md`, then `references/mixed-architecture.md`, then @@ -81,7 +84,9 @@ classical method you already trust, substitute it, classify the win person; a judgment-class model may gate, route, or verify around that call. Proof, model-checking, contracts, and DST stay with their tools — a Noul is a sensor, not a discharged proof obligation - (`references/formal-methods.md`). + (`references/formal-methods.md`). A capability kernel keeps + secrets out of the agent and leaves BLOCK/ASK/ALLOW in policy; + type-safe is not the same as correct. 4. Give the provider one narrow judgment per question (a knowledgeable person could answer in a second given the state). Split multi-factor judgments; fuse in code with visible weights. Jev's option/envelope @@ -90,7 +95,8 @@ classical method you already trust, substitute it, classify the win only where the answer *is* a span in the text (`judgment-class.md` species map). 5. Exploit the family's cheap fan-out: batch independent questions in - one request when the provider supports it; encode text + all labels + one request when the provider supports it (measurement economics: + 200 calibration questions in 2 requests; ~100 Battleship noul/turn); encode text + all labels once for GLi\* heads; score many prompts against one embedding for dual-encoder vision. Sequence a second call only when its state or options depend on an earlier answer. @@ -110,47 +116,52 @@ classical method you already trust, substitute it, classify the win | Familiar method | Judgment shape | Detail | |---|---|---| -| Mental models across domains (not SWE-only) | EU, abstention, VOI, MCDA, SDT, search/control, Leveson, NATM/Norman/snap-fit | `references/mental-models.md` | -| Judgment-model class (Jev is exemplar, not monopoly) | Species: decide / locate (GLiNER) / categorize (GLiClass) / rank / perceive; open heads include encoder DeBERTa, LoRA distill, openjev-lm | `references/judgment-class.md` | -| Open weights vs constrained decoding vs proprietary API | Three open paths: encoder open-jev / AR constrained decode (TypeAR + pcdServer) / trained decision-only (Laya, Nimble, Archer Watch). Softmax over allowed tokens ≠ Noul | `references/judgment-class.md` (when-to-use table) | +| Mental models across domains (not SWE-only) | EU, abstention, VOI, MCDA, SDT, search/control, Leveson, NATM/Norman/snap-fit; **extractable-from-state boundary map** (self-contained vs needs outside knowledge) | `references/mental-models.md` | +| Judgment-model class (Jev is exemplar, not monopoly) | Species: decide / locate (GLiNER) / categorize (GLiClass) / rank / perceive; open heads include encoder DeBERTa, LoRA distill, **domain specialist LoRA on independent gold**, openjev-lm, kev. Compaction job is backend-agnostic (Jev Score/Noul vs GLiNER2.5 encoder). Indexer cousin: GLiNER extract + escalate-S2 (10–50× unfilled). Computer-use observe→score-among-candidates→code-acts is backend-agnostic (Jev Ultrafast ↔ GLiNER2 Ultrafast ↔ Cua-S1 specialist ↔ Stagehand experimental Jev stack; **OCR+AX desktop:** typesafe-computer-use hosted Jev, never ships a screenshot for the *decision*; Cua-S1 is not TypeSafe Jev; Stagehand pick is a fast path, not a replacement; **≠** jev-macos-loop OmniParser **≠** camoufox; **ASR voice-browser:** jev-voice-browser hosted Jev, never ships a waveform). GLiFormer encoder serving `/v1/systemone` is a class-backend (jeff), not a Jev replica. Local MLX PCD is O(1) constrained-AR speed, **not** a calibrated Noul (system-one-benchmark Brier 0.3884 vs Jev 0.1096). **jevmlx** is the productized Apple Silicon one-pass schema→JSON+probs library (softmax ≠ Noul; no local leaderboard yet). **openvons** is an independent open-Jev class (LM/vision/voice finite-choice+prob; JevPick 3.2–4.8×; Flutter on-device; `/v1/systemone` wire-compat, not a TypeSafe replica). **OpenJev** (IamBusy) is a local 0.6B LoRA+scalar head on `/v1/decide` (45/60 *theirs*; **not** a TypeSafe drop-in; distinct from hraness/sysone OpenJev runners). **semif-serve** puts SemIf behind `/v1/systemone` (runoff, no option ceiling; 1164 vs 178 ms *theirs*; wire-compat ≠ replica). **grande** Rust/WebGPU kev-shaped branches (JGLUE *theirs* JNLI 0.614 ECE→0.088 / JCQA 0.853; 270M 0.710/0.710; isolation 0.098/0.996). **laya-jolt** Clojure/Jolt byte parity vs Python Laya. **JEV-CPU** SemIf on CPU (leesk212; Meanblock 404). **local-jev** ONNX ModernBERT measured not-equivalent (done 30% / shape 57% *theirs*). **GLiNER2→Choice/Score/Noul spec** (Eran-BA; design only, ≠ jeff GLiFormer). **openJev-verdict-2.0** competing NAR claims as **audit object not endorsement** (77.10%/0.0636/0.0144 *theirs*; PR #1; ≠ IamBusy/OpenJev `/v1/decide`). **chakuho** 1-token logprob local `/v1/systemone` (softmax ≠ Noul; coverage ≠ correctness; GUI 336 *theirs* 27B 95%/92% vs Jev 89%/82%). **jevinf** open replica engine (NanoJev/decider-2b/Laya; 2.57×/2.27× 100% argmax; MPS only). **laya-multilingual** mmBERT-base 322M (MASSIVE 0.366/0.387 vs English laya 0.227/0.733; Khmer 0.000@0.952 conf; ships uncalibrated). **schema-scorer** DeBERTa-v3-large scalar head (Hub; GitHub 404; v2 Choice 0.841 *theirs*; peaked ranking ≠ calibration). **githubnext/localjev** prompted-JSON `/v1/systemone` (MIT **261★**; TypeSafe SDK drop-in; wire-compat ≠ logit-equiv vs razorback16 structured-read; **≠** kunchenguid/local-jev; 1,200-req bake-off *theirs* Qwen3.6 76.7% / Gemma 26B 75.0% / DiffusionGemma 74.2% short; not calibrated). **NandhaKishorM/laya** PyPI+Router packaging of Hub Laya (Apache-2.0; **710★**; not a new species; T4 32.8 ms *theirs*; post-T ECE 0.081 vs Jev 0.246; Banking77 0.425 vs Jev 0.870; 0.766 is fine-tune not zero-shot; 0.85 still soft; **≠** TypeSafe `/v1/systemone`). **external openjev census** (@airesearch12 tweet ≠ jevbench v1.1; GLiNER2+routers class-boundary; incomplete vs watch). **JevBench v1.2 scored board** (geo-mean I/C/S/K 25% each; Jev 75.3 / SemIf 74.6 *theirs*; cal ON rank; instruction models in the table; ≠ v1.1 87.6; ≠ tweet census; Laya absent gap; Qwen3.8 27B ≠ Archer). **Hourly 0842:** already-folded class as a recipe (wire≠logit · product+FALLBACK · packaging honesty · pointer-not-generator · leaderboard VOI); skip thin noise | `references/judgment-class.md` | +| Open weights vs constrained decoding vs proprietary API | Three open paths: encoder open-jev / AR constrained decode (TypeAR + pcdServer; decision-token LoRA; packed one-forward logprob on open LLMs; **CUDA/PyTorch local replica** jevify — uncalibrated likelihoods ≠ Noul) / trained decision-only (Laya + ONNX port, Nimble, kev, **blackwood-rlcd** multimodal now, Archer Watch still Watch). **Domain LoRA specialist on independent gold** (not a Jev teacher-copy): train when downstream reads p; few-shot hosted when only argmax. Local `/v1/systemone` surfaces: jev-local (stub until `hf`), kev (trained pointer), von (tiny SAN — not a replica), **jeff** (GLiFormer-400M encoder, typesafe-sdk drop-in — not a Jev replica), **local-jev** (ModernBERT approximation — not equivalence). Laya ONNX: Mattepiu port vs **gqgs** complete browser int8 (distinct). Loopback **gateway** (sysone) routes hosted + local; does not run weights. Softmax over allowed tokens ≠ Noul. **jevmlx:** MLX one-pass schema→JSON+probs (Apple Silicon replica economics; not a Jev replica). **openvons:** independent open-Jev class (Apache-2.0 code; GitHub SPDX NOASSERTION); wire-compat `/v1/systemone`; NOTA + execute/confirm/reject. **OpenJev** `/v1/decide` ≠ TypeSafe. **semif-serve:** SemIf runoff wire (MIT pyproject / GitHub SPDX null). **grande** Rust/WebGPU `/v1/systemone` (license null; softmax ≠ Noul until T). **laya-jolt** Clojure Apache-2.0 byte-parity Laya. **JEV-CPU** CPU SemIf. **local-jev** ONNX NLI approximation (confidence omitted). **chakuho:** 1-token logprob constrained-AR endpoint (uncalibrated; coverage is format-mass). **jevinf:** Jev-kind engine + wire (not a replica). **laya-multilingual:** mmBERT-base for non-English; route by script. **githubnext/localjev:** Bun Chat Completions bridge; prompted JSON + entropy confidence; **≠** kunchenguid/local-jev; **≠** razorback16 logits. **NandhaKishorM/laya:** PyPI `laya` + Router over the three Hub ckpts; packaging ≠ new species; Jev still leads >20 options / soft-acc / raw ECE. **@airesearch12 census:** list ≠ rank; GLiNER2/routers counted as openjevs are a class-boundary. **JevBench v1.2:** same class table includes Luna/Gemini/DeepSeek/Qwen3.8; OpenJev on board = razorback16 DiffusionGemma ≠ IamBusy; SemIf formerly OpenJev | `references/judgment-class.md` (when-to-use table) | | Entropy as allocator (low / medium / high) | Typed low+medium decisions → System One marginals; high-entropy synthesis → frontier decoder. Product rhetoric, not a meter. **Hypothesis** | `references/judgment-class.md` | -| Formal / semi-formal (proof vs judgment) | Sensor vs constraint vs searchlight; Alloy vs Apalache; DST trio; TOCTOU-of-Noul, AI×FM | `references/formal-methods.md` (one-screen: `references/formal-semi-formal.md`) | -| Mixed architecture (judgment model + LLM) | Provider judges, LLM writes, code owns control; not a stack replacement | `references/mixed-architecture.md` | -| Context sieve | Relevance Noul per block; always-keep set in code; stub + recall key | `references/applied-mappings.md#1-context-sieve` | -| Exact-text keep / drop | Choice include/exclude/mixed over candidates code already holds | `references/applied-mappings.md#2-exact-text-keep--drop` | -| Environment / harness triage | Scan every step for env failure; LLM autopsy only on flags | `references/applied-mappings.md#3-environment--harness-triage` | -| Moderation and ranking | Hold-before-publish vs graded rerank; fail policy per action | `references/applied-mappings.md#4-moderation-and-ranking` | -| Skill / tool routing | Choice over a closed catalog + whether-anything-fits; code dispatches | `references/applied-mappings.md#5-skill--tool-routing` | +| Formal / semi-formal (proof vs judgment) | Sensor vs constraint vs searchlight; Alloy vs Apalache; DST trio; TOCTOU-of-Noul, AI×FM. Skills→oxlint: AST/precheck compose with remainder judgment without hard-gating a Noul as a proof. PR attention ≠ correctness (anti-soundness-theater). Engine owns truth / Jev owns judgment (Stockfish+Jev chess coach). Capability kernel: type-safe ≠ correct; irreversible behind threshold AND human. Eval integrity: check the instrument, not just the score (`dinostomp jev` tests a question like an if-statement). Effect contracts, not surface tokens (construct-auto-classifier; privilege ≠ verdict). **Jev supplies evidence, code owns authority** (actiongate-jev; a positive score never overrides a deterministic security failure). **Turnstile clone:** deterministic policy + Jev remainder + replay (evidence ≠ authority). **Type-safe ≠ correct as jaggedness receipts** (atlas; schema-valid ≠ picked-right). **TLA+ compose with a Jev-class oracle:** never confidently wrong; escalate is the safety valve (jev-labs; inverse of soundness theater is hard-gating without escalate). **Advance/coverage ledger:** Jev answers questions; SEAL answers whether the world may change (coverage.path auto|code|human|escalate; mint ≠ product brain). **Conflict ≠ ignorance:** Noul collapses both; named Choice escape separates (typed-evaluation-collapse; schema-as-interface). **Sentence-as-rule lint:** ast-grep matcher silent × Jev `ask:` loud (mizchi/jevlint ≠ huntedman/JevLint). **Whole-repo intent:** VERIFIED/VIOLATION/UNKNOWN; empty search ≠ proof (jev-intent-review) | `references/formal-methods.md` (one-screen: `references/formal-semi-formal.md`) | +| Mixed architecture (judgment model + LLM) | Provider judges, LLM writes, code owns control; not a stack replacement. Advisory sidecar never changes host routing. Dual-process: S1 decides, S2 generates (routing accuracy unmeasured); **Harbor-shaped cousin:** decide→policy→LLM leftover (shared Answer schema; jev vs gen-json vs gen-logprob; Noul 0.5 never rounded). **Closed-vote CU:** no planner LLM; code builds options, Jev only picks (JevOnly). **Host-owned product:** app retains handlers/permissions; Jev over live typed actions (waymode). S1 specialists + S2 coordinator is the same split (description-only greenfield). Internals ≠ FSM; placement is a **component node**. Browser-use strength = DOM-as-text + speculative fan-out, not vision. Fail polarity is per act: skip-wake fail-open vs merge-gate BLOCK fail-closed; OMP/pi acceptance+route **fail-open** (contrast pi-jev-approver fail-closed). OMP prompt suppression: operator owns the bar; not a sandbox (omp-greenlight). Skill-broker outline: Jev never grants access. Hybrid local decide + remote fill; `DONE` ≠ verified success. Specialist computer-use: plan ≠ execute, dry-run default (Cua-S1; not TypeSafe Jev). Judgment as a language primitive (Ruby `almost_certain?`/`pick`/`rate`). Decision-native RAG: retrieve wide → decide → evidence set → LLM. Classify-first MCP (topology A): content to the judge without entering main agent context first. Generative UI: model decides, compiler emits. Draft-gate silence ≠ safer (heartbeat). Living class-pattern atlas (not a 342-title dump). Stagehand experimental Jev: pick-and-copy extract + act tree + observe/cache-check; LLM fallback; draft stack. Public judgment wall (six parallel questions; policy-in-code; cost-to-1M). PR attention ≠ correctness. Session-sticky first-prompt route (fail-closed fallback). Capability kernel (LLM ring 3 / Interlock ring 0; secrets never in agent; Jev SENSOR; policy.py BLOCK/ASK/ALLOW; type-safe ≠ correct). Typed control plane around DSPy (drafts AFTER route+action). Engine owns truth / Jev owns judgment. Human-confirmed kill (mapped explanations; identity re-check). Wire-compat encoder backend (GLiFormer `/v1/systemone` drop-in; cheaper, less accurate on reasoning-heavy). Loopback gateway routes hosted + local (not a model). **Constrained optimizer + S1 features** (slo-router: Jev never the sole hot-path gate; fail-open local features). **Effect-based shell gate** (construct: privilege ≠ verdict; fail-closed). **Attention filter / VOI** (jev-lens: never blocks the agent; never green unless sure). **Jev supplies evidence, code owns authority** (actiongate-jev). **Measurement owns endorsement** (jev-packs evidence-gated). **Ranking ≠ calibration** (does-jev-confidence; never hard-threshold raw p). **Hot-click CU** (ego-jev: indexed table → operation+target; text model only for type). **Jev judges relevance, code decides structure** (jev-compactor; never rewrite; regex floor). **Local rules first / never auto-train on own hides** (x-reply-filter). **Control-plane combinators** (Then/Gate/Vote/Cascade/Weighted; not chat turns). **Skill VOI / abstention** (skillranker; hook fail-open). **Receipts not leaderboard** (atlas + frontier-100 + OOD). **Turnstile** evidence≠authority + replay. **MLX one-pass replica economics** (jevmlx; softmax ≠ Noul). **TLA+ consensus kernel** (jev-labs; never confidently wrong; escalate). **SEAL advance/coverage** (no seal, no advance; exception queue visible). **Sureness bands** (how-sure-is-jev; max_prob is generous). **JevBench v1.1** (calibration reported, not scored). **CI typed gate** (ci-gatekeeper before expensive review). **Codex MCP adapter** (jev-in-codex; ranking unbenchmarked; lexical fallback). **Stop-hook attention redirect** (jev-preflight; eight axes; assist=one reinspect; fail-open; not a merge blocker). **Pre-send view selection** (dizk/jev-lens; 79% fewer tokens *theirs*; compress-before-first-send; distinct from rashedInt32/jev-lens). **Observational memory** (pi-om; keep/kind verbatim; model-free compact). **tools≠use** (carryforward 0/4 recall; SessionStart > hoping). **Physical-world S1** (HA-Jev; sensors from typed answers; not for locks/heaters). **Judgment outside the store** (jevql CLI; DB never sees `jev()`). **Landed-script / headless≠auto-approve** (construct); **digital-design combinators** (jev-combinators rename + Router/Loop/Retry/Fallback/Memory; metaphor ≠ literal AND/OR); **VOI cache admission** (jevcache same-intent; 0 FP/100 *theirs*; fail-open); **worth-your-attention VOI** (ThinkyMiner/Winnow 80%/90%; ≠ kevinpita/winnow); **Jev WHETHER / Python HOW / LLM WHAT** (hermes-jev-router; license null); **typed escalate/continue/abort baton** (jev-handoff; inverted loop; gate never grants; fail-open); **Playwright executes, Jev chooses** (browser-jev; sample-from-distribution); **OpenJev `/v1/decide` ≠ drop-in** + **SemIf runoff wire**; **conflict ≠ ignorance** (named Choice escape); **decision-as-memory flywheel** (DGUI_HYPERMEM-JEV 6-row schema); **record/replay CI** (jevassert landed; accuracy+ECE+cost gates offline); **failure-finding arena** (chenmingtang830/jevarena ≠ meetr1912/jev-arena); **BBQ** 97.28%/0.04/0.34/$0.3429 *theirs*; **decider≠executor** (jeffrey: Jev next-tool, LLM fills args); **sentence-as-rule** (jevlint); **VOI hunk prune** (prune-review ~20% target; 1.18% with outlier); **persist constraints across compaction** (pi-heed); **Harbor SGR-judge contract** (jev-judge-bench; canaries ≠ quality; ≠ jevarena/jevbench); **hand no-text steps** (jev-use; Vercel drops confidence; margin 0.4; fail-open gate; ≠ jev-ultrafast); **Pi System-One control plane** (pi-jev-control; GUI never force-click; compaction never writes session); **never free-generates** (jev-gpt tree of Choices; 400 calls / 75 s / 2¢ *theirs*); **OpenRouter recipe atlas** (jev-cookbook; 16–36 item samples not benches; 425 calls / $0.015); **personal-history feed** (jevfeed; no social graph; one request per batch of ten); **competing NAR claim-audit** (openJev-verdict-2.0; dual-channel ECE; PR #1; ≠ IamBusy/OpenJev); **empty compaction-proxy skip** (IPECTER context-pruner **and** jev-runway LICENSE-only); **1-token logprob endpoint ≠ Noul** (chakuho; coverage ≠ correctness); **open replica engine** (jevinf argmax-parity); **unofficial Elixir SDK ≠ OTP peer** (typesafe-elixir-sdk ≠ dannote/jev); **jevex n=16 files-to-read VOI** (rename of jev-semantic-explorer); **commit pre-review attention≠verdict** (commitjev; middle band never rounded); **Hermes plugin is Agnes not TypeSafe** (hermes-plugin-jev); **pi-jev-compact ≠ pi-jev-compaction** (verbatim summarizer replacement); **decision-native inbox** (mailordinal; humans own ambiguity); **unofficial jev-cli not ready** (≠ jevql); **laya-multilingual** English checkpoint confident-wrong OOD; **schema-scorer peaked ranking ≠ calibration**; **productized System One HTTP** (classifier.dev; label+p; batch ~1000; Jev primary / LLM fallback); **escalate-under-threshold** (smart single-label <0.7; multi-label ignores); **silent FALLBACK** (granite 0.546 vs advertised 0.800; rh-guard owns the gate); **systematic-review pointer (choxos/jev-reviewer ≠ egma-ai)** two-pass Choice+Noul; *Not found* is an answer; human check is the product; **githubnext/localjev** prompted JSON ≠ structured-read logits (wire-compat ≠ logit-equiv; **≠** kunchenguid/local-jev; GitHub Next **261★**; 1,200-req caveats *theirs*); **NandhaKishorM/laya packaging** Router script-before-p; post-T ECE ≠ raw ECE; 0.85 still soft; Banking77 token-budget; **≠** TypeSafe drop-in; **external census ≠ scored bake-off** (@airesearch12; GLiNER2+routers class-boundary; incomplete vs Laya/localjev/kev; Harbor honesty watch); **JevBench v1.2 geometric-mean product** (I/C/S/K 25% each; cal ON rank; Luna I=96.8 rank #7; ×2/est. Harbor honesty; option-order 72→21; Laya absent gap; Qwen3.8 27B ≠ Archer); **hourly 0842 apply-the-five** (already §73–§78; do not re-card); **skip thin noise**; **hard-gate a Noul as a PR/quality gate is soundness theater** (totally-tim/jev-gate / claude-jev-warden; ≠ jev-gateway / MongLong0214/jev-gate / jev-gate-student-b); **S1 keeps flying / S2 one-use** (khordoo delta: escalate without stall; purple = consumed; Local controller ≠ githubnext/localjev; seed = geometry; 20% still soft; no pixels; S2 never grants). **OCR+AX desktop CU:** typesafe-computer-use (hosted Jev; never screenshot-to-frontier for the decision; overlapping options = doubt; writer/decider; 155× *theirs* one screenshot; 0.4/0.5 still soft; **≠** jev-ultrafast **≠** cua-s1 **≠** camoufox). **ASR voice-browser CU:** jev-voice-browser (partial-speech VOI; pointer spans; spoken confirm ≠ auth; 27/27 *theirs* fixtures; **≠** jev-voice-control **≠** nikolas-j **≠** OCR desktop). **Wrap-as-execution ALLOW/ASK/DENY:** wrap *is* the tool function; rules first; ASK throws; fail-closed (AgentGhost; **≠** actiongate **≠** jev-use fail-open; rh-guard owns the gate cousin). **JP genre atlas:** apps by hole; stars research-time; not verified evals (@studio_yebisu; **≠** class census §77 **≠** v1.2). **External pedagogy:** Akshay “Jev Clearly Explained”; LLM hammer; schema-safe ≠ correct; 200×/400× TypeSafe ceiling; shadow + questions-as-code; **≠** official docs **≠** Flavio **≠** AgentGhost. **Meaning-grep dedicated:** proposition≠embedding; boolean composition of thresholded Nouls; Semgrep.dev collision; not a gate (jev-semgrep §86) | `references/mixed-architecture.md` | +| Context sieve | Relevance Noul per block; always-keep set in code; stub + recall key. Encoder cousin: GLiNER2.5 retention Choice + char-offset spans (gliner25-compaction); fail-closed keep_full; shadowMode default. Stdout cousin: jev-pruner (Jev Noul after hard ≤10k/JSON-diff envelope; fail-safe original; archive). Session-ledger cousin: carryforward (verbatim facts; Jev scores recall; rules never judged; fail-open dump). Classify-first MCP cousin: jev-sift (batch path/url/text → Jev without entering main agent context; uncertain/errors/truncation ≠ irrelevant). **Framework-agnostic compact+gate:** Jev judges relevance, code decides structure; never rewrite; regex floor in code; compaction fail-open if Jev down, safety gate fail-closed (jev-compactor later bench **73%** / 350 ms / 4 of 4 *theirs*; OpenCode port fast-jev-opencode already §62). **Pre-send views:** dizk/jev-lens (79% fewer tokens on 500 SWE-rebench trajectories; code full unless confident). **Observational memory:** pi-om keep/kind verbatim. **tools≠use:** carryforward SessionStart hook > MCP sitting there. **Hermes WHETHER/HOW/WHAT:** compact original chunks; skip next main-model when evidence is enough (hermes-jev-router; needs core patch; fail-open). **Empty compaction-proxy skip:** IPECTER/jev-context-pruner **and** IPECTER/jev-runway LICENSE-only. **Pi verbatim summarizer replacement:** pi-jev-compact (keep/drop tool calls; fail-open to LLM summary; ≠ vava-nessa/pi-jev-compaction) | `references/applied-mappings.md#1-context-sieve` | +| Exact-text keep / drop | Choice include/exclude/mixed over candidates code already holds. Extractive quotes / pointer-not-generator (model never writes the excerpt; char-offset compaction same species). Observed a11y/DOM controls: score among them; code clicks (Jev or GLiNER2 or Cua-S1 option-attention). Harness productization: Stagehand extract pick-and-copy (Jev picks; code copies; schema/gate else LLM). **Closed-vote CU:** code builds options, Jev only picks, no planner LLM (JevOnly); host-owned handlers/permissions (waymode). **Hot-click CU:** indexed element table → operation+target in one request; code owns observe/execute/verify; text model only for type (ego-jev; ~2× vs per-step LLM, n=3, not a bench). **OCR+AX desktop CU:** numbered OCR+AX items; hosted Jev; writer only for free text (typesafe-computer-use; exclusive actions; split kind/item/site; **≠** jev-ultrafast). **ASR voice-browser CU:** regex spans; Jev picks; code copies (jev-voice-browser; numbered overlay, no second model). Meaning-as-spec: Cucumber .feature only; Jev picks among observed controls (jevcumber). **Evidence-synthesis pointer (choxos/jev-reviewer, ≠ egma-ai):** two-pass Choice (which line) + Noul (does this line itself answer); *Not found* is an answer; human check never overwritten. **Adversarial browser:** Playwright executes, Jev chooses next act; sample from the distribution not argmax; fail only high conf **and** high severity (browser-jev) | `references/applied-mappings.md#2-exact-text-keep--drop` | +| Environment / harness triage | Scan every step for env failure; LLM autopsy only on flags. Merge-gate cousin: cluster in code, judge labels cause, policy owns PASS/BLOCK (latch; judge never says ignore alone). **Pre-review typed gate:** should_review/risk/route/touches_secrets → auto-approve|human-review|block (ci-gatekeeper; cheap before expensive LLM/human; operator-owned thresholds). **Stop-hook attention redirect:** eight risk axes, assist=one reinspect, fail-open, uncalibrated 0.85 (jev-preflight; not a merge blocker). **VOI hunk prune:** Jev scores PR hunks before expensive generative review (prune-review; 22-run 1.18% with 305% outlier *theirs*; ~20% target). **Whole-repo intent:** VERIFIED/VIOLATION/UNKNOWN beyond the diff (jev-intent-review). **Commit pre-review:** message vs diff / cohesion / omitted change; middle band is review not a verdict; regex proves literals (commitjev) | `references/applied-mappings.md#3-environment--harness-triage` | +| Moderation and ranking | Hold-before-publish vs graded rerank; fail policy per action. Meaning-search without embeddings (jevgrep packed parallel; 79% top-5 on stripped repos; keyword still wins exact strings). **Meaning-grep** AND/OR/NOT over *thresholded* line Nouls; proposition≠embedding; Semgrep.dev collision; not a gate (jev-semgrep §86). **Evidence-packet explorer:** index-once ask-many, citable source_of_truth/tests/callers (jevex). Measured RAG rerank vs generative rerank (Jev-RAG one-run ≥70% cost / 72% latency vs Spark rerank; full-context Spark still faster). **Local rules first then remainder Nouls:** never auto-train on the model's own hides (x-reply-filter). Minimal consumer labels: bohutang/sift ~$0.00003/post. **Worth-your-attention VOI:** ThinkyMiner/Winnow read/skim/save/skip from typed answers (80%/90% *theirs*; always qualify vs kevinpita/winnow). **Personal-history ranking without a social graph:** jevfeed (one Jev request per batch of ten; distribution *is* ranking; history never uploaded). **OpenRouter recipe atlas:** jev-cookbook (triage/PII/rerank/browser; samples not benches). **Decision-native inbox:** mailordinal (nine typed signals → 100-point policy; humans own ambiguity). **Evidence-packet explorer delta:** jevex n=16 SWE 160s→69s / $8.74→$3.13 *theirs* (rename of jev-semantic-explorer). **Productized classification API:** classifier.dev (label+calibrated p; spam/inbox/feedback; **185★**). **Open NAR packaging:** NandhaKishorM/laya PyPI+Router (**710★**; Hub weights; not HTTP Jev) | `references/applied-mappings.md#4-moderation-and-ranking` | +| Skill / tool routing | Choice over a closed catalog + whether-anything-fits; code dispatches. Route ≠ memory: cheap intent gate skips memory tours on easy routes. Session-sticky first-prompt classification (lock for the session; fail-closed to a declared fallback). OMP/pi: `jev_route` topology/tier + `jev_acceptance_gate` before done (**fail-open**; contrast pi-jev-approver fail-closed). OMP prompt suppression: Jev grades gated calls; operator owns the bar; plugin never self-tunes (omp-greenlight; not a sandbox). **Outline only:** Hermes pre-agent skill broker — code owns grants; Jev never grants access (skill-broker; not a production recipe). **Constrained optimizer:** Jev supplies task/exactness/evidence features; controller owns SLO/quality floors; fail-open local features (slo-router; measured p95 77.93→490.38 same routes). **VOI over skill library:** two-pass + none-of-these; hook never blocks (skillranker 52★). **Sibling contrast:** skill-broker grants in code (fail-closed foundation-only) vs skillranker advisory vs turnstile runtime authorize. **Codex MCP adapter:** jev_select_capability / jev_search / jev_triage; caller supplies catalog; lexical fallback (jev-in-codex). **Harbor roster-size harness:** BM25 vs Jev at 50–500 (pi-jev-skill-bench; 43 gold; no live numbers this pass). **Pi strip-roster:** two-stage gate 0.30 / fits 0.40; no key → no-op; tool mode is tools≠use cousin (pi-jev-skill-suggestion). **Pi System-One control plane:** pi-jev-control (task/model/skill/memory/review/GUI; no live quality numbers; license null). **Host-adapter surface delta:** jev-routing adds Cursor Agent CLI / Devin CLI (still not MCP). **Hermes plugin ≠ TypeSafe:** hermes-plugin-jev is Agnes 3.0 Flash chat-completions branded as Jev | `references/applied-mappings.md#5-skill--tool-routing` | | Expensive observation router | Structural prove (text layer) ∩ remainder Noul (needs OCR?) | `references/applied-mappings.md#6-expensive-observation-router` | -| Agent preference lint / semantic gates | Project-defined rules as criteria; provider classifies evidence; code maps outcome | `references/mixed-architecture.md#preference-lint-and-gates` | +| Capability kernel / human-confirmed gate | Secrets never in the agent; closed action space; Jev SENSOR; policy BLOCK/ASK/ALLOW. Distinct from pre-exec toolgate. Human is the only kill trigger; identity re-check; shields override; mapped explanations not raw model prose. **Permission vs probability:** operator owns auto-approve thresholds; plugin never self-tunes the safety bar; host deny stays above (omp-greenlight; not a sandbox). **Spoken confirm ≠ auth:** jev-voice-browser `destructive` spoken "confirm" is convenience not a guarantee (control-port reach is the grant; rh-guard owns the gate cousin). **Privilege ≠ verdict:** effect-based shell gate; fast structural rules then Jev Choice + independent risk Nouls; fail-closed (construct-auto-classifier; 0 dangerous / 975 *theirs*). **Jev supplies evidence, code owns authority** (actiongate-jev; a positive score never overrides a deterministic security failure). **Turnstile:** policy first, Jev after permit, missing Jev → Review, replay thresholds. **SEAL:** no seal, no advance; coverage.path visible; mint ≠ product brain. **Never confidently wrong:** TLA+ kernel may escalate, must not return a confident wrong (jev-labs). **Landed-script trust / headless≠auto-approve:** construct (byte-identical to remote default branch; headless escalation is deny-and-report, not auto-approve). **Typed baton:** escalate/continue/abort; gate `allow` never grants (jev-handoff). **Persist constraints across compaction:** conversational policy as structured state; Jev never writes policy; fail-open (pi-heed 98.5%/0 false block *theirs*). **Pi control-plane tool gate / GUI:** pi-jev-control (deterministic fast-path then Jev; GUI < threshold → unknown, never force-click). **jev-use PreToolUse gate:** deny/ask, fail-open, 12/12 *theirs*; Vercel margin fallback. **Commit-msg hook:** commitjev blocks only on a warning; instrument failure is not a refuse. Toolbelt notes: jev-security-scan / jev-decisions / TeoMastro; unofficial jev-cli not ready (≠ jevql); actiongate slogan already §64; **rh-guard owns reward-hack**. **Wrap-as-execution ALLOW/ASK/DENY:** the wrap *is* the tool function; rules first; ASK throws; fail-closed (AgentGhost; **≠** actiongate **≠** toolgate **≠** jev-use; rh-guard owns the gate cousin). **Tool-risk as placement (not a wrap card):** Akshay pedagogy cites LangChain middleware; AgentGhost owns wrap-as-execution | `references/applied-mappings.md#7-capability-kernel--human-confirmed-gate` | +| Decide → policy → LLM leftover | Typed decide backends share one Answer schema; policy in code routes auto/review/llm; generator writes leftover text only. Noul 0.5 = cannot-tell, never rounded. Score conf 0.0 = flat, never acted on. Hard flags always review. **Inbox cousin:** mailordinal (typed signals → deterministic priority; no leftover LLM required). **Public decide-backend:** classifier.dev (HTTP classification API; leftover LLM is fallback when Jev is down). **Open NAR cousin:** NandhaKishorM/laya (self-hosted Choice/Score/Noul; Router picks ckpt; policy still in the caller) | `references/applied-mappings.md#8-decide--policy--llm-leftover-cascade` | +| Closed-vote computer-use | Code builds every option from observation/goal/facts; decision model only picks; code acts/verifies/undoes. No planner LLM. Host-owned product cousin: app retains handlers/permissions; Jev over live typed actions. **Hot-click cousin:** indexed viewport table → operation+target; code owns the loop; generator only for type (ego-jev). **Adversarial cousin:** Playwright executes, Jev chooses explore/continue (browser-jev; sample not argmax). **Decider≠executor cousin:** Jev picks next-tool/progress/risk/done; LLM only fills args (jeffrey; risk≥0.5 pause; stuck ladder). **Hand no-text steps:** jev-use (plugin; writing stays with the LLM). **Never free-generates:** jev-gpt (WordNet tree of Choices; architecture demo). **OCR+AX productized cousin:** typesafe-computer-use (macOS; exclusive kinds; split questions; post-type Noul still soft; 155× *theirs* one screenshot). **ASR voice-browser cousin:** jev-voice-browser (Playwright; partial-speech wait; spoken confirm ≠ auth) | `references/applied-mappings.md#9-closed-vote-computer-use` | +| Agent preference lint / semantic gates | Soft project rules as criteria; linter owns hard rules; one Score per rule on the diff; bands + fail-open; name the observation window (edit vs turn). Named Choice escape when conflict ≠ ignorance (typed-evaluation-collapse). Sentence-as-rule cousin: ast-grep `rule:` × Jev `ask:` (mizchi/jevlint; 13/15 1.00/1.00 *theirs*; fail-open no-verdict). Independent of huntedman/JevLint | `references/mixed-architecture.md#preference-lint-and-gates` | | Dual orchestration (Jev ∩ LLM ∩ MCP) | Jev-as-tool vs Jev-as-outer-loop; schemas are exact state | `references/mixed-architecture.md#dual-orchestration-jev--llm--mcp` | -| "It's just classification" / "not probabilistic programming" / stack-replacement FAQ | Typed judgment is a software primitive, not a new task; marginals are not a joint; Jev is not the only model | `references/faq.md` | +| "It's just classification" / "not probabilistic programming" / stack-replacement FAQ | Typed judgment is a software primitive, not a new task; marginals are not a joint; Jev is not the only model. **Public pedagogy:** Akshay “Jev Clearly Explained” — LLM hammer; schema-safe ≠ correct; 200×/400× TypeSafe ceiling. **Meaning-grep:** jev-semgrep ≠ semgrep.dev; proposition≠embedding; not a gate | `references/faq.md` | | Feature engineering / multi-criteria analysis | Nouls + Score distributions as named features, weights in code | `references/mappings.md#1-semantic-judgments--features-and-explicit-utility` | -| Selective classification / decision theory | Thresholds from action costs, abstention paths | `references/mappings.md#2-probabilistic-judgments--cost-sensitive-decisions` | -| Decision tables / circuits / state machines | Judgment predicates, code owns transitions | `references/mappings.md#3-semantic-predicates--decision-circuits` | -| Retrieve + expensive relevance fn | Bounded rerank of a retrieved shortlist | `references/mappings.md#4-retrieval--bounded-semantic-reranking` (independent TREC DL2019 benchmark: Jev zero-shot best MAP 0.4748, nDCG@10 0.683 vs tuned monoBERT 0.718 — competitive, not dominant) | -| Value of information / gather as an act | Pay for another observation only if EV(decision) improves more than cost; abstain from calling *any* model when a regex already answers (meta-VOI) | `references/mappings.md#6-value-of-information--gather-as-an-enumerated-act` (**Hypothesis** until a labeled act/outcome log; 149-row receipt is Empirical as a shape) | -| Signal detection / ROC | Criterion and operating point from costs and base rate, not accuracy | `references/mappings.md#7-signal-detection--criterion-not-accuracy` (**Hypothesis** for non-SWE plots) | -| Org / safety control structure | Sensor ≠ constraint (Leveson); STPA if the sensor lies | `references/mappings.md#8-control-structure--sensor--constraint-leveson` | -| Search / control loops (any domain) | Algorithm stays yours; judgment substitutes one classifier step | `references/mappings.md#9-search--control-loops--one-substituted-classifier-step` | +| Selective classification / decision theory | Thresholds from action costs, abstention paths. Train a domain specialist when downstream code **reads the probability**; few-shot hosted API when only **argmax** matters (calibration/VOI, not an accuracy bake-off). **Active-learning triage:** high conf accept / middling expensive teacher / low-or-boundary human; log full distributions. **Do not distill Jev as teacher of record** (~68% ceiling compounds errors; real outcomes stay the targets) | `references/mappings.md#2-probabilistic-judgments--cost-sensitive-decisions` | +| Decision tables / circuits / state machines | Judgment predicates, code owns transitions. Language primitive: Ruby `chance`/`pick`/`rate` as control flow (hunch; English-as-config; fail polarity per action) | `references/mappings.md#3-semantic-predicates--decision-circuits` | +| Retrieve + expensive relevance fn | Bounded rerank of a retrieved shortlist. Decision-native RAG: retrieve wide → decide explicitly → evidence set → conflict resolve → reason only over kept evidence (embeddings stay candidate generators; no universal benchmark). Classify-first MCP: same sandwich on agent I/O (path/url/text → judge; main LLM opens survivors). Meaning-search without embeddings (packed parallel relevance; two-stage outline→zoom). **Meaning-grep** AND/OR/NOT over *thresholded* line Nouls; proposition≠embedding; not a gate (jev-semgrep §86). **Evidence-packet explorer:** index-once, BM25 shortlist, Jev ranks, citable source_of_truth (jevex). Measured pointwise rerank vs a generative reranker (one-run; name the no-RAG arm) | `references/mappings.md#4-retrieval--bounded-semantic-reranking` (independent TREC DL2019 benchmark: Jev zero-shot best MAP 0.4748, nDCG@10 0.683 vs tuned monoBERT 0.718 — competitive, not dominant) | +| Store as semantic index (SQL / SQLite / zoxide / dataframe) | Cheap exact predicates first; typed questions on the remainder. In-engine extension (sqlite-jev / pg-jev) vs **judgment outside the store** (jevql CLI; vanilla Postgres never sees `jev()`) vs path index (joxide) vs dataframe columns (jevpandas / jevframe). **ORDER BY over probs is a ranking job:** calibration ≠ sortable; measure pairwise inversion / Score ordinality / two-decimal ties (jev-orderby-bench) | `references/mappings.md#4-retrieval--bounded-semantic-reranking` | +| Soft judgment inside a hard envelope | Model may only match the deterministic policy or be more conservative (bitrate ABR; query-planner override-when-confident; compaction mutations/shell operators → keep_full; stdout prune: ≤10k/JSON-diff-whole-doc untouched, then Noul; Cua-S1: plan≠execute, dry-run, fail-closed checkbox/fill; Stagehand extract: schema/completion-gate/screenshot-always-LLM then pick, else LLM; skills→oxlint: AST/precheck prove, guidance whole-file in state, remainder Noul — not a hard gate; pre-exec toolgate: allow/block/review — Jev is not authorization; guard error/timeout stops; distinct from capability kernel interlock — secrets never in agent, closed action space. OMP prompt suppression: operator owns thresholds; plugin never self-tunes; host `bash.patterns: deny` stays the floor (omp-greenlight — not a sandbox). Human-confirmed kill: Jev recommends, human is the only trigger, identity re-check, shields override, mapped explanations (port-cleanup). Effect-based shell: fast-allow/deny prove, then Jev remainder; privilege ≠ verdict (construct-auto-classifier). Wrap-as-execution: rules prove allow/deny, Jev remainder, ASK throws, fail-closed (AgentGhost). SLO router: exactness raises quality floor, never overrides capability; Jev features fail-open to local (slo-router)) | `references/mappings.md#12-runtime-assurance-sandwich-hypothesis`; `references/mappings.md#18-structural-prove--soft-remainder-hypothesis-as-domain-general-empirical-as-named-shapes` | +| Value of information / gather as an act | Pay for another observation only if EV(decision) improves more than cost; abstain from calling *any* model when a regex already answers (meta-VOI). Fail-open wake/resume: skip the LLM turn only if the judge answers and p(wake) is low (Horvitz). Selective memory: verbatim ledger + scored recall (carryforward; rules never judged; 9×3 is a hint). Classify-first read: pay for a full agent open iff relevance (or typed question) says it might change the act (jev-sift; errors/truncation ≠ irrelevant). **Training-data VOI:** spend expensive teacher/human labels only where confidence says they change the outcome (jev-triage); do not distill Jev as teacher of record. **Decision-model latency cost:** measure p95 of sync Jev on the hot path before claiming routing (slo-router 77.93→490.38 same routes). **Human-review VOI:** attention filter that never blocks the agent and never says green unless sure (jev-lens). **Skill-library VOI:** load a skill iff it changes the next step; abstention first-class (skillranker). **tools≠use:** an MCP memory tool the agent never calls is not VOI (carryforward 0/4; SessionStart hook). **Pre-send token-econ:** compress before first send (dizk/jev-lens; post-send prune cost 17% more *theirs*). **Same-intent cache admit:** skip the LLM iff Jev says same intent (jevcache 0 FP/100 *theirs*; fail-open). **Human-feed VOI:** worth-your-attention before click (ThinkyMiner/Winnow). **Second-call VOI:** skip the narrating main-model turn (hermes-jev-router WHETHER/HOW/WHAT). **Hunk-review VOI:** score PR hunks before the generative reviewer (prune-review). **Intent-search VOI:** pay to judge places the diff did not touch (jev-intent-review). **Empty compact-proxy skip:** IPECTER context-pruner **and** jev-runway slogan only. **Batch ranking VOI:** one request per ten (jevfeed). **No-text-step VOI:** hand the judgment, keep writing on the LLM (jev-use 186 vs 2,672 ms *theirs*). **Files-to-read VOI:** jevex n=16 160s→69s *theirs*. **Commit-attention VOI:** commitjev middle band never rounded. **Pi compact VOI:** pi-jev-compact keep/drop vs LLM summary. **Escalate-under-threshold VOI:** classifier.dev smart re-asks single-label <0.7; multi-label ignores (re-judge worse). **Escalate-without-stall VOI:** khordoo S2 is async one-use; S1 never waits; log consumed not arrived. **Split-question CU VOI:** kind/item/site/offscreen in one request; exclusive actions; perception rebuild is the expensive gather (typesafe-computer-use; 155× *theirs* one screenshot). **Partial-speech VOI:** complete Noul; closed-set may fire; free-text waits (jev-voice-browser). **Evidence-synthesis two-pass VOI:** choxos/jev-reviewer fan-out then absolute Noul; *Not found* cheaper than paraphrase (≠ egma-ai). **Script-before-p VOI:** NandhaKishorM/laya Router routes before the forward pass because Khmer 0.000@0.952 gating cannot catch | `references/mappings.md#6-value-of-information--gather-as-an-enumerated-act` (**Hypothesis** until a labeled act/outcome log; 149-row receipt is Empirical as a shape; wakegate 21/21 is smoke; jev-triage is Empirical as README architecture; slo-router live analysis is Empirical as a *negative* on sync Jev; jev-lens is Empirical as README architecture) | +| Signal detection / ROC | Criterion and operating point from costs and base rate, not accuracy. Operator owns the criterion; a plugin must not self-tune the safety bar (omp-greenlight 40.9% / 0 of 94 *theirs*). Exactness raises a quality floor — it must not override capability/context (slo-router). Privilege ≠ verdict (construct; `sudo status` can be safe). **Ranking ≠ calibration:** never hard-threshold raw p as a frequency (does-jev-confidence; AUC ~0.91, stated ~75% vs human ~10%). **OOD / AUC ≠ ECE:** sign of miscalibration flips by type (jev-ood-calibration; unknowable policy label mean p 0.74). **Sureness over vectors:** max_prob/margin/entropy/gini/perplexity → CERTAIN|…|CLUELESS; Choice confidence = max_prob (most generous) (how-sure-is-jev). **Conflict ≠ ignorance:** Noul collapses both; named Choice separates (typed-evaluation-collapse). **BBQ stereotype/uncertainty:** 97.28% / bias 0.04/0.34 *theirs*; not a general bias cert (jev-bbq-experiment). **Cookbook moderation dial:** 5 Nouls; samples not benches (jev-cookbook). **Dual-channel ECE / like-for-like:** openJev-verdict-2.0 claim-audit (correctness-head ≠ distribution ECE; PR #1). **Commit middle band:** commitjev never rounds 0.35–0.65. **English-checkpoint OOD:** laya-multilingual (Khmer 0.000@0.952; gating cannot catch). **Coverage ≠ correctness:** chakuho 8B coverage 1.00 while `__none__` collapses. **Peaked schema-scorer:** rankings not calibrated confidences. **Escalate-under-threshold:** classifier.dev 0.7 is *theirs*; multi-label does not share it. **Laya 0.85 gating still soft:** README recipe, not Harbor-calibrated; post-T ECE ≠ raw ECE; Banking77 token-budget (NandhaKishorM/laya) | `references/mappings.md#7-signal-detection--criterion-not-accuracy` (**Hypothesis** for non-SWE plots) | +| Org / safety control structure | Sensor ≠ constraint (Leveson); STPA if the sensor lies. Capability kernel: Jev SENSOR, policy.py constraint; secrets never in agent (interlock). Host deny fires ahead of Jev prompt-suppression (omp-greenlight). Judgment ≠ permission: Jev never grants skill access (skill-broker outline). Contracts on effects, not tokens (construct). Attention filter ≠ permission gate (jev-lens). **Jev supplies evidence, code owns authority** (actiongate-jev; positive score never overrides a deterministic failure). **SEAL coverage ledger** (exception queue visible; mint ≠ product brain). **Never confidently wrong** (jev-labs; escalate instead of hard-gate). **Physical-world S1:** typed answers as HA sensors; not for locks/heaters (HA-Jev). **Gate never grants** (jev-handoff). **Conversational constraint sensor** (pi-heed; Jev never writes policy). **Pi control-plane sensors** (pi-jev-control: router/gate/retry/sieve/review/click; GUI never force-click). **jev-use gate never grants** (fail-open PreToolUse). **Inbox policy in code** (mailordinal; model never sets queue order). **Agnes branded as Jev is not a sensor** (hermes-plugin-jev). Toolbelt sensors ≠ policy (jev-security-scan / jev-decisions); unofficial jev-cli not ready; actiongate slogan §64; rh-guard owns reward-hack. **Wrap-as-execution:** wrap *is* the actuator (AgentGhost; ASK throws; fail-closed; rh-guard owns the gate cousin). **Meaning-grep is not a gate** (jev-semgrep ranking fail-open; rh-guard skip) | `references/mappings.md#8-control-structure--sensor--constraint-leveson` | +| Search / control loops (any domain) | Algorithm stays yours; judgment substitutes one classifier step. **Decider≠executor:** Jev picks next-tool/progress/risk/done; LLM only fills args (jeffrey; pick ≠ fill; still not a planner-writer). **Tree-of-Choices writer:** jev-gpt never free-generates (one question per word). **Pick≠write plugin:** jev-use. **1-token selector:** chakuho reads next-token mass over declared labels (not a writer). Robotics text-state (geometry-as-text, not pixels); khordoo: no graphical input; S1 flies, S2 advises without stalling; physics owns collisions; do not replace A* / a solver with a Noul. **OCR+AX desktop CU:** typesafe-computer-use (hosted Jev; never screenshot-to-frontier for the decision; exclusive actions; split kind/item/site). **ASR voice-browser:** jev-voice-browser (Playwright executes; partial-speech wait; spoken confirm ≠ auth) | `references/mappings.md#9-search--control-loops--one-substituted-classifier-step` | | Spec property pipeline | Rank candidate props; checker owns validity | `references/mappings.md#10-spec-property-pipeline-hypothesis` (**Hypothesis**) | | Alloy instance loop | Cluster CEXs; Analyzer owns in-scope truth | `references/mappings.md#11-alloy-instance-loop-hypothesis` (**Hypothesis**) | | Runtime assurance sandwich | Abstain → RV/monitor → act | `references/mappings.md#12-runtime-assurance-sandwich-hypothesis` (**Hypothesis**) | | DST multiverse triage | Cluster failing seeds/timelines; regress on the same seed | `references/mappings.md#13-dst-multiverse-triage-hypothesis` (**Hypothesis**) | | Durable agent control | Resonate protocol settles promises; Jev gates inside a step | `references/mappings.md#14-durable-agent-control-hypothesis` (**Hypothesis**) | -| Assignment hybrid | Soft affinity + hard solver | `references/mappings.md#15-assignment-hybrid--soft-affinity--hard-solver-hypothesis` (**Hypothesis**) | +| Assignment hybrid | Soft affinity + hard solver. **Empirical as shape:** slo-router constrained min-cost s.t. quality+SLO floors; Jev features, not the sole gate (p95 77.93→490.38 same routes *theirs*) | `references/mappings.md#15-assignment-hybrid--soft-affinity--hard-solver-hypothesis` (**Hypothesis**; slo-router Empirical as *shape*) | | Situated density (Shirky) | Aggressive soft loops only inside a named community | `references/mappings.md#16-situated-density-shirky-hypothesis` (**Hypothesis**) | | Input brittleness / paraphrase stability | Synonymous wording that swings p → abstain or rewrite | `references/mappings.md#17-input-brittleness--sensitivity-calibration-selective-abstention-hypothesis` (**Hypothesis**) | -| Structural prove ∩ soft remainder | Allowlist/text-layer/law first; judge only leftovers | `references/mappings.md#18-structural-prove--soft-remainder-hypothesis-as-domain-general-empirical-as-named-shapes` (**Hypothesis**; jevgate/OCR shapes Empirical) | +| Structural prove ∩ soft remainder | Allowlist *proves* the easy verbs; judge only unlisted leftovers; fail-open (cannot block). Effect-based cousin: fast-allow/deny <1ms then Jev remainder, **fail-closed** on execution (construct-auto-classifier). **Wrap cousin:** rules first then Jev remainder, ASK throws, **fail-closed** on execution (AgentGhost) | `references/mappings.md#18-structural-prove--soft-remainder-hypothesis-as-domain-general-empirical-as-named-shapes` (**Hypothesis**; jevgate/OCR shapes Empirical; construct Empirical as certification; AgentGhost Empirical as README wrap) | | Effect-oriented state-machine loops | Soft predicates on transitions; code owns the transition | `references/mappings.md#19-effect-oriented-state-machine-loops-hypothesis` (**Hypothesis**; ZIO client, not Effect.ts) | -| Agent self-supervision / on-track detection | Pre-gate → output judge → done-check → supervisor nouls | `references/agent-self-assessment.md` | -| Optimizer/program frameworks (Ax, DSPy) | Typed fields → one provider request; judge metrics; threshold discipline. Ax and DSPy climb LM-program knobs only | `references/optimizer-integration.md` | -| Perception → decision pipeline / measure / hill-climb | Stages with a versioned state contract; frozen taskset; DSPy/Ax only on the LM-program slice. **Hypothesis**. Same section as the row below | `references/validation.md#eval--hill-climb` | -| Eval & hill-climb | Decision-stage jevals hygiene; Harbor taskset × harness × runtime; one score-composition table | `references/validation.md#eval--hill-climb` | +| Agent self-supervision / on-track detection | Pre-gate → output judge → done-check → supervisor nouls. S1 reflex keeps control; optional S2 is one-use advice (khordoo delta: never stall; log consumption; Local ≠ localjev; 20% still soft). **OCR+AX desktop CU:** typesafe-computer-use (never screenshot-to-frontier for the decision; `done` ≠ success; 0.4/0.5 still soft). **ASR voice-browser:** jev-voice-browser (never waveform-to-Jev; spoken confirm ≠ auth). Claim/evidence Stop (anti-hallucinated-done); bounded Pi supervisor (shadow recovery, never generates commands). OMP/pi `jev_acceptance_gate` + `jev_route` fail-open (confidence: 0 on missing Jev). OMP prompt suppression: operator owns bar; host deny stays above (omp-greenlight). Closed-vote CU: no planner LLM (JevOnly); host-owned product (waymode). Effect-based pre-gate fail-closed (construct-auto-classifier). Stop-hook attention filter: never blocks the agent; never green unless sure (jev-lens). Runtime authorize: Jev supplies evidence, code owns authority (actiongate-jev; fail-closed on deterministic security failure). Turnstile clone: observe/enforce + replay; Jev never grants what policy denied. SEAL: generation fills Candidates; only a Seal advances. jev-labs: quorum+stability; escalate when evidence degrades. jev-handoff: typed escalate/continue/abort; inverted loop; fail-open. hermes-jev-router: WHETHER/HOW/WHAT; skip next main-model when evidence is enough. browser-jev: Playwright executes, Jev chooses; sample-from-distribution. jeffrey: Jev→tool→Jev; LLM fills args. pi-heed: persist user constraints across compaction; check side-effecting calls. jev-decisions: reviews are advice, never stop commands. **jev-use:** fail-open PreToolUse gate; Vercel margin fallback. **pi-jev-control:** control-plane sensors; GUI unknown never force-click. **Wrap-as-execution:** ALLOW/ASK/DENY on the actuator path; ASK throws; fail-closed (AgentGhost; **≠** jev-use fail-open; rh-guard owns the gate cousin) | `references/agent-self-assessment.md` | +| Optimizer/program frameworks (Ax, DSPy) | Typed fields → one provider request; judge metrics; threshold discipline. Ax and DSPy climb LM-program knobs only. Typed control plane sits *around* DSPy (ontology/security/confidence/state/allow-list); DSPy drafts AFTER route+action | `references/optimizer-integration.md` | +| Perception → decision pipeline / measure / hill-climb | Stages with a versioned state contract; frozen taskset; DSPy/Ax only on the LM-program slice. Specialist composition stays **Hypothesis**; open multimodal decide (blackwood-rlcd) is a named receipt. Structured observe→decide→verified-act (no screenshots) is a computer-use speed-layer receipt (Jev or GLiNER2 or Cua-S1 specialist; Cua-S1 source-only, not TypeSafe Jev). Stagehand experimental Jev is the same job inside a major harness (pick-and-copy extract; 37/75 no-LLM ~0.5s vs 4.37s is *their* card; pick ≠ replacement). **OCR+AX desktop:** typesafe-computer-use (MIT **427★**; hosted Jev; $0.0002/155× *theirs* one screenshot, not a taskset). **ASR voice-browser:** jev-voice-browser (MIT **103★**; 27/27 fixtures *theirs*). Same section as the row below | `references/validation.md#eval--hill-climb` | +| Eval & hill-climb | Decision-stage jevals hygiene; Harbor taskset × harness × runtime; one score-composition table. Shared bake-off exemplar: open-jev-laya-bench (ECE/NLL/Brier; LLM-as-judge is not the score). Harbor-style frozen protocol vs constrained LLMs: DMB (accuracy/calibration/latency/cost; raw logs). Feedstock: jevals-data CC-BY-4.0 boards + JSONL (recompute-from-logs). Collab-arm curriculum: llm_autonomous vs scripted_plus_jev vs llm_plus_jev (Wilson / McNemar). Negative: combinatorial grid assembly ≠ extractive (ARC-AGI Direct Jev 4/400). Harbor on/off routing: chess-engine tasks, hidden perft verifier, one-run preliminary (jev-gateway-bench). Pair CI merge-gate with Harbor + rh-guard. Harbor needle/noise stdout prune: jev-pruner (manual sweep theirs; plugin eval cannot reach Jev → fail-safe original; Terminal-Bench pilot is integration not a full bench). Cua-S1 specialist form: source-only (metric names, no checkpoint scores; not TypeSafe Jev). Stagehand extract pick-and-copy: 37/75 no-LLM ~0.5s vs baseline 4.37s (*their* 25×3; pick is a fast path, not a replacement; draft stack #2951–#2955). Pre-registered independent eval: jev-baselines-eval (**both AMBIGUOUS**; cascade sign-flip at exact parity; confidence=1.0 theater; encoder-with-labels wins; serving-path ≠ model-speed). Healthcare Harbor-shaped: explore-typesafe-ai (synthetic FHIR; not clinically validated). Honest-negative PDF: databricks-jev-pdf-lab (no quality-equivalent Jev payoff). Meaning-search Harbor-shaped: jevgrep 79% top-5 vs BM25 40% / grep 20% on stripped repos (keyword still wins exact strings). Measured RAG rerank one-run: Jev-RAG ≥70% cost / 72% latency vs Spark rerank (full-context Spark still faster). jevals-shaped oxlint: Phoenix fixtures agree with the human answer key. Native-probability calibration arena (jev-arena Brier 0.0059 / ECE 0.0620 *theirs*; overconfident in low bins; fan-out 2 requests). Typed control-plane bake-off metrics (dspy-control-plane; offline stubs ≠ quality). Tetris Jev vs Haiku demo (not a rigorous eval). Fan-out suite: sonar heatmap-as-policy; vickrey CDF then code bids; bracket Brier vs Elo (live trailed Elo). Domain specialist vs few-shot hosted: Domain-jev-maker matched-precision KL / r / McNemar (independent CLINC gold). Cascade compare arms: jav-email-cascade jev vs gen-json vs gen-logprob (mock gen-json flat-confidence is *their mock*). ORDER BY ranking family: jev-orderby-bench six gates; Score ordinal 0.143 weak link; 53-way 0.99 tie; recodelabs batch-40 fails ranking. Class-backend economics: jeff GLiFormer ~$2.6 vs ~$15.6 (~6×) L4 HTTP; A10G direct ~$0.65 (~24×); AG News 75.5% vs 90.5%. Harbor Jev vs local MLX PCD vs AR JSON: system-one-benchmark toxic-chat n=50 Jev 84.0% / Brier 0.1096 vs PCD 52% / Brier 0.3884 (O(1) speed ≠ calibrated Noul). Evidence-packet explorer: 1/8→6/8 SWE-bench Verified finish *theirs* (n=8, empty-as-miss, author-run). Meaning-grep LLM-as-judge: jev-semgrep precision 0.94 / recall 0.98 *theirs* (10×51; not Harbor; dedicated §86; **51★** ephemeral). OMP prompt-suppression: omp-greenlight 1,013 calls / 10 sessions; default **40.9%** prompts removed / **0 of 94** unsafe auto-approvals on labelled corpus (operator owns bar). Eval-instrument hygiene: dinostomp `jev` if-statement tests (accuracy / p(yes) cut / ECE / blank lean / rewording); FINDINGS 189 / 99 against itself. SLO routing latency cost: slo-router p95 **77.93 → 490.38 ms**, same routes/accuracy *theirs* (eight-row demo is not a benchmark). Effect-gate certification: construct-auto-classifier Jev **0** dangerous / 975; every chat model leaked. INSTRUCT_JEV 119-row Choice/Noul/Score instruct seed. **Evidence-gated packs:** jev-packs nine verified packs (accuracy/ECE/cost/latency on pinned jev-1.13; no numbers, no endorsement). **Ranking ≠ calibration:** does-jev-confidence 8,000 human-annotated judgments; AUC ~0.91; stated ~75% vs human ~10%; two-parameter recalibration removes ~96% ECE without changing rank. **Hot-click CU:** ego-jev HN/wiki n=3 medians ~2× vs per-step LLM (high variance; not a bench). **Verbatim compact vs summarize:** jev-compactor 64.5% / 366 ms / $0.0004 / 0 hallucinated paths / 4 of 4 facts vs Sonnet summary 96.2% / 6.1 s / 1 invented path (one synthetic session). **Eval integrity cluster (no leaderboard theater):** atlas receipts (type-safe ≠ correct); jev-frontier-100 Jev 77.0% vs Qwen3.5 4B/2048 96.7% (budget-attached; exploratory); jev-ood-calibration 900 tickets ECE 0.107 = 4.4× floor, Choice/Score T~3.3 vs boolean T 0.66. **JevBench v1.1:** Capability/Speed/Cost → Main Score; calibration reported not scored; native vs verbalized; partial runs not ranked. **Never confidently wrong protocol:** jev-labs 1,080 golden 0 wrong under chaos (escalate; not a proof of zero). **Sureness 60-q:** Choice confidence = max_prob. **Pre-send views:** dizk/jev-lens 79% fewer tokens / 500 SWE-rebench *theirs*. **Product-arm compact:** jev-compactor 73% / 350 ms / 4 of 4 vs shipped summarizers. **tools≠use:** carryforward 0/4. **openvons** LM 0.916 vs 27B 0.875 / JevPick 3.2–4.8× *theirs*. **jevcache** 0 FP / precision 1 / recall 0.38 / n=100 *theirs*. **zeroshot-vs-bert** +0.05–+0.13 vs DeBERTa-c; contamination DiD; ~230 labels *theirs*. **ThinkyMiner/Winnow** 80%/90%. **OpenJev** 45/60. **semif-serve** 1164 vs 178 ms. **pi-jev-skill-bench** harness (43 gold; no live numbers this pass). **typed-evaluation-collapse** Noul band vs Choice p=1.0. **DGUI_HYPERMEM-JEV** 6-row flywheel. **jevassert landed** record/replay CI (accuracy/ECE/Brier/cost/latency; exit 0/1/2; McNemar). **jev-packs pairing** 2,990-case matrix: Jev/Sonnet 5 accuracy tie, Jev better calibrated 7/9, ~250× cheaper *theirs*. **jevarena** failure-finding (≠ jev-arena; harness not findings). **BBQ** 58,492 / 97.28% / $0.3429 *theirs*. **jevlint** 13/15 1.00/1.00 *theirs*. **grande** JGLUE 0.614/0.853 + 270M 0.710/0.710 *theirs*. **local-jev** done 30%/shape 57% vs Jev. **pi-heed** 98.5%/0 false block *theirs*. TeoMastro summary.md 404 this pass; **Harbor SGR-judge contract** (jev-judge-bench; 21 offline tests; canaries ≠ quality; **no quality headline yet**; ≠ jevarena/jevbench); **cookbook samples not benches** (jev-cookbook 16–36; 425/$0.015); **competing NAR claim-audit** (openJev-verdict-2.0 77.10%/0.0636/0.0144 *theirs* unverified; PR #1 throughput≠latency / Laya parity / like-for-like ECE; ≠ IamBusy/OpenJev); **chakuho GUI 336** 27B 95%/92% vs Jev 89%/82% *theirs* (coverage ≠ correctness); **jevinf** 2.57×/2.27× 100% argmax (not ECE); **jevex n=16** 160s→69s / $8.74→$3.13 / 16/16 both arms *theirs*; **commitjev** 13 labelled / 0 false on 5 clean *theirs* (small control); **laya-multilingual** MASSIVE 0.366/0.387 vs 0.227/0.733; ECE 0.314→0.106 after T *theirs*; **schema-scorer** v2 Choice 0.841; **open-jev-laya-bench / jev-tree-choice-cap / INSTRUCT_JEV HF 401** this pass; jevlogs GitHub 404 + HF 401; **classifier.dev vs_jev** tracked JSON (F1 0.887 / 230–232 ms; AG News 87.7%; granite 0.546 vs advertised 0.800 *theirs*; n=7 train-on-test; read eval/README); **githubnext/localjev** 1,200-req prompted-JSON bake-off (Qwen3.6 76.7% / Gemma 26B 75.0% / DiffusionGemma 74.2% short *theirs*; not logits; not calibrated; ≠ kunchenguid/local-jev); **NandhaKishorM/laya vs-Jev** third-party unpublished-here (Banking77 0.425 vs Jev 0.870; post-T ECE 0.081 vs 0.246; 0.766 fine-tune not zero-shot; T4 32.8 ms *theirs*); **external openjev census** (tweet ≠ v1.1 GitHub artifact; watch jev-models; likes ephemeral); **JevBench v1.2** (534 decisions; geo-mean; Jev 75.3 / SemIf 74.6 *theirs*; cal ON; v1.1 87.6 not comparable; option-order 72→21; ×2 latency assumption; many costs est.; Laya absent gap); **OCR+AX desktop CU cost table:** typesafe-computer-use $0.0002 vs Opus $0.032 (155×) *theirs* one screenshot; 0.4/0.5 still soft; **≠** Harbor taskset; **ASR voice-browser fixtures:** jev-voice-browser 27/27 / ~300 ms / ~$0.0002 *theirs* not a Harbor taskset; **wrap-as-execution (not a quality bench):** AgentGhost MIT **2★**; ASK throws; fail-closed; no Harbor numbers; **JP genre atlas (tweet, not scores):** @studio_yebisu stars research-time; not verified evals; likes ephemeral; **external pedagogy (article, not scores):** @akshay_pachaar 200×/400× TypeSafe ceiling; schema-safe ≠ correct; likes ephemeral | `references/validation.md#eval--hill-climb` | | (meta) Finding new mappings & applications | Toolbox sweep: judgment-shaped component of a known method, substituted + falsified | `references/toolbox-mapping.md` | | Named methods / operators / theorems | Substitution tiers: operand-judgments, preconditioned theorems, non-substitutable | `references/methods-catalog.md` | | (meta) Where a judgment model sits relative to any construct | 11 positions + logical-operator rules + position×construct traversal as the application generator | `references/composition-algebra.md` | -| Question mechanics & debugging | Instruction/criteria/state shape, budgets, diagnosis table, revision discipline | `references/question-design.md` | +| Question mechanics & debugging | Instruction/criteria/state shape, budgets, diagnosis table, revision discipline. Missing `other` → confident wrong Choice (confidence gating cannot catch); lint the request (`wellposed` recipe; `tenbin` owns the skill) | `references/question-design.md` | | Heuristic search over a taxonomy | Parallel beam over Choice distributions | `references/mappings.md#5-hierarchy--bounded-heuristic-search` | | Existing-system insertion / code-smell audit | Opportunity map, fit test, smallest boundary, policy centralization | `references/boundary-audit.md` | @@ -202,7 +213,7 @@ Desired behavior and non-judgment baseline: Semantic judgment(s) and what each output means: Pillar (EU / VOI / MCDA / SDT / search / safety / formal): Hole (sieve / keep-drop / triage / rank / route / gate / perceive / abstain / gather): -Family (closed decision API / open head / encoder open-jev / constrained-AR surface / GLiNER locate / GLiClass categorize / listwise ranker / vision scorer): +Family (closed decision API / open head / open multimodal RLCD / encoder open-jev / constrained-AR surface / GLiNER locate / GLiClass categorize / listwise ranker / vision scorer): Evidence/candidate source and known coverage gaps: Deterministic policy, constraints, and action ownership: Batchable vs genuinely dependent steps: diff --git a/.agents/skills/augustus/references/agent-self-assessment.md b/.agents/skills/augustus/references/agent-self-assessment.md index d36cbda..f6dd169 100644 --- a/.agents/skills/augustus/references/agent-self-assessment.md +++ b/.agents/skills/augustus/references/agent-self-assessment.md @@ -12,7 +12,87 @@ relying: foreman, pi-jev, pi-warden, winnow, fast-jev-compaction, jev-judgment. parallel, thresholded in code: `destructive` (noul, hold ≥0.90), `exfiltrates_secrets` (≥0.70), `beyond_request_scope` (≥0.85), `impact_if_unwanted` (Score 4 levels, escalate ≥2.5). Ship shadow mode - first; enforce only after observing real traffic. + first; enforce only after observing real traffic. Productized pre-exec + cousin: [toolgate](https://github.com/fdemir/toolgate) — `allow` / + `block` / `review` before execution; guard error/timeout **stops**. + Jev is not authorization. 72-case synthetic, not independently + annotated. Distinct from the ndolinschi *vocabulary* (allow / + ask_human / deny) below (`notes.md` §55). + **Capability kernel, different trust model (2026-09-18 ~19:48):** + [interlock](https://github.com/somoore/interlock) — secrets never + enter the agent; closed action space; Jev is SENSOR; `policy.py` + decides BLOCK/ASK/ALLOW. Do not merge with toolgate. Type-safe ≠ + correct (`notes.md` §59). + **Human-confirmed cousin:** + [port-cleanup](https://github.com/epiphany-dynamics/port-cleanup) + — Jev recommends; human is the only kill trigger; identity + re-check; shields override; mapped explanations (`notes.md` §59). + **OMP prompt suppression (permission vs probability, + 2026-09-18 ~22:38):** + [omp-greenlight](https://github.com/SemetricLabs/omp-greenlight) + — auto-approve is a criterion the **operator** owns; Jev + is not a grant. Not a sandbox (`notes.md` §62). + **Effect-based shell pre-gate (2026-09-18 ~23:40):** + [construct-auto-classifier](https://github.com/godspede/construct-auto-classifier) + — fast-allow/deny then Jev Choice + independent risk + Nouls. Privilege ≠ verdict. Fail-closed on missing / + low-conf / high-risk. Jev 0 dangerous / 975 *theirs*. + Distinct from toolgate / greenlight / interlock + (`notes.md` §63). Landed-script trust / headless ≠ + auto-approve (`notes.md` §68). + **Runtime authorize — evidence ≠ authority + (2026-09-19 ~00:39):** + [actiongate-jev](https://github.com/omkarghugarkar007/actiongate-jev) + — Jev supplies evidence; code owns ALLOW/REVIEW/BLOCK. + Positive score never overrides a deterministic security + failure. Fail-closed on financial/destructive/credential + if Jev is down (`notes.md` §64). + **Persist constraints across compaction (2026-09-19 + ~05:46):** + [pi-heed](https://github.com/Nyarlathoteppppp/pi-heed) + — conversational policy as structured state; replayed + after compaction without calling Jev again. Jev + classifies KEEP/LIFT/…; **never writes policy**. + Side-effecting calls checked before they run. Fail-open. + Shadow default. *Theirs:* recall 98.5% / false block + 0.0% / $0.000058. Distinct from actiongate (RBAC/schema) + (`notes.md` §70). + **jev-use PreToolUse gate (2026-09-19 ~06:43):** + [jev-use](https://github.com/shitianfang/jev-use) + — deny/ask; **fail-open**; 12/12 *theirs*; Vercel + reconstructs confidence as margin (default 0.4). Gate + never grants (`notes.md` §71). + **Pi control-plane gates (2026-09-19 ~06:43):** + [pi-jev-control](https://github.com/goodruizhan/pi-jev-control) + — deterministic fast-path then Jev; GUI < threshold + → unknown, never force-click. License null + (`notes.md` §71). + **Turnstile clone (2026-09-19 ~01:47):** + [turnstile](https://github.com/zyphr-labs/turnstile) — + policy first; Jev remainder; receipts + replay; missing + Jev → Review. Experimental alpha. Same doctrine as + actiongate (`notes.md` §66). + **Never-confidently-wrong consensus (2026-09-19 + ~02:38):** + [jev-labs](https://github.com/copyleftdev/jev-labs) + — TLA+ protocol; Jev oracle; escalate when unstable. + 1,080 golden 0 wrong *theirs*; not a proof of zero + (`notes.md` §67). + **Advance/coverage (2026-09-19 ~02:38):** + [seal](https://github.com/Reasonofmoon/seal) + — generation fills Candidates; only a Seal advances; + coverage.path visible; mint ≠ product brain + (`notes.md` §67). + **Email / ticket leftover cascade (2026-09-18 ~20:43):** + [jav-email-cascade](https://github.com/skiingfalcon/jav-email-cascade) + — typed decide, policy auto/review/llm, generator only on + leftover text. Noul 0.5 never rounded into auto. Distinct + from dual-process-ai (routing unmeasured). `notes.md` §60. + **Closed-vote / host-owned CU (2026-09-18 ~21:39):** + [JevOnly](https://github.com/buluoray/JevOnly) — no planner + LLM; code builds options; Jev only picks. [waymode](https://github.com/mossburgh/waymode) + — host retains handlers/permissions; Jev over live typed + actions; `completed` ≠ server-state success. `notes.md` §61. 2. **Post-action output judge** (after the tool result exists, not before): `leaks_secret` (noul ≥0.90) and `failure_class` (Choice ~6 options). The gate sees intent; only the output judge sees what the command printed. @@ -22,23 +102,216 @@ relying: foreman, pi-jev, pi-warden, winnow, fast-jev-compaction, jev-judgment. answers it — one Noul only for the semantic remainder ("does this reply claim the work is finished?"), one threshold. Spending the model on the countable half is the `/bin/ls`-as-first-tier pattern - (`mappings.md` §18). + (`mappings.md` §18). Claim/evidence cousin + ([clear-head](https://github.com/VladyslavHontar/clear-head), + ~16:48): check factual claims against **what was actually read this + session**; keyword retriever, not semantic; below `JEV_FIRM` 0.6 + never blocks; true-but-unread still flags unsupported + (`notes.md` §51). Anti-hallucinated-done, not a test runner. + **Attention-filter Stop (never blocks the agent; + 2026-09-18 ~23:40):** + [jev-lens](https://github.com/rashedInt32/jev-lens) — + calibrated “do I need to look?” in the background; the + Stop hook returns at once. Never edits files; never says + green unless sure. **Hunch:** VOI for human review, not a + permission gate. Distinct from jev-gates (stops writes) + (`notes.md` §63). Distinct from + [dizk/jev-lens](https://github.com/dizk/jev-lens) + (pre-send views; 79% fewer tokens *theirs*) and from + [jev-preflight](https://github.com/muse0509/jev-preflight) + (agent attention redirect; fail-open; not a merge + blocker) (`notes.md` §68). + Computer-use cousin of the same honesty: + [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) + — loop `DONE` is termination, not verified success; apps inspect + the actual result (`notes.md` §52). 4. **Stuck-detector**: three failures with the same strategy → ask for a - new hypothesis, not another retry. + new hypothesis, not another retry. **Silence is not safer:** a draft + gate that treats a missing Jev answer as "don't send" holds forever. + Missing verdict needs a fail-open / heartbeat — not a block and not a + pass. Contrast Abide `<0.5` silence (the *edit proceeds*). Showcase + class pattern on [jevable.com](https://jevable.com/) (`notes.md` §56). + **Stuck ladder cousin (2026-09-19 ~05:46):** + [jeffrey](https://github.com/thomasbrueggemann/jeffrey) — withhold + the looping tool, re-ask Jev (2 Jev calls / 0 steps per recovery). + Jev owns stuck/progress/done; LLM only fills args. `notes.md` §70. 5. **Supervision during long runs** (foreman): separate concurrent loop estimates `meaningful_progress`, `implementation_complete`, `tests_sufficient`, `worker_stuck`, `work_off_track`, `ready_to_finish`; deterministic policy with hysteresis (retry counts, verification history) gates continue/stop/retry/verify. The model never - commands; it estimates named probabilities. + commands; it estimates named probabilities. Same split as + [jev-reflex-autonomy-lab](https://github.com/khordoo/jev-reflex-autonomy-lab): + **S1 keeps control**; optional S2 is one-use advice on low confidence + and does not fly the drone (`notes.md` §46). **Delta (`notes.md` + §80):** escalate-under-threshold **without stalling**; purple + confidence = that Jev decision **consumed** returned S2 (purple + S2 bar = arrival; red = fail); Local controller is rule-based + **≠** githubnext/localjev; 20% starting gate still soft; no + pixels to either provider; seed = geometry not async replay. + Experimental viz, not a production supervisor. Computer-use speed layer of the same split: + [solari-reflex](https://github.com/hitakshiA/solari-reflex) — one + structured observation → one typed decision → one verified action; + **no screenshots**; model output never becomes a selector. Harbor-style + task score (Stripe API / answer key). Author table vs Codex on Solari: + 60.2 s vs 194.9 s; 66 s vs 460 s; 24.2 s vs 98.4 s (`notes.md` §48). + Encoder backend of the same hole: + [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) + — local GLiNER2 scores observed controls; `DONE` ≠ verified success + (`notes.md` §52). + Specialist-form cousin, **not TypeSafe Jev:** + [Cua-S1](https://github.com/trycua/cua/tree/main/libs/cua-s1) — + option-attention among observed elements; plan ≠ execute; dry-run + default; source-only (`notes.md` §54). + Harness cousin (draft stack): + [Stagehand #2955](https://github.com/browserbase/stagehand/pull/2955) + — pick-and-copy extract + act tree; LLM fallback; pick ≠ + replacement (`notes.md` §57). + OCR+AX desktop product (hosted Jev; MIT **427★**): + [typesafe-computer-use](https://github.com/awlevin/typesafe-computer-use) + — never ships a screenshot for the *decision*; + overlapping options = doubt; `done` ≠ verified + success; 0.4 / 0.5 still soft (`notes.md` §81). + ASR voice-browser product (hosted Jev; MIT **103★**): + [jev-voice-browser](https://github.com/moritzkremb/jev-voice-browser) + — never ships a waveform; partial-speech wait; + spoken confirm ≠ auth; numbered overlay, no second + model (`notes.md` §82). + Wrap-as-execution product (MIT **2★**): + [AgentGhost](https://github.com/reddpy/AgentGhost) + — ALLOW/ASK/DENY *is* the tool function; rules + first; ASK throws; fail-closed; judge swappable + (`notes.md` §83). **≠** actiongate **≠** toolgate + **≠** jev-use. rh-guard owns the gate cousin. + External pedagogy (not a product): + [@akshay_pachaar](https://x.com/akshay_pachaar/status/2101037514945597645) + — LLM hammer; schema-safe ≠ correct; shadow + first; 200×/400× TypeSafe ceiling (`notes.md` + §85). **≠** official docs **≠** Flavio. + Meaning-grep is **not** a self-supervision gate: + [jev-semgrep](https://github.com/uehaj/jev-semgrep) + ranks lines; rh-guard skip (`notes.md` §86). + Closed-vote extreme (no planner LLM): + [JevOnly](https://github.com/buluoray/JevOnly) (`notes.md` §61). + Host-owned product: + [waymode](https://github.com/mossburgh/waymode) (`notes.md` §61). + Productized Kahneman cascade for *any* cheap-decide / expensive-write + loop (business/life, not only SWE): + [dual-process-ai](https://github.com/taro1985/dual-process-ai) — + `conf ≥ τ` S1 decides else S2 generates; routing fails open; safety + fails closed; **routing accuracy not measured**; keyword fallback is + not S1 (`notes.md` §49). Tune τ on your escalation log. + **Distinguish names:** the existing "foreman" *shape* here is the + Kevthetech143/super-jev loop (named probabilities → hysteresis + table). [`reification-labs/foreman`](https://github.com/reification-labs/foreman) + (~16:48) is a **description-only** Elixir/Phoenix scaffold claiming + parallel S1 specialists + one S2 coordinator with typed + `{value, probability}` — README is stock Phoenix; `mix.exs` has no + Jev dep. Do not invent an Elixir API (`notes.md` §51). + Bounded Pi supervisor of the same lifecycle: + [jevons](https://github.com/LilDojd/jevons) — not a second agent; + default recovery **shadow**; steering never generates commands. + Distinguish from pi-jev-approver (fail-closed remainder) and + pi-jev-context (sieve). + **OMP/pi fail-open cousins (2026-09-18 ~21:39):** + [omp-jev-extensions](https://github.com/luw2007/omp-jev-extensions) + — `jev_acceptance_gate` before done; `jev_route` topology/tier. + Missing Jev **fails open** (`confidence: 0`). Contrast + pi-jev-approver fail-closed without a key (`notes.md` §61). + **OMP prompt suppression (2026-09-18 ~22:38):** + [omp-greenlight](https://github.com/SemetricLabs/omp-greenlight) + — grades gated calls; suppresses the prompt when Jev says + allow. Operator owns the bar; plugin never self-tunes. + Default 40.9% / 0 of 94 *theirs*. Not a sandbox. Host deny + fires first (`notes.md` §62). + **Skill pack grant (outline only):** + [skill-broker](https://github.com/adamjralph/skill-broker) + — Jev scores relevance; code owns grants; never broaden + access. Not a production recipe (`notes.md` §62). 6. **Context economy**: the context-sieve card (`references/applied-mappings.md#1-context-sieve`). Judge every large tool result with one relevance Noul before it enters context. Hide confident-no blocks behind a stub + recall key; always keep current - instruction, recent turns, errors, and opaque blocks. winnow hides at - relevance ≤0.22; fast-jev-compaction asks two nouls per tool call - (should the call stay knowing it was made? should the result stay - verbatim?). + instruction, recent turns, errors, and opaque blocks. [kevinpita/winnow](https://github.com/kevinpita/winnow) hides at + relevance ≤0.22 (agent context sieve; distinct from [ThinkyMiner/Winnow](https://github.com/ThinkyMiner/Winnow) worth-your-attention VOI, `notes.md` §69); fast-jev-compaction asks two nouls per tool call + (should the call stay knowing it was made? should the result stay + verbatim?). Encoder-backend cousin: + [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) + — GLiNER2.5 retention Choice + exact character-offset copies; mutating + tools stay `keep_full`; low-confidence fails closed to `keep_full`; + `shadowMode` default true. Not a summarizer. Not Jev (`notes.md` §50). + Stdout-prune cousin, same family, different job: + [jev-pruner](https://github.com/tamaratran/jev-pruner) — Jev Noul on + Bash chunks after a hard envelope; fail-safe original; archive + (`notes.md` §53). Marketplace id still `fast-jev-output`. + Session-ledger cousin: + [carryforward](https://github.com/Dharundp6/jev-carryforward) — + verbatim JSONL; Jev scores which facts are still live; constraints + and corrections always return; fail-open dump if the scorer is + down. Nine entries × three tasks is a hint, not proof + (`notes.md` §55). Eval: `recall` **0/4** — SessionStart + hook > hoping (`notes.md` §68). Do not copy mcp add. + Observational-memory sibling: + [pi-observational-memory-jev](https://github.com/willfish/pi-observational-memory-jev) + — keep/kind verbatim; model-free compact (`notes.md` + §68). + Pre-send cousin: + [jev-lens](https://github.com/dizk/jev-lens) — views + before first send; 79% fewer tokens *theirs* + (`notes.md` §68). + Classify-first cousin: + [jev-sift](https://github.com/kbhuw/jev-sift) — batch path/url/text + → Jev **before** the main agent reads; uncertain/errors/truncation + ≠ irrelevant. Transport tests ≠ accuracy. No LICENSE this pass + (`notes.md` §56). Do not copy plugin how-to. + Framework-agnostic compact+gate cousin + (2026-09-19 ~00:39; was empty skip §61): + [jev-compactor](https://github.com/edwardyen724-g/jev-compactor) + — Jev judges relevance; code decides structure; never rewrite; + regex floor independent of Jev. Compaction fail-open if Jev + down; safety fail-closed on pending destructive. OpenCode + fail-open port already §62: fast-jev-opencode (`notes.md` §65). + Later product-arm bench **73%** / 350 ms / 4 of 4 + (`notes.md` §68). + WHETHER/HOW/WHAT cousin (license null; 2026-09-19 + ~04:39): + [hermes-jev-router](https://github.com/rsdkrasen/hermes-jev-router) + — compact original chunks; skip next main-model when + evidence is enough (needs core patch); fail-open + (`notes.md` §69). + Pi summarizer-replacement cousin: + [pi-jev-compact](https://github.com/dev-willbird1936/pi-jev-compact) + — verbatim keep/drop of paired tool calls; fail-open to + LLM summary if <25% saved. **≠** + vava-nessa/pi-jev-compaction (`notes.md` §72). + Typed baton cousin: + [jev-handoff](https://github.com/shitianfang/jev-handoff) + — escalate/continue/abort; gate never grants; inverted + loop (`notes.md` §69). + Advice-only cousin: + [jev-decisions](https://github.com/bojansandhaus/jev-decisions) + — 25 prepared reviews; **never stop commands** + (`notes.md` §69). + Adversarial-browser cousin: + [browser-jev](https://github.com/DowLucas/browser-jev) + — Playwright executes, Jev chooses; sample from the + distribution (`notes.md` §69). + Decider≠executor cousin: + [jeffrey](https://github.com/thomasbrueggemann/jeffrey) + — Jev→tool→Jev; LLM fills args; risk≥0.5 pause + (`notes.md` §70). + Persist-constraints cousin: + [pi-heed](https://github.com/Nyarlathoteppppp/pi-heed) + — user constraints survive compaction; Jev never + writes policy (`notes.md` §70). + Plugin no-text cousin (2026-09-19 ~06:43): + [jev-use](https://github.com/shitianfang/jev-use) + — fail-open PreToolUse gate; Vercel margin fallback + (`notes.md` §71). + Pi control-plane cousin: + [pi-jev-control](https://github.com/goodruizhan/pi-jev-control) + — named sensors; GUI unknown never force-click + (`notes.md` §71). ## Non-negotiable boundaries @@ -51,6 +324,23 @@ relying: foreman, pi-jev, pi-warden, winnow, fast-jev-compaction, jev-judgment. only safe because a hard interlock or sandbox sits underneath. A gate that *selects* or *authorizes* a side effect fails closed instead (`mixed-architecture.md` prefilter table; `mappings.md` §18). + Compaction *drop* is that second kind: + [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) + fails closed to `keep_full` (`notes.md` §50). Stdout prune is the + same polarity: + [jev-pruner](https://github.com/tamaratran/jev-pruner) fails closed + to original output (`notes.md` §53). Tool *execution* is the + other polarity: [toolgate](https://github.com/fdemir/toolgate) + stops on block / review-without-approval / guard error + (`notes.md` §55). Capability kernel + ([interlock](https://github.com/somoore/interlock)) never lets + the secret into the agent in the first place (`notes.md` §59). + Human-confirmed kill + ([port-cleanup](https://github.com/epiphany-dynamics/port-cleanup)) + re-checks identity before SIGTERM. Session-memory *omit* fails open (dump the + ledger): [carryforward](https://github.com/Dharundp6/jev-carryforward). + Draft-gate *silence* is the opposite mistake: treating no-answer as + a hold. Missing verdict needs a heartbeat (`notes.md` §56). - Cache identical judgments (~120s) and deduplicate sibling calls into one in-flight request. - pi-warden measured cost makes continuous guarding viable: ~$0.00004 and @@ -62,6 +352,20 @@ relying: foreman, pi-jev, pi-warden, winnow, fast-jev-compaction, jev-judgment. - Grounding of a generated claim: one Choice per claim–evidence pair (supports / contradicts / unrelated) + a confidence review flag; judge against the cited source text, never against another model's prose. + Pointer-not-generator: the model points at line ids; code copies + verbatim with place; *not found* is an answer + ([choxos/jev-reviewer](https://github.com/choxos/jev-reviewer); + systematic-review Jev Reviewer; **12★**; human tick never + overwritten; `notes.md` §48, §74). + Distinct: [egma-ai/jev-reviewer](https://github.com/egma-ai/jev-reviewer) + assigns **attention** P0/P1/P2, not correctness; OpenAI writes + deltas (`notes.md` §58). + Compaction: point at character offsets in the tool result + ([gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction); + `notes.md` §50). Session-evidence Stop + ([clear-head](https://github.com/VladyslavHontar/clear-head)): claims + vs **what the assistant actually read**; keyword retriever; below + firm-confidence never blocks (`notes.md` §51). - Self-report fidelity: compare the agent's claimed action with its actual trace via decomposed Nouls (right tool? args match schema? result matches call?). Escalate on low confidence; never auto-retry. @@ -75,8 +379,17 @@ When the "judge" is really "does this change violate a rule we already wrote?", do not ask Jev whether the code is good. Load `references/mixed-architecture.md#preference-lint-and-gates`. The transferable contract (`doeixd/jev-pref`): the project defines the rule, Jev classifies -visible evidence, code maps the outcome, the agent acts. Shadow-mode the gate -first; permit remains a separate axis from confidence. +visible evidence, code maps the outcome, the agent acts. +[`coldteadotai/abide`](https://github.com/coldteadotai/abide) is the +fuller productized path of that contract (compile / calibrate / tune / +replay; one Score per rule on the diff; bands + fail-open). Soft +rules → soft judgment; the linter owns hard rules. Shadow-mode the +gate first; permit remains a separate axis from confidence. +`notes.md` §47. Plain-English PR-check cousin +([if-ai](https://github.com/Victor-Casado/if-ai)): one condition + +required min-confidence; fail-closed on error / empty / low +confidence. [jev-marshal](https://github.com/LightningK0ala/jev-marshal) +is Watch / empty repo this pass (`notes.md` §51). ## Using Jev to test and optimize the skill suite itself @@ -101,5 +414,9 @@ first; permit remains a separate axis from confidence. gets this variance check first, over frozen outputs, before its numbers mean anything. - Gate vocabulary is converging across implementations; reuse it rather - than inventing: allow / ask_human / deny (toolgate), ok / retry / + than inventing: allow / ask_human / deny (ndolinschi *vocab*), + allow / block / review ([toolgate](https://github.com/fdemir/toolgate) + *product* — Jev is not authorization; `notes.md` §55), BLOCK / ASK / + ALLOW ([interlock](https://github.com/somoore/interlock) *kernel* — + Jev is SENSOR, policy decides; `notes.md` §59), ok / retry / escalate / stop (harnessjudge). Same shape as the lifecycle gates above. diff --git a/.agents/skills/augustus/references/applied-mappings.md b/.agents/skills/augustus/references/applied-mappings.md index cedb1d9..51ecbc6 100644 --- a/.agents/skills/augustus/references/applied-mappings.md +++ b/.agents/skills/augustus/references/applied-mappings.md @@ -3,7 +3,7 @@ These cards are *where a judgment-class model sits* in running software. They are family-agnostic: the **typed judgment provider** is TypeSafe Jev by default (live docs / `typesafe-ai`); an open Choice/Score/Noul head -(e.g. Laya) is a substitute you must self-eval (`research/notes.md` §18); +(e.g. Laya, kev) is a substitute you must self-eval (`research/notes.md` §18, §45); GLiNER (locate) / GLiClass (categorize) / listwise rankers / vision scorers are cousin species with different objectives (`judgment-class.md`). Do not copy request fields from this file. @@ -36,24 +36,161 @@ code: hide if p ≤ t and not always_keep; stub + recall key fail open on missing verdict → keep ``` -**Example**: winnow hides at relevance ≤0.22; fast-jev-compaction asks two +**Example**: [kevinpita/winnow](https://github.com/kevinpita/winnow) +hides at relevance ≤0.22 (agent **context sieve** — always qualify +the owner; distinct from [ThinkyMiner/Winnow](https://github.com/ThinkyMiner/Winnow) +worth-your-attention VOI, `notes.md` §69); fast-jev-compaction asks two Nouls (should the *call* stay? should the *result* stay verbatim?); `ibrahemid/jevprune` keeps last-N + error signatures in code, then judges the rest per line; `kevinpita/pi-jev-context` hides (does not delete) older Pi history, always-keep user/system/todos, `/jev off` restores. Pi compaction cousins (`tamaratran/fast-jev-compaction`, `vava-nessa/pi-jev-compaction`) keep verbatim drop, never summarize. +**pi host port this hour (Empirical as README + their bench, +2026-09-18 ~19:48):** +[fast-jev-compaction-pi](https://github.com/zaycruz/fast-jev-compaction-pi) +— same verbatim job on pi's `session_before_compact`; fallback to +the built-in summary on any failure. Their large-session card: +compaction ~50× faster than pi's LLM summary; pure mode drops old +calls; `preserveCallInputs` restores commands/paths. Do not copy +`pi install` (`notes.md` §59). +**Encoder backend, same job (Empirical as README behavior, 2026-09-18 +~16:22):** +[gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) +— GLiNER2.5 (`fastino/gliner2.5-base-v1`) chooses +`keep_full` / `keep_evidence` / `keep_call_only` / `drop` and copies +exact character-offset spans. Not a summarizer. Mutating tools / unknown +shell / control operators → `keep_full`. Low-confidence or invalid +evidence **fails closed to `keep_full`** (the reduction is the +irreversible act; from the evidence side this *looks* like keep-on-error). +`shadowMode` defaults true. Characters, not tokens; no published +retention-quality rates (`notes.md` §50). Family: +`judgment-class.md`. Do not copy the plugin. +**Stdout prune, same family, different job (Empirical as README / +evals README, 2026-09-18 ~17:15):** +[jev-pruner](https://github.com/tamaratran/jev-pruner) — after Bash +runs, Jev Noul-prunes stdout chunks **before** the main LLM sees +them; no summary. Hard envelope first (≤10k estimated tokens; +JSON/XML/YAML/diff/binary; whole-document commands untouched), then +soft Noul. Fail-safe keep original on any failure; full archive for +recovery. Marketplace id still `fast-jev-output`. Codex is opt-in +wrapper, not automatic interception. Same author as +fast-jev-compaction; complementary, not a duplicate. Do not copy +the plugin (`notes.md` §53). +**Session-ledger cousin, same family, different job (Empirical as +README behavior, 2026-09-18 ~17:48):** +[carryforward](https://github.com/Dharundp6/jev-carryforward) — +verbatim JSONL facts (`record`); Jev Noul-scores `recall` against +the current task. Nothing summarised or deleted. Constraints and +corrections **always return in full** (Jev never votes on a rule). +Fail-open: no key → whole list. Thresholds 0.60 full / 0.30–0.60 +one line are *theirs*. Nine entries × three tasks is a **hint, not +proof** (`notes.md` §55). Eval finding *theirs*: agent +called `recall` **0/4** with tools available — SessionStart +hook > hoping (`notes.md` §68). Do not copy `mcp add`. +**Classify-first MCP, same family, different job (Empirical as +README / schema, 2026-09-19 ~00:38):** +[jev-sift](https://github.com/kbhuw/jev-sift) — batch path / public +URL / inline text (or a tool description) → Jev relevance or 1–8 +typed questions **before** the main agent reads. Content goes to +the judge without entering main agent context first (paths/URLs). +Uncertain → closer look; errors and truncation ≠ irrelevant. Hard +envelope (theirs): 50 items, 60k char, 2 MB / 20 s, public-IP only, +no JS/cookies/login, PDFs unsupported. Transport tests ≠ accuracy. +No LICENSE this pass. Same retrieve-wide → decide → evidence-set +family as decision-native-rag-skills. Cousins: typesafe-screening-mcp, +kazuhideoki/jev-search, jev-pruner (after Bash), carryforward (ledger +you already hold). Not jev-routing (host adapter). Topology A MCP +(LLM outer loop). Do not copy plugin / `mcpServers` / key-file +how-to (`notes.md` §56). Local teacher-copy for the same hole: [`SargeDev/jev-gate-student-b`](https://huggingface.co/SargeDev/jev-gate-student-b) (Qwen2.5-0.5B LoRA; P(relevant) from yes/no logits; 148,160-row [`jev-distill-corpus`](https://huggingface.co/datasets/SargeDev/jev-distill-corpus); -card: fail-open on errors). Agreement with Jev labels is not independent +card: fail-open on errors; HF card **unchanged** this pass vs +`notes.md` §33). That is System One as a **memory/context +gate**, not an action permit: vector recall → local yes/no → inject or +stub (`notes.md` §33, §44, §55). Agreement with Jev labels is not independent gold (`notes.md` §33). Official cousin: classifying RAG passages cookbook (**Contract**). **Counterexample**: one Noul "is this log useful?" over 3k lines — that is nine judgments pretending to be one. **Test**: recall of must-keep lines (failures, the current instruction); tokens saved; timeout leaves the artifact in context. Fail-open: a false drop loses evidence. +**Framework-agnostic compact + same-pass safety (Empirical as +README + one-session bench; 2026-09-19 ~00:39; was empty +skip §61):** +[jev-compactor](https://github.com/edwardyen724-g/jev-compactor) +— **Jev judges relevance. Code decides structure.** Keep +messages byte-for-byte; never rewrite; tool pairs never +split. Regex floor in code (`rm -rf` / force-push / `DROP +TABLE` / `curl | sh`) independent of Jev. Dual fail +polarity: compaction **fails open** if Jev is down +(history unchanged) unless `failClosed`; pending-action +destructive/exfil **fails closed**. One synthetic 12.7k- +token session *theirs*: earlier vs-Sonnet card **64.5%** / +**366 ms** (`notes.md` §65); later product-arm table +**73%** (53–76%) / **350 ms** / **$0.0004** / **4 of 4** +facts vs shipped summarizers (30–250× cheaper) — two +synthetic sessions, not a survey. Foreman safety in the +same ~300 ms pass. Claude Code shorter path remains +fast-jev-compaction. OpenCode fail-open port: +fast-jev-opencode (§62 MED). Observational-memory sibling: +[pi-observational-memory-jev](https://github.com/willfish/pi-observational-memory-jev) +— keep/kind only; verbatim ledger; model-free compact +(`notes.md` §68). Do not copy npm (`notes.md` §65, §68). +**Pre-send view selection (Empirical as 500-trajectory +bench; 2026-09-19 ~03:38):** +[jev-lens](https://github.com/dizk/jev-lens) — **not** +[rashedInt32/jev-lens](https://github.com/rashedInt32/jev-lens) +(§63 human Stop filter). Code builds outline/focus/ +testlog/… views from the tool result's own lines; Jev +picks the smallest view that still serves the next step; +`recall` restores dropped lines. 500 SWE-rebench +trajectories *theirs*: **79%** fewer tokens (11.6M → +2.4M); 88% command / 31% code. Compress **before** first +send — post-send prune broke cache and cost **17% more**. +Claude plugin unmeasured. Do not copy npm (`notes.md` +§68). +**Jev WHETHER / Python HOW / LLM WHAT (Empirical as README ++ offline pytest; license null; 2026-09-19 ~04:39):** +[hermes-jev-router](https://github.com/rsdkrasen/hermes-jev-router) +— compaction keeps original chunks (never rewrite); +duplicate observational tools suppressed; skip the next +main-model call when evidence is enough (**needs a Hermes +core patch**). Fail-open. Offline pytest: **2 vs 1** +main-model call pattern. Aggressive defaults. Community +plugin, not vendor. Cousin of dizk/jev-lens + +jev-compactor. Do not copy patch/plugin how-to +(`notes.md` §69). +**tools≠use / SessionStart over hoping (Empirical as +eval finding; 2026-09-19 ~03:38):** +[carryforward](https://github.com/Dharundp6/jev-carryforward) +— with tools + skill installed, the agent called +`recall` **0/4**. SessionStart hook injects rules +unconditionally; an MCP tool sitting there is not enough +(`notes.md` §68). 9×3 remains a hint. +**Empty compaction-proxy skip (2026-09-19 ~06:43 and +~07:49):** +[jev-context-pruner](https://github.com/IPECTER/jev-context-pruner) +— description-only; `contents/` 409 empty. +[jev-runway](https://github.com/IPECTER/jev-runway) — +LICENSE only; created≈pushed 1s; README 404. Sibling of +fast-jev-compaction / jev-compactor / dizk/jev-lens / +jev-pruner. Do not invent files (`notes.md` §71, §72). +**Pi verbatim summarizer replacement (Empirical as +README + latency table; 2026-09-19 ~07:49):** +[pi-jev-compact](https://github.com/dev-willbird1936/pi-jev-compact) +(MIT) — keep-windows/pins in code, then one Noul per +paired tool call; Pi stores original characters, not a +paraphrase. Fail-open to the built-in LLM summary +(off / no key / <25% saved / HTTP error). **Distinct +from** +[pi-jev-compaction](https://github.com/vava-nessa/pi-jev-compaction). +FB-Scanner *theirs*: kept 1 of 261; replay 0.6 s vs +first UI spinner 26 s (host cost, not judge cost). Pair +pi-jev-control / pi-heed. Do not copy `pi install` +(`notes.md` §72). ## 2. Exact-text keep / drop @@ -79,6 +216,166 @@ the subset operation you already had hunks unstaged; lines never split; atomic apply after confirm. Line-by-line search cookbook (**Contract**): score existing line ids, do not generate ids. lizard-agent: pick among visible elements; answers are *located*. +Omni cousin this hour: [blackwood-rlcd](https://huggingface.co/BlackwoodAI/blackwood-rlcd) +picks among **letters drawn on the screenshot**; code still clicks +(`notes.md` §46). Text-only cousin: [jev-e2e](https://github.com/perixtar/jev-e2e) +— Jev selects observed controls; Playwright independently checks; +a confident model cannot substitute for checked expectations. +**Extractive quotes (Empirical as named receipts, 2026-09-18 ~14:52):** +[testimonial-miner](https://github.com/AppitStudio/testimonial-miner) — +code numbers sentences; one broadcast (Choice/Noul/Score + per-sentence +Nouls); the model never writes; `redecide` retunes thresholds on the +log. **[choxos/jev-reviewer](https://github.com/choxos/jev-reviewer)** +(systematic-review Jev Reviewer; MIT; **12★**; +https://jevreviewer.xera.ac; **≠** egma-ai) — the model +**points at line ids**; code copies verbatim quotes with +file/page/row; *Not found* / *Unclear* are answers. Two-pass: +relative Choice (which line?) then absolute Noul (does this line +itself answer?); quotes = Noul ≥ 0.5 *theirs*. Human tick is the +product: checked answers never overwritten (`notes.md` §48, §74). +Claim/evidence Stop cousin +([clear-head](https://github.com/VladyslavHontar/clear-head), ~16:48): +the model judges claims against **keyword-retrieved session lines**, +not against another model's prose; `JEV_FIRM` below 0.6 never blocks +(`notes.md` §51). Compaction cousin +([gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction), +~16:22): the model points at **character offsets** in a tool result; +code copies those bytes; a generator summary is the rejected species +(`notes.md` §50). Stdout-prune cousin: +[jev-pruner](https://github.com/tamaratran/jev-pruner) — the model +scores chunks of observed Bash stdout; code keeps verbatim lines and +archives the rest (`notes.md` §53). Computer-use cousin: +[solari-reflex](https://github.com/hitakshiA/solari-reflex) — structured +observation → typed decision → verified act; **no screenshots**; model +output never becomes a selector (`notes.md` §48). Encoder-backend +cousin of the same hole: +[gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) +— local GLiNER2 (`fastino/gliner2-multi-v1`) scores observed a11y/DOM +controls; code clicks; remote text helper only for TYPE; `DONE` ≠ +verified success (`notes.md` §52). Open-head cousin: +[laya-mind2web](https://huggingface.co/ShaunSpark/laya-mind2web-browser-agent) +— Laya operation + target index over interactive DOM elements (not +screenshot multimodal). Contrast blackwood-rlcd (letters on a +screenshot). Specialist-form cousin, **not TypeSafe Jev:** +[Cua-S1](https://github.com/trycua/cua/tree/main/libs/cua-s1) — +option-attention among observed elements (fill/check/click/skip); +code owns execution order; dry-run default; source-only +(`notes.md` §54). +**Meaning-as-spec (Empirical as README resolver, 2026-09-18 +~19:48):** +[jevcumber](https://github.com/RubyBrewsday/jevcumber) — Cucumber +`.feature` only; no step-definition glue. Jev picks among +**observed controls** and **literals already in the step**; never +writes code or invents values. Lockfile makes replay +deterministic (`--frozen` CI, no key). Refuse below 0.6. Same +pointer family as Stagehand pick-and-copy / jev-e2e. Do not copy +the tarball install (`notes.md` §59). +**Harness pick-and-copy (Empirical as PR-body architecture + their +local eval, 2026-09-19 ~00:48; draft stack):** +[Stagehand #2955](https://github.com/browserbase/stagehand/pull/2955) +(5/5 of [#2951](https://github.com/browserbase/stagehand/pull/2951)–#2955, +all OPEN draft) — Jev **picks** observed a11y elements; **code copies** +text. `extract` `"off"` | `"judge"` | `"pick"`. Judge replaces the +metadata LLM `completed` check (throw → LLM). Pick: schema plan +(scalars / bools-enums / lists of flat objects; else LLM); must +validate + completion gate else LLM; screenshot extract always LLM. +Their card (gemini-3.8-flash, 25×3): **37/75** no-LLM in **~0.5 s** +vs baseline **4.37 s** / two LLM calls; 69/75 vs 23/25 (**92% both**); +LLM-off **36/75** — pick is a **fast path, not a replacement**. +Same observe→score-among-candidates→code-acts *job* as jev-ultrafast / +gliner2-ultrafast / solari-reflex / cua-s1, inside a major harness. +Do not merge clocks. Do not copy `experimentalJevAct` +(`notes.md` §57). +**Closed-vote harness, no planner LLM (Empirical as README +architecture, 2026-09-18 ~21:39):** +[JevOnly](https://github.com/buluoray/JevOnly) — code builds +every option from observation / goal / fact register; **Jev +only picks**. No planner LLM, no helper LLM, no free text. +Type without generation. Verify then undo. Irreversible +`risk ≥ 0.50` never default. Worked example *theirs*: 11 +steps, 43 Jev calls, ~340k tokens, ~$0.014, 17 s. Distinct +from Stagehand (LLM fallback). Do not copy `run.sh` +(`notes.md` §61). +**Host-owned product surface (Empirical as README + their +eval suite):** +[waymode](https://github.com/mossburgh/waymode) — the **app** +retains handlers, permissions, validation, and state; Jev +selects among live typed actions. `completed` is Jev's +reading — prove durable effects via server state. Default +p ≥ 0.7. Evidence *theirs*: 24/26 public suite, 34/36 +completion regression — **bounded development evidence, not +proof every app is self-driving**. Not on npm. Do not copy +AI_GATEWAY how-to (`notes.md` §61). +**Hot-click CU on an indexed viewport (Empirical as README ++ n=3 medians; 2026-09-19 ~00:39):** +[ego-jev](https://github.com/jiangkoumo/ego-jev) — drive +ego-lite with Jev. Indexed element table in; one request +answers operation **and** per-op target (speculative, +compatible-only heads). Code owns observe / execute / +stale-ref / loop / `--until` exit. Text model only when +typing is needed; malformed fill → `text_model_failed`, +never guess. Jev `done` ≠ business success. Measured +*theirs*: HN **4.9 s vs 9.7 s**, wiki **5.4 s vs 10.1 s** +(~2× vs per-step `kimi-k3`; n=3; high variance; not a +benchmark). Cousin of jev-ultrafast. Distinct from JevOnly +/ waymode / Stagehand. Do not copy `install.sh` +(`notes.md` §65). +**OCR+AX desktop CU (Empirical as README; MIT **427★**; +2026-09-19 ~09:51):** +[typesafe-computer-use](https://github.com/awlevin/typesafe-computer-use) +— macOS Vision OCR + AX → numbered items → TypeSafe +Choices (`kind` / `item` / `site` / optional +`offscreen`) → code clicks/types. **Never ships a +screenshot for the *decision*.** Writer only for +`type_text` / `site: other` / the one-shot **answer** +(the answer reader *may* receive the capture — a +writer packet, not the Choice). Overlapping options +read as doubt; keep the set exclusive. AX is a bonus, +never sole (Spotify 0 *theirs*). Post-type Noul 0.5 +and `--min-confidence` 0.4 stay product copy, not +Harbor τ. $0.0002 vs Opus $0.032 (155×) *theirs* on +**one screenshot**, not a taskset. **≠** jev-ultrafast +**≠** cua-s1 **≠** jev-macos-loop **≠** camoufox. Do +not copy `uv sync` / `.env` (`notes.md` §81). +**ASR voice-browser CU (Empirical as README; MIT **103★**; +2026-09-19 ~10:01):** +[jev-voice-browser](https://github.com/moritzkremb/jev-voice-browser) +— Web Speech partials → snapshot ≤100 → one 9–11-question +Jev request → Playwright. **Jev never generates.** Regex +spans; Jev picks; code copies. Closed-set may act on a +partial; free-text waits. Numbered overlays, spoken +digit, no second model. Spoken confirm is convenience +not auth. 27/27 fixtures *theirs*. **≠** +chris-wozniczek/jev-voice-control **≠** +nikolas-j/jev-voice-browser **≠** typesafe-computer-use. +Do not copy `npm` / `.env` / `run.sh` (`notes.md` §82). +**Adversarial browser, Playwright executes / Jev chooses +(Empirical as README; license null; 2026-09-19 ~04:39):** +[browser-jev](https://github.com/DowLucas/browser-jev) — +one Jev call per step (six oracle Nouls + severity Score ++ next-action Choice). Code-only checks first. **Sample +from the distribution, not argmax.** Fail only high conf +**and** high severity. Visual blind. Demo lesson: narrow +questions (untranslated 0.30 inside "confusing" vs 0.99 +on its own Q). CI exit 1 on non-baselined findings. Do +not copy playwright / `.env` (`notes.md` §69). +**Score-among-observed atlas (Empirical as public showcase class +pattern, 2026-09-19 ~00:38):** +[jevable.com](https://jevable.com/) — candidates already on the +page (a11y/DOM, ads, on-screen posts); the model scores; **code** +clicks / filters. Not a 342-title dump. Cross-link: jev-ultrafast / +gliner2-ultrafast / solari-reflex / cua-s1 / laya-mind2web. Your +Signal: score posts already on screen, apply rules locally — same +judge-once / re-policy family as Near Here. Do not merge Flights +7 s / $0.0039 with gliner2-ultrafast 12.20 s; computer-use "100×" +is a **claim** (`notes.md` §56). +**DOM-as-text + fan-out (Empirical as atlas browser-use *shape*):** a +screenshot task translated into a structured DOM snapshot as `state`, +then speculative questions over numbered candidates — not vision +(`notes.md` §49; `mental-models.md` §boundary). +**Axis check:** if the kept byte / cited fact / click target is not +already in the candidates you numbered, this card does not apply — +retrieve or parse first; do not ask recall. **Counterexample**: "write the patch that matches this sentence" — that is generation. **Test**: every kept byte occurs in the input; mixed never auto-included; snapshot stale → abort, don't guess. @@ -119,7 +416,83 @@ confidence goes back to the agent, not into silent mitigation. Noul workaround; run status silent / disclosed / recovered / clean. Heuristic fixture 12 traces: P=R=0.86; Jev on that fixture not yet measured. Pre-mortem: scan a new sandbox *before* users meet it, fail -the build on *silent* env-breaks — a shape, not a CLI. **Counterexample**: sampling 2% of production with an LLM judge — +the build on *silent* env-breaks — a shape, not a CLI. +**Merge-gate cousin (Empirical as README behavior / offline demo, +2026-09-18 ~16:48):** +[latch](https://github.com/CaseReed/latch) — cluster a finished red +run (Playwright / Jest / pytest / JUnit) **in code**; Jev labels each +cause (≤8 calls; cached free); **policy** returns Gate: PASS (infra +noise) vs Gate: BLOCK (real failure). The judge never says "ignore" +alone; `ignore_as_infra` needs `env_cascade` + an infra fingerprint. +Reporter never fails Playwright (missing key → `needs_human`); +`--gate` is a separate CI step. Demo: 8 connection errors → PASS; 5 +assertions → BLOCK. Message-based grouping fragments logic +regressions (`pallets/click`: 13 failures → 10 clusters). Pair with +Harbor (frozen CI artifacts × PASS/BLOCK) and rh-guard +(eval-integrity). Their policy thresholds are not class constants +(`notes.md` §51). Do not copy the reporter. +**Pre-review typed gate (Empirical as README + own-repo +latencies; 2026-09-19 ~02:38):** +[ci-gatekeeper-bot-jev](https://github.com/NemanjaManic/ci-gatekeeper-bot-jev) +— four questions (`should_review` / `risk` / `route` / +`touches_secrets`) → `auto-approve | human-review | block` +*before* expensive LLM/human review. Operator-owned +thresholds; conservative default (`cosmetic`) escalated +trivial diffs to human-review in practice. Measured +*theirs*: Jev **504–629 ms**; secondary review ~4–5 s +only on human-review + elevated risk. Comment never +includes raw diff. `package.json` MIT / GitHub SPDX +**null**. Cousin of latch (that one is flaky-vs-real on a +*finished* red run). Distinct from egma attention ≠ +correctness. Do not copy `action.yml` / secrets +(`notes.md` §67). +**Stop-hook attention redirect, not a merge blocker +(Empirical as owner-run hook smoke; 2026-09-19 ~03:38):** +[jev-preflight](https://github.com/muse0509/jev-preflight) +— eight risk axes in one request on a redacted turn +diff; `assist` = at most one reinspect then finish; +**fail-open**; default 0.85 **uncalibrated**. Not an +autofix, not a merge blocker, not a replacement for +tests/SAST. Distinct from latch / ci-gatekeeper +(authorize-or-block) and from rashedInt32/jev-lens +(human attention filter). Owner-run Claude Code 2.1.267: +no-key fail-open PASS; key-enabled exactly one +continuation *theirs*. Do not copy marketplace / key +(`notes.md` §68). +**VOI hunk prune before generative review (Empirical as +22-run cost table; 2026-09-19 ~05:46):** +[prune-review](https://github.com/shubhangi013/prune-review) +— Jev scores each hunk; only a smaller packet reaches +the generative reviewer. Safety escarpment always keeps +concurrency/auth/a11y/startup. Target ~20% cost cut. +*Theirs:* winning-only 27.9% (post hoc); all 22 incl. +305% outlier **1.18%**; excl. outlier 15.9%. Cost, not +quality. Source preview. Cousin of ci-gatekeeper (that +one auto-approve/human-review/block), not a clone. Do +not copy pnpm (`notes.md` §66, §70). +**Whole-repo intent beyond the diff (Empirical as CLI; +under construction):** +[jev-intent-review](https://github.com/yottayoshida/jev-intent-review) +— stated intent → search the repo after the change → +one small question per place → +VERIFIED/VIOLATION/UNKNOWN/NOT_APPLICABLE. Empty search +≠ proof. CLI works; GitHub Action not written. Do not +copy Cloudflare how-to (`notes.md` §70). +**Commit pre-review attention≠verdict (Empirical as 13 +labelled + own-history; 2026-09-19 ~07:49):** +[commitjev](https://github.com/yodablocks/commitjev) +(MIT) — seven Nouls + one headline Choice per commit; +six regex checks never reach the model. Middle band is +**"review"**, never rounded. Nouls decide; Choice only +headlines at confidence ≥0.50. Hook **blocks only on a +warning**. Calibration *theirs*: every rule fires on its +defect; **0 false on 5 clean** (small control); own 16 +commits 3 warn / 4 review / $0.0017. Same owner as +jev-orderby-bench (one commit per call; never sort +two-decimal probs). Cousin prune-review / ci-gatekeeper +/ jev-preflight. Do not copy hook install +(`notes.md` §72). +**Counterexample**: sampling 2% of production with an LLM judge — the economics inversion is the point. **Test**: planted harness bugs recovered; false-flag rate on known-clean runs; LLM never runs on the clean majority. High-stakes cousin: `luantak/is-malicious` is a *pre-run* @@ -152,9 +525,171 @@ re-filters); LlamaIndex Jev rerank **Empirical** BEIR nfcorpus MiniLM zero-shot MAP 0.4748 / nDCG@10 0.683 vs monoBERT 0.718 (`mappings.md` §4). Realtime ~200ms chat claims remain **Hypothesis** as a number. **Counterexample**: using top-1 Choice as a relevance score across queries; dropping RAG -chunks fail-closed so a timeout empties the context. **Test**: moderation +chunks fail-closed so a timeout empties the context. +**Decision-native RAG (Empirical as architecture; Hypothesis as a +universal win, 2026-09-18 ~17:48):** +[decision-native-rag-skills](https://github.com/emergency-lee/decision-native-rag-skills) +— retrieve wide → decide explicitly → evidence set → conflict +resolve → reason only over kept evidence. Embeddings stay candidate +generators. Provider-agnostic; no bundled Python harness; **no +universal benchmark**. Default migration gates are starting +targets. Offline replay → shadow → canary → A/B (`notes.md` §55). +Do not ship because an LLM judge prefers it. +**Classify-first agent I/O of the same sandwich (Empirical as +README, 2026-09-19 ~00:38):** +[jev-sift](https://github.com/kbhuw/jev-sift) — file lists and +public URLs stay candidate generators; the judge sees content; the +main LLM opens only items worth a closer look. Inline text the +agent already read cannot recover that cost. Uncertain/errors/ +truncation ≠ irrelevant. Mocks ≠ accuracy (`notes.md` §56). +**Living applied-mappings atlas (Empirical as showcase; 342 is +*their* count):** +[jevable.com](https://jevable.com/) — class patterns (intent +columns, score-among-observed, VOI gates, generative UI decide, +robotics text-state, draft-gate fail modes), not a hit list. +JSON-LD first page is 36; `pageSize` 36. No public API this pass. +Maker clocks stay claims unless already a named receipt +(`notes.md` §56). +**Recursive file search (MED; distinguish from federated web):** +[kazuhideoki/jev-search](https://github.com/kazuhideoki/jev-search) +scores local files then fzf — **not** +[superagents-lab/jev-search](https://github.com/superagents-lab/jev-search) +(web lanes). Max-over-chunks is not a calibrated whole-file +probability. No LICENSE this pass. +**Meaning-search without embeddings (Empirical as a named +stripped-repo card, 2026-09-18 ~18:46):** +[jevgrep](https://github.com/Bentlybro/jevgrep) (`jgrep`) packed- +parallel Jev relevance; two-stage outline then zoom top 30; no +index. 228 questions on docstring-stripped Flask/httpx/Django/ +AutoGPT: **79% top-5** vs BM25 40% / grep 20%. Keyword still wins +exact wording (BM25 top-10 96% vs 85%). Packed+parallel 0.9 s vs +serial ~23 min on AutoGPT 4,329 files. Distinct from kazuhideoki +(file+fzf), superagents-lab (web), and jev-sift (classify-first +MCP). Do not copy `install.sh` (`notes.md` §58). +**Meaning-grep AND/OR/NOT over line Nouls (Empirical as README ++ their LLM-as-judge test, 2026-09-18 ~21:39; dedicated +2026-09-19 ~16:30):** +[jev-semgrep](https://github.com/uehaj/jev-semgrep) — zero-dep +Node; one Noul per line × meaning; `-e`/`-a`/`-v` boolean +over *thresholded* bits; 30 lines × 8 concurrent. +Proposition ≠ embedding (cross-encoder line+question). +Contrast-set: all six “about a refund”; only customer +asking pass; angry-agent cosine ~1. Calibrated ~0.5 vs +top-k cosine. No index (vector index wins for repeated +large fixed corpora). Cross-lingual JP↔EN plus +FR/RU/DE/ES/ZH/KO; EN safer near threshold. Name collides +with [Semgrep.dev](https://semgrep.dev) SAST. **Not a +gate** (ranking fail-open; rh-guard skip). Distinct from +jevgrep (file/chunk packed search), jev-sift, jevex, +bohutang/sift, pg-jev, kazuhideoki/jev-search, +superagents-lab/jev-search, and jev-combinators +(metaphor ≠ literal AND/OR). LICENSE MIT (GitHub +NOASSERTION). **51★** this pass (ephemeral; SIGNAL ★42; +§61 0★). Their judge test: precision 0.94, recall 0.98 +on 10×51 lines — not Harbor. Do not copy npm / `npx` / +`.env` / marketplace how-to (`notes.md` §61, §86). +**Evidence-packet explorer (Empirical as their +`docs/performance.md`, author-run):** +[jev-semantic-explorer](https://github.com/jimmyhealer/jev-semantic-explorer) +(**jevex**) — index once (chunks + BM25), Jev ranks a +shortlist, MCP returns `source_of_truth` / tests / callers / +line ranges. Read-only, not a patcher. Claude Code A/B: +6.8→2.2 files, 8.6→3.2 tools. SWE-bench Verified n=8: +**1/8 → 6/8** finish (empty output = miss); packet n=50 +HitFile 0.233 vs BM25 0.159 is diagnostic, **not** the +product KPI. Distinct from jevgrep / jev-sift / +s1-graphify-indexer. Do not copy MCP how-to (`notes.md` +§61). +**Evidence-packet explorer delta (rename + n=16 SWE +card; 2026-09-19 ~07:49):** +[jevex](https://github.com/jimmyhealer/jevex) **is** +`jimmyhealer/jev-semantic-explorer` renamed (same +`created_at`; GitHub redirects). New *theirs*: SWE-bench +Verified **n=16**, 160s → **69s**, $8.74 → **$3.13**, +patch **16/16 both arms**. 90s cap 1/16 vs **11/16** +finished. Keep n=8 finish 1/8 → 6/8. Claude Code n=5 +6.8 → 2.2 unchanged. (`notes.md` §72). +**Measured RAG rerank vs a generative reranker (Empirical as +one-run; Hypothesis as a transfer, ~18:46):** +[Jev-RAG](https://github.com/Max-sm-yc/Jev-RAG) — same search +~30k tokens: RAG+Jev+Spark $0.00122838 / 62.3 s vs RAG+Spark- +rerank+Spark $0.00421838 / 228.14 s vs Spark full-context $0.0032 +/ **10.60 s**. ≥70% cost and 72% latency cut vs Spark *rerank*, +not vs no-RAG. Costs include embeddings. License null this pass. +Do not invent a bake-off (`notes.md` §58). +**Local rules first, then remainder Nouls; never auto-train +on the model's own hides (Empirical as README + small e2e; +2026-09-19 ~00:39):** +[x-reply-filter](https://github.com/zhuyansen/x-reply-filter) +— Chrome MV3. `rules.js` proves easy junk (zero cost); +batched Jev four Nouls on the rest (promo / bait / +off-topic / AI filler; default ≥0.75 collapses, does not +delete). Auto-hides sit in a confirm queue; only +user-confirmed examples become few-shot (10/10) plus a +"same class as marked junk" Noul. E2E *theirs*: three +samples → 0.90 / 0.93 vs 0.08 / 0.10. Cousin of +[bohutang/sift](https://github.com/bohutang/sift) +(§62 MED; ~$0.00003/post *theirs* — Substance/Humor/Chit-chat/Promo/Junk + AI-written). Cheap hold-before-show cookbook. +Do not copy wrangler (`notes.md` §65). +**Worth-your-attention VOI (Empirical as unreviewed goldens; +always qualify the owner; 2026-09-19 ~04:39):** +[ThinkyMiner/Winnow](https://github.com/ThinkyMiner/Winnow) +— Chrome extension: read / skim / save / skip from typed +answers; templates never prose. **Distinct from** +[kevinpita/winnow](https://github.com/kevinpita/winnow) +(context sieve). Feed batches ≤12; 7-day cache. +Unreviewed goldens *theirs*: **80%** verdict / **90%** +content-type. HN 30 links ~$0.0015. Not on the Chrome Web +Store. Do not copy unpacked-extension how-to +(`notes.md` §69). +**Personal-history feed without a social graph (Empirical as +README + 17 offline tests; 2026-09-19 ~06:43):** +[jevfeed](https://github.com/fengyiqicoder/jevfeed) +(MIT) — last 200 history pages stay local; outbound links +via Jina Reader; cheap LLM summarizes/filters; **one Jev +request per batch of ten** (the distribution *is* ranking). +No likes/follows/accounts. Distinct from ThinkyMiner/Winnow +(grade an existing feed) and kevinpita/winnow (sieve). Do +not copy `npm start` (`notes.md` §71). +**OpenRouter recipe atlas (Empirical as 15 small samples, +not benches; 2026-09-19 ~06:43):** +[jev-cookbook](https://github.com/nexibeo/jev-cookbook) +(MIT; 1★) — triage / PII / rerank / moderation / browser / +Gmail. Code prepares, Jev answers narrow questions. Samples +16–36 handmade; authors say **not benchmarks**. Recipes +01–13: 425 calls / $0.015; browser 5/6 *theirs*. Do not +copy OpenRouter tilde-id (`notes.md` §71). +**Decision-native inbox (Empirical as README + tests; +life/business, not SWE-only; 2026-09-19 ~07:49):** +[mailordinal](https://github.com/Milo318/mailordinal) +(MIT) — nine typed questions in one request, then a +**100-point deterministic policy** (SLA + account tier +in code). Does not ask "how urgent is this?" Humans own +ambiguity: low routing confidence never silently lowers +priority. Demo labelled `demo`; live Jev optional and +server-side. Cousin jav-email-cascade. Independent, not +affiliated with TypeSafe. Do not copy `npm run dev` +(`notes.md` §72). +**Public classification API (Empirical as README + +eval/README; life/business; 2026-09-19 ~08:37):** +[classifier-dev](https://github.com/mrmps/classifier-dev) +(MIT; **185★**; https://classifier.dev) — the +categorization *product* those inbox apps would call. +Caller labels in, label + calibrated confidence out; +batch `{id, text}[]` ~1000. Jev primary (`src/jev.ts`); +LLM fallback only. Distinct from ask-jev-ai's +six-question wall. Do not copy wrangler / `npm i -g` +(`notes.md` §73). +**Open NAR packaging cousin (Empirical as README; +2026-09-19 ~09:07):** +[NandhaKishorM/laya](https://github.com/NandhaKishorM/laya) +(Apache-2.0; **710★**) — self-hosted Choice/Score/Noul +with a script-before-p `Router` over Hub checkpoints. +Not a classification HTTP API. Not a TypeSafe drop-in. +Do not copy `pip install laya` (`notes.md` §76). +**Test**: moderation cost/coverage + false-hold vs false-publish; ranking recall *separate* -from nDCG; select misroute rate. +from nDCG; select misroute rate; required-evidence recall vs Top-K. ## 5. Skill / tool routing @@ -180,24 +715,179 @@ fail closed on side effects; no-match option when coverage is open ``` **Example**: skill_suggestion cookbook (**Contract**); GodsBoy 94.4% vs -70.8% lexical (exploratory: questions revised after the first full run); `Dicklesworthstone/skillranker` from live session context; +70.8% lexical (exploratory: questions revised after the first full run); +`Dicklesworthstone/skillranker` from live session context +(**VOI / abstention; 52★; hook fail-open** — see below); LlamaIndex selectors fail closed or a declared default. [`rajdhakad9826/routeKit`](https://github.com/rajdhakad9826/routeKit): Jev estimates task *requirements*; code applies hard constraints and a deterministic cost/quality/latency policy — Jev does not pick the model (**Hypothesis** until measured on *your* catalog; `notes.md` §33). +**Session-sticky first-prompt route (Empirical as README machine, +2026-09-18 ~18:46):** +[jev-adaptive-thinking](https://github.com/jxu-dev-c/jev-adaptive-thinking) +classifies the first user prompt for `jev-auto`, then **locks** +provider/model for the process-local session; later requests never +reclassify. Timeout / missing first-round text / no stable session +ID → lock `gpt-5.6-sol` (fail-closed fallback, not passthrough). +Same family as routeKit. License null this pass. Live testing left +to the deployer. Do not copy dylib/YAML (`notes.md` §58). +**Harness plugins (brief, ~19:48):** +[dsh-jev](https://github.com/buberlo/dsh-jev) — DeepSeek Harness +decision layer; a model answer can only gate, never widen a +permission; failure never produces an allow; not on npm. +[opencode-system-one](https://github.com/emirbartu/opencode-system-one) +— OpenCode plugin; every Jev call fails open; license null. +Do not copy plugin JSON (`notes.md` §59). +**OMP/pi acceptance + route (Empirical as README fail +polarity, 2026-09-18 ~21:39):** +[omp-jev-extensions](https://github.com/luw2007/omp-jev-extensions) +— `jev_acceptance_gate` before done (Choice `{accepted, +rejected}`, not a boolean); `jev_route` subagent topology + +tier. **Fail-open** if Jev missing/timeout/malformed; +fail-open paths `confidence: 0`. Distinct from +pi-jev-approver (fail-closed without a key). Do not copy +bun / `~/.omp` (`notes.md` §61). +**OMP prompt suppression (Empirical as measured traffic + +labelled corpus, 2026-09-18 ~22:38):** +[omp-greenlight](https://github.com/SemetricLabs/omp-greenlight) +— grades gated tool calls; suppresses the approval prompt +when Jev says allow. **1,013 calls / 10 sessions / 8.95 +session-hours.** Default preset **40.9%** prompts removed; +**0 of 94** unsafe auto-approvals on a 140-row labelled +corpus (live traffic has no labels). Operator owns +thresholds; the plugin **never self-tunes** the safety bar. +Not a sandbox; host `bash.patterns: deny` fires ahead. +Agent prose never sent (0→3 corpus misses). Never shadows a +built-in tool. Distinct from specpi-jev-guard (single danger +score; static fast-path miss), toolgate, interlock, and +omp-jev-extensions (fail-open *route*). Composes with +waymode. Do not copy `omp plugin` / YAML (`notes.md` §62). +**Skill-library VOI (Empirical as README architecture; +correction vs §7 fail-closed; 2026-09-19 ~01:47):** +[skillranker](https://github.com/Dicklesworthstone/skillranker) +— Jev two-pass (wide Choice then fit Nouls) from live +session context; both passes include **"none of these"**. +Advisory: the agent follows user instructions. Claude +prompt-hook maps recommendation failures to **quiet +exit-zero** (never blocks the agent). CLI keeps meaningful +exit codes. Libraries >254: Quill lexical prefilter admits +≤254 + none. Explicit requests resolve locally first. +Local feedback / replay without a new Jev call. **Hunch:** +pay to load a skill iff it changes the next step. +Compose with decision-combinators (control plane, not chat). +Distinct from skill-broker (grants). Do not copy cargo +(`notes.md` §66). +**Hermes pre-agent skill intervention (Hypothesis / outline +only — not a production recipe):** +[skill-broker](https://github.com/adamjralph/skill-broker) +— `PROJECT-OUTLINE.md` is authoritative. Deterministic code +owns catalog, profile policy, limits, and **grants**. Jev +scores relevance/confidence over authorised candidates and +**never grants access**. Candidates ≠ grants. Jev down → +foundation-only; never broaden access. Replayable route +evidence. Distinct from jev-hermes (route ≠ memory) and +from shipped routers on this card. Language/license null +this pass. Do not copy an install (`notes.md` §62). +**Sibling contrast this hour (delta, not a re-fold; +2026-09-19 ~02:38):** README now restates the outline; +still project-definition (`docs/adr` appeared; no +runtime). Same evidence≠authority doctrine as +[turnstile](https://github.com/zyphr-labs/turnstile) +(runtime authorize after policy; missing Jev → Review) +and opposite polarity from +[skillranker](https://github.com/Dicklesworthstone/skillranker) +(advisory VOI; hook fail-open). skill-broker **grants** +live in code. `notes.md` §67. +**Codex MCP host adapter (Empirical as README +architecture; ranking unbenchmarked; 2026-09-19 +~02:38):** +[jev-in-codex](https://github.com/teempai/jev-in-codex) +— `jev_select_capability` / `jev_search` / `jev_triage`. +Caller supplies the catalog; server never executes +capabilities or sees Codex internals. Independent Nouls, +batched four; rec ≥ 0.5 is a **provisional heuristic**. +Absent key / errors → **lexical fallback** (local scores +are not model probabilities). Experimental MVP; MIT. +Distinct from jev-routing (Go host adapter, **not MCP**) +and jev-sift (topology A classify-first). Do not copy +npm / `config.toml` (`notes.md` §67). +**Harbor roster-size harness + Pi strip-roster (Empirical +as README architecture; no live Jev numbers this pass; +2026-09-19 ~04:39):** +[pi-jev-skill-bench](https://github.com/iamdin/pi-jev-skill-bench) +— BM25 vs Jev at roster **50 / 100 / 200 / 500**; 43 gold; +token/USD = chars/4; experiment harness not a production +claim. Cite only after `out/results-*.md` exists. +[pi-jev-skill-suggestion](https://github.com/iamdin/pi-jev-skill-suggestion) +— strip ``; two-stage (mean Noul 0.30 → +chunked Choice ≤254+none → shortlist 3 → fits 0.40); +fail-open; **no key → no-op**. Tool mode is a tools≠use +cousin (hopes the agent calls `skill_suggest`); auto mode +runs every prompt. Contrast skillranker (advisory VOI) / +skill-broker (grants) / jev-in-codex (caller catalog). +Do not copy `pi install` (`notes.md` §69). +**Constrained optimizer + S1 features (Empirical as live +analysis *shape*; 2026-09-18 ~23:40):** +[slo-router](https://github.com/zeeshan8281/slo-router) +— Jev supplies bounded task / exactness / external-evidence +features; a constrained controller picks the cheapest +backend meeting quality + latency SLO floors. Fail-open to +deterministic local features on timeout/invalid. Exactness +raises the quality floor; **never overrides** context or +capability. On their fixture, Jev preserved the same +routes/accuracy as the local path and raised p95 E2E +**77.93 → 490.38 ms** (~6.3×). 3/8 task-label disagreements +did not change routes. Eight-row demo is **not** a model +benchmark. License null this pass. **Hunch:** System One +belongs on the feature side of a constrained optimizer, +never as the sole hard gate on the hot path. Distinct from +routeKit (unmeasured) and bitrate-advisor (soft affinity +inside a cap). Do not copy uvicorn / OpenRouter +(`notes.md` §63). +[`trietphan/jev-claw`](https://github.com/trietphan/jev-claw) is the +same split for OpenClaw (classify axes; `decide()` maps the route; path +regex floors risk). [`nekowasabi/jev-routing`](https://github.com/nekowasabi/jev-routing) +is a host adapter, not an MCP plugin: compact, then one Choice + done, +then one schema (`notes.md` §44). +**Host-adapter surface delta (2026-09-19 ~07:49):** +same [jev-routing](https://github.com/nekowasabi/jev-routing) +binary now lists **Cursor Agent CLI** and **Devin CLI** +beside Claude Code / Codex / Grok Build. Still not MCP. +Do not re-card; do not copy ports (`notes.md` §72). +**Hermes plugin branded as Jev is Agnes (identity lock; +2026-09-19 ~07:49):** +[hermes-plugin-jev](https://github.com/Mrmimee/hermes-plugin-jev) +(README MIT / GitHub SPDX null) — Choice/Noul/Score +*shape* over **Agnes 3.0 Flash** chat-completions. +Distinct from hermes-jev-router (TypeSafe +WHETHER/HOW/WHAT). Do not copy `~/.hermes` +(`notes.md` §72). [`TheoOliveira/pi-jev`](https://github.com/TheoOliveira/pi-jev) is the same selector hole inside Pi (tools + skills); fail-open to a keyword shortlist; **not** `kevinpita/pi-jev-context` (sieve). [`ddfeyes/jev-mode`](https://github.com/ddfeyes/jev-mode) is the latency-class split: bulk triage/tag/route off the frontier context (synthetic 1,000: −77.8% tokens; accuracy claim is **parity**). +**Route ≠ memory** ([jev-hermes](https://github.com/de-niji/jev-hermes)): +a cheap intent Choice skips memory/tool *tours* on `calendar` / `mail` / +`status`; `complex` keeps memory. Savings are skipped tours, not +turning memory off (`notes.md` §48). Toolrouter / open JevRouter: **Hypothesis** until measured on *your* catalog. **Counterexample**: the agent looping "pick a tool, call it, pick again" with the provider as the planner. **Test**: callability (literal / paraphrase / near-miss neighbor); reject-all when nothing fits; calibre reminder — thresholds do not transfer (`validation.md`). +**Pi System-One control plane (Empirical as README + npm +tests, no live quality numbers; 2026-09-19 ~06:43):** +[pi-jev-control](https://github.com/goodruizhan/pi-jev-control) +— task/model router, skill/memory/context gates, review +gate, GUI action router in one Pi extension. License null; +v0.3.0 private. Compaction never modifies the on-disk +session; GUI < threshold → unknown, never force-click. +Distinct from omp-jev-extensions / jevons / pi-heed / pi-om. +Do not copy `pi install` (`notes.md` §71). ## 6. Expensive observation router @@ -222,4 +912,383 @@ the OCR bill it authorises**; 9 false-skips vs 28 for rules-only OCR, so a watermark talks a scan into "has text." **Test**: planted scans are sent; planted born-digital pages are not billed; page order preserved. Re-measure on *your* documents. Same sandwich as jevgate (Proven / Refused / -Unknown). +Unknown). Same VOI as retrieve-then-state: if the answer is not in the +cheap text layer, **pay for the passage / OCR**, then judge +(`mental-models.md` §boundary; atlas history suite). + +## 7. Capability kernel / human-confirmed gate + +**Method**: change the *trust boundary*, not the after-the-fact +"is this dangerous?" question. Two named shapes this hour +(`notes.md` §59): + +1. **Capability kernel.** The LLM is ring 3; a kernel it cannot + talk to is ring 0. Secrets never enter the agent (canaries and + placeholders only). The action space is closed. A judgment-class + model is a **sensor**; ordinary policy code decides BLOCK / ASK + / ALLOW. Type-safe ≠ correct; irreversible stays behind a + threshold **and** a human. +2. **Human-confirmed kill.** The model recommends; the operator is + the only actuator. Re-check identity immediately before the + irreversible signal. Shields override the judge. Displayed + explanations are app-owned mapped text, not raw model prose. + +**Transfers**: Leveson sensor ≠ constraint (`mappings.md` §8); +structural prove ∩ remainder (`mappings.md` §18); fail-closed on +the irreversible act (`mixed-architecture.md`). **Does not +transfer**: a launch-week firewall that asks "dangerous?" after +the LLM already decided with **real secrets in scope**; treating +toolgate (pre-exec of a *proposed* call) as the same product as a +kernel that never showed the secret; letting mapped UI copy be +the model's free-form reason. + +```text +stunt_double = canaries + placeholders + allowlisted actions # code +sensor = parallel Nouls / Choice on the proposed act # model +policy = BLOCK | ASK(human) | ALLOW + placeholder swap # code +kill = human confirm after identity re-check # not the model +``` + +**Example (Empirical as README architecture, 2026-09-18 ~19:48):** +[interlock](https://github.com/somoore/interlock) — twelve-hazard +Noul battery ~100 ms; `policy.py` is the product; 38-case +regression set tunes the local judge, **not a blind paper**. +Distinct from [toolgate](https://github.com/fdemir/toolgate) +(allow/block/review on a proposed tool; Jev is not authorization; +real args may already be in scope). rh-guard crossover: eval- +integrity is a different hole from a ring-0 kernel; do not merge +products. Do not copy pip / `INTERLOCK_ARMED` how-to. +**Human-confirmed cousin (Empirical as README safety model):** +[port-cleanup](https://github.com/epiphany-dynamics/port-cleanup) +— Jev Stop/Keep/Your-decision; kill recs need conf ≥ 0.8; +identity re-check before SIGTERM; shields override; TCP only; +tiny final race (no pidfd). Gate UX for Augustus + rh-guard. +Do not copy Keychain how-to. +**Permission vs probability (Empirical as measured +suppression, not a kernel; 2026-09-18 ~22:38):** +[omp-greenlight](https://github.com/SemetricLabs/omp-greenlight) +— Jev is still the sensor; the **operator** owns the +auto-approve bar; the plugin never self-tunes it. Host deny +rules remain the constraint and fire first. Default 40.9% / +0 of 94 *theirs* on the labelled corpus. Not a sandbox; not +interlock (secrets may already be in the agent's world). +Do not copy YAML (`notes.md` §62). +**Spoken confirm ≠ auth (Empirical as README; voice; +2026-09-19 ~10:01):** +[jev-voice-browser](https://github.com/moritzkremb/jev-voice-browser) +— `destructive ≥ 0.5` → say "confirm". README *theirs*: +convenience, not a guarantee. Anyone who can reach the +control port drives the browser. Sensor, not an +interlock. rh-guard owns the gate cousin. Do not copy +`npm` (`notes.md` §82). +**Wrap-as-execution ALLOW/ASK/DENY (Empirical as +README; 2026-09-19 ~10:20):** +[AgentGhost](https://github.com/reddpy/AgentGhost) +— the wrap *is* the tool's execution function; +rules first; ASK/DENY throw; `failMode: closed`. +Judge swappable. Provider-hosted tools out of +reach. **≠** actiongate **≠** toolgate **≠** +jev-use. rh-guard owns the gate cousin. Do not +copy `npm` (`notes.md` §83). +**Privilege ≠ verdict / effect-based shell gate +(Empirical as certification; 2026-09-18 ~23:40):** +[construct-auto-classifier](https://github.com/godspede/construct-auto-classifier) +— fast-allow/deny <1 ms, then Jev Choice allow/deny + nine +independent risk Nouls. Allow only if Choice allow at +operator-owned `minConfidence` (0.6) **and** every risk +below `riskThreshold` (0.7). Missing/low-conf/high-risk/ +failed call = deny (fail-closed). `sudo status` can be a +safe read. Jev: **0** dangerous allowed / 975 decisions; +every chat model leaked (16–104). Pair with dinostomp +(audit the instrument) and omp-greenlight (operator-owned +dial). Distinct from toolgate / greenlight / jevgate / +interlock. **Hunch:** contracts on effects, not surface +tokens. **Landed-script trust** (byte-identical to the +remote default branch; default on; trusts whoever +controls that remote). **Headless ≠ auto-approve:** +escalation becomes deny-and-report, not a pending prompt +and not an allow (`notes.md` §68). Do not copy bun / agy +(`notes.md` §63). +**Typed escalate/continue/abort baton (Empirical as README ++ 40 tests; 2026-09-19 ~04:39):** +[jev-handoff](https://github.com/shitianfang/jev-handoff) +— MCP two-way handoff. Escalation reasons: +needs_generation / not_typeable / low_confidence / +backend_error. Gate `allow` **never grants**. Fail-open +(Jev down → typed escalate). Inverted loop: executor +enumerates, Jev picks, LLM woken only on escalate. +Vercel drops confidence (margin fallback not calibrated). +No independent quality bench. Same author as wakegate. +Do not copy npx / mcp.json (`notes.md` §69). +**Toolbelt sensors, not policy (notes only; 2026-09-19 +~04:39):** +[jev-security-scan](https://github.com/win4r/jev-security-scan) +(MIT; stdlib; local rules then nine Nouls; dual p≥0.85 + +locate; four synthetic samples; not a cert; cousin +is-malicious). +[jev-decisions](https://github.com/bojansandhaus/jev-decisions) +(MIT; 1★; 25 prepared reviews; **advice, never stop +commands**; auto hooks off). +[jev-vs-llm-guardrails-intent-router](https://github.com/TeoMastro/jev-vs-llm-guardrails-intent-router) +(license null; 218 labelled items; `summary.md` **404 +this pass** — do not invent numbers; README block ≥0.70). +**rh-guard owns the reward-hack angle.** +**Jev supplies evidence, code owns authority (Empirical as +README slogan; 2026-09-19 ~00:39):** +[actiongate-jev](https://github.com/omkarghugarkar007/actiongate-jev) +— deterministic policy / RBAC / schemas / limits own +ALLOW | REVIEW | BLOCK. Jev (via OpenRouter) is semantic +evidence only. A positive model score **never overrides** a +deterministic security failure. Six narrow questions, never +one vague "is this safe?" Financial / destructive / +credential **fail closed** if Jev is down. 500-case eval is +label-baseline integrity, **not** model accuracy. Early MVP. +Distinct from toolgate / interlock / construct / greenlight. +**Hunch:** canonical sensor≠constraint slogan for the +class. Do not copy pnpm (`notes.md` §64). +**Persist constraints across compaction (Empirical as +79-session bench; 2026-09-19 ~05:46):** +[pi-heed](https://github.com/Nyarlathoteppppp/pi-heed) +— conversational policy as structured state; replayed +after compaction without calling Jev again. Jev +classifies KEEP/LIFT/…; **never writes policy**. +Side-effecting calls checked before they run. Fail-open. +Shadow default. *Theirs:* v0.8.0+Jev recall 98.5% / +false block 0.0% / $0.000058; mid-session rule change +8/13 off vs 0/13 on. Distinct from actiongate +(RBAC/schema). Do not copy `pi install` (`notes.md` §70). +**Pi control-plane tool/GUI gates (Empirical as README; +license null; 2026-09-19 ~06:43):** +[pi-jev-control](https://github.com/goodruizhan/pi-jev-control) +— deterministic fast-path then Jev on uncertain tools; +GUI confidence below threshold → `unknown`, **never +force-click**. Distinct from pi-heed (constraint ledger) +(`notes.md` §71). +**jev-use PreToolUse gate (Empirical as 12/12 + Vercel +margin; 2026-09-19 ~06:43):** +[jev-use](https://github.com/shitianfang/jev-use) +— deny/ask only; **fail-open**; 12/12 *theirs*; Vercel +drops confidence so margin default 0.4. Same author as +jev-handoff. Gate never grants (`notes.md` §71). +**Commit-msg hook: fail-open on instrument failure +(Empirical as 13 labelled; 2026-09-19 ~07:49):** +[commitjev](https://github.com/yodablocks/commitjev) +— hook **blocks only on a warning**; a failed check is +not a reason to refuse. Middle band is review, never a +verdict. Credential regex cannot tell a fixture from a +key (correct failure direction). Do not copy hook +install (`notes.md` §72). +**Turnstile clone (Empirical as README architecture; +2026-09-19 ~01:47):** +[turnstile](https://github.com/zyphr-labs/turnstile) +— deterministic policy first; Jev semantic remainder +only after permit; allow/review/deny; receipts + +threshold replay with no new model calls. Jev **never +grants** authority policy denied. Missing Jev / timeout +→ **Review**. Starting 0.85 deny / 0.35 review are **not +calibrated**. Experimental alpha; no npm; Claude adapter +defaults observe + Jev off. Demo uses fixed judgments +(enforcement, not accuracy). Same doctrine as actiongate; +different product. Do not copy bun (`notes.md` §66). +**Advance/coverage ledger (Empirical as README + +BEYOND-JEV.md; 2026-09-19 ~02:38):** +[seal](https://github.com/Reasonofmoon/seal) +— **No seal, no advance.** Jev answers questions; SEAL +answers whether the world may change and **shows the +exception queue**. Every seal stamps `coverage.path` ∈ +`{auto | escalate | human | code}`. Mint ≠ product +brain. Deterministic first (`provider: code:…`). Effects +stay locked while escalations are open. MIT; zero runtime +deps. Do not copy `scripts/demo.sh` (`notes.md` §67). +**Never-confidently-wrong kernel (Empirical as TLA+ + +1,080 golden chaos table; 2026-09-19 ~02:38):** +[jev-labs](https://github.com/copyleftdev/jev-labs) +— TLA+ owns the protocol (quorum 3 of 5; stability gate +above identity noise floor 0.042; TLC 1,049,750 states / +0 errors). Jev is the noisy oracle. Golden pharmacy +rounds *theirs*: **0** wrong across none/realistic/severe +(severe: 314 correct / 46 escalated). Escalate is +allowed; a confident wrong is not. Synthetic, **not** +clinical. Inverse of soundness theater: do not hard-gate +a soft judgment without that path. Do not copy +`verify.sh` (`notes.md` §67). +**Counterexample**: post-decision "is this dangerous?" with AWS +keys still in the prompt. **Test**: delete the sensor — the +constraint and the closed action space still hold; a canary use +is a catch; a human still confirms the irreversible act. + +## 8. Decide → policy → LLM leftover cascade + +**Method**: one typed decide contract, ordinary code as the +router, a generator only where leftover *text* must be written. +Three Harbor-shaped compare arms share the contract so the +policy never knows the backend: native-probability decide (one +call), verbalized JSON confidence (one call), constrained +single-token logprobs (one call per question). **Transfers**: +Noul 0.5 is cannot-tell and is **never rounded** into auto; +Score confidence 0.0 is a flat distribution and is **never +acted on**; hard flags (injection) always review even with an +LLM configured. LLM hook optional — unset degrades to human +review, the run still completes. **Does not transfer**: treating +self-reported `"confidence"` as a Noul; multiplying eight +parallel answers into a joint; hosted Jev as a default for +real customer mail (compliance first; the `DecisionBackend` +seam is the local-head answer). Distinct from dual-process-ai +(S1/S2 metaphor; routing accuracy unmeasured). + +```text +prepare = strip quoted history / signature / cap # code +decide = 8 typed questions, one shared Answer schema # model (any backend) +policy = auto | review | llm # code, thresholds +leftover = draft category/priority/summary JSON # LLM only if policy says so +``` + +**Example (Empirical as README architecture + mock compare, +2026-09-18 ~20:43):** +[jav-email-cascade](https://github.com/skiingfalcon/jav-email-cascade) +— 74 labelled synthetic emails; jev / gen-json / gen-logprob +arms; ~$0.034/1k emails *theirs*. Mock finding: gen-json +confidence essentially flat → almost none cleared the acting +threshold; gen-logprob works at 8× calls. A live jev vs +generative comparison is the point of running it, not a table +to invent here. License null this pass. Do not copy uv / +`.env`. `notes.md` §60. +**Counterexample**: rounding Noul 0.49/0.51 into auto-act; +asking a chat model the eight questions and treating the +JSON `"confidence"` as calibrated. **Test**: the same emails +through all three backends; report raw accuracy, acted +accuracy, and mean confidence on wrong answers; injection +fixtures never auto. +**Inbox cousin without leftover LLM (Empirical as README; +2026-09-19 ~07:49):** +[mailordinal](https://github.com/Milo318/mailordinal) +— same sandwich, no generator required: nine typed +signals → 100-point policy → ranked queue. Humans own +the review lane. Life/business. Independent of TypeSafe +(`notes.md` §72). +**Public decide-backend cousin (Empirical as README; +2026-09-19 ~08:37):** +[classifier-dev](https://github.com/mrmps/classifier-dev) +— spam/inbox/feedback over HTTP; leftover LLM is +*fallback when Jev is down*, not the product. Policy +(hold / route / act) still lives in the caller. +`notes.md` §73. +**Open NAR cousin (Empirical as README; +2026-09-19 ~09:07):** +[NandhaKishorM/laya](https://github.com/NandhaKishorM/laya) +— same sandwich without a hosted Jev: Router picks +the checkpoint, policy stays in the caller, 0.85 is +still soft. `notes.md` §76. + +## 9. Closed-vote computer-use + +**Method**: drive a task with nothing but closed votes. Code +constructs every available option from the environment's +state, the goal, and an explicit fact register. A +judgment-class model assigns probabilities and picks. Code +acts, verifies the effect, undoes what did not work, and +keeps values read off the environment. **Transfers**: type +without generation (values from goal / facts / page only); +irreversible risk votes never default; inspectable rejected +alternatives. **Does not transfer**: a planner LLM that +proposes free text; inventing fill values; treating Jev's +`completed` / done vote as verified success when the host +has server state to check. Distinct from Stagehand (LLM +fallback when pick abstains) and Cua-S1 (not TypeSafe Jev). + +```text +observe = numbered controls / typed actions the host already owns +options = code builds from observation ∪ goal ∪ fact register +vote = Choice / Noul: done? off-path? which action? +act = existing handler or Chromium; never generated selectors +verify = before/after vs intended effect; undo reversible misses +``` + +**Example — harness (Empirical as README architecture, +2026-09-18 ~21:39):** +[JevOnly](https://github.com/buluoray/JevOnly) (Apache-2.0) — +no planner LLM. Chromium first env; core is +environment-agnostic. Worked Wikipedia compare: 11 steps, 43 +Jev calls, ~$0.014, 17 s *theirs*. `risk ≥ 0.50` never +default. Do not copy `run.sh` (`notes.md` §61). +**Example — product (Empirical as README + eval suite):** +[waymode](https://github.com/mossburgh/waymode) (MIT) — the +**app** keeps handlers, permissions, validation, and state. +Jev selects among live typed actions; a new control enters +the next snapshot without a matching model tool. Default +p ≥ 0.7; 8 steps. Evidence 24/26 and 34/36 *theirs* — +bounded development evidence, not a self-driving proof. Jev +selects the field; it does not generate fill text. Not on +npm. Do not copy AI_GATEWAY (`notes.md` §61). +**Hot-click cousin (not closed-vote; same observe→score→act +hole; 2026-09-19 ~00:39):** +[ego-jev](https://github.com/jiangkoumo/ego-jev) — indexed +viewport table; Jev picks operation+target; code owns the +loop; optional text model only for type. `--until` in code +beats Jev `done`. n=3 medians ~2×, not a bench. Distinct +from this card's no-planner extreme (`notes.md` §65). +**Adversarial cousin (Playwright executes, Jev chooses; +license null; 2026-09-19 ~04:39):** +[browser-jev](https://github.com/DowLucas/browser-jev) — +code-only checks first; sample from the distribution not +argmax; fail only high conf **and** high severity. Visual +blind. Same inverted-loop family as jev-handoff +(`notes.md` §69). +**Decider≠executor cousin (Empirical as README; 2026-09-19 +~05:46):** +[jeffrey](https://github.com/thomasbrueggemann/jeffrey) +(MIT) — Jev owns next-tool / progress / risk / done; the +LLM **only fills args**. Loop `Jev → tool → Jev`. Risk +Score ≥ 0.5 pauses mutating tools. Stuck ladder withholds +the looping tool (2 Jev / 0 steps). Distinct from this +card's no-LLM extreme and from jev-handoff (host baton). +Do not copy npm (`notes.md` §70). +**Hand no-text steps (Empirical as 95-call card; 2026-09-19 +~06:43):** +[jev-use](https://github.com/shitianfang/jev-use) +(MIT v0.4.1) — plugin hands did-it-work / which-next / +severity / safe to Jev; writing stays with the LLM. p50 +220 ms; batched 186 vs 2,672 ms; gate 12/12 *theirs*. +**≠** jev-ultrafast. Do not copy `npx` (`notes.md` §71). +**Never free-generates (Empirical as README demo; +2026-09-19 ~06:43):** +[jev-gpt](https://github.com/florian-hoenicke/jev-gpt) +(license null) — one typed question per word over a +WordNet / jina tree, then rank texts. ~400 calls / 75 s / +2¢ *theirs*. Architecture demo, not a product. Distinct +from jeffrey (pick next-tool). Do not copy API-key how-to +(`notes.md` §71). +**Continuous-control cousin (Empirical as README delta; +2026-09-19 ~09:50):** +[khordoo/jev-reflex-autonomy-lab](https://github.com/khordoo/jev-reflex-autonomy-lab) +— code owns physics/collisions; Jev only picks the next +typed flight action; optional S2 never flies and never +grants. Escalate without stalling. Local controller +**≠** githubnext/localjev. Experimental viz, not this +card's no-LLM extreme. Do not copy npm (`notes.md` §80). +**OCR+AX productized cousin (Empirical as README; MIT +**427★**; 2026-09-19 ~09:51):** +[typesafe-computer-use](https://github.com/awlevin/typesafe-computer-use) +— hosted Jev picks among numbered OCR+AX items; code +acts. Writer is leftover generation, not a planner. +`done` is loop termination, not verified success. +Exclusive action set; split questions. 155× *theirs* +one screenshot. Distinct from this card's no-planner +browser extreme and from ego-jev hot-click. Do not +copy `uv` (`notes.md` §81). +**ASR voice-browser cousin (Empirical as README; MIT +**103★**; 2026-09-19 ~10:01):** +[jev-voice-browser](https://github.com/moritzkremb/jev-voice-browser) +— Playwright executes; Jev only picks among snapshot +ids and regex spans. Partial-speech wait policy. +Spoken confirm ≠ auth. Distinct from this card's +no-planner extreme and from OCR desktop §81. Do not +copy `npm` (`notes.md` §82). +**Counterexample**: Stagehand extract `"pick"` with LLM +fallback sold as "no LLM" — pick is a fast path, not this +card. **Test**: every typed character exists in goal, facts, +or observed text; every irreversible act had a separate +risk vote; durable writes proved from host state, not from +Jev `completed`. + + diff --git a/.agents/skills/augustus/references/composition-algebra.md b/.agents/skills/augustus/references/composition-algebra.md index ebefb8b..1c1a5fb 100644 --- a/.agents/skills/augustus/references/composition-algebra.md +++ b/.agents/skills/augustus/references/composition-algebra.md @@ -16,8 +16,8 @@ mappings.md conventions. |---|---|---|---|---|---| | 1 | **Operand** | F(Jev(...)) — judgment's number feeds the function | Nouls as CatBoost features; Score as PUCT leaf value | It's a calibrated belief in your rubric's units, not a natural quantity; version feature/question defs with the consumer | **Empirical recipe** | | 2 | **Post-judge** | F(x) → Jev judges the result | Output judge (leaks_secret, failure_class); citation check on generated text | Only the post-judge sees what the call printed; the pre-gate cannot | **Empirical recipe** | -| 3 | **Gate** | if Jev(x): apply F — Jev decides *whether* F runs, or whether F's result is admitted | Pre-action gates (destructive .90/exfil .70); winnow context sieve | A gate is a filter, not authorization — validate operation+target in code. Error paths fail **per action**, not always open: an advisory guard fails open *because* a hard interlock or sandbox sits underneath; a gate that selects or authorizes a side effect fails closed (`mixed-architecture.md` prefilter table; `mappings.md` §18) | **Empirical recipe** | -| 4 | **Selector (of F or its parameters)** | Jev picks which F runs: Choice over functions/models/effort levels | jev-router (cheapest model), jev-codex-router (model+effort), DiffJury review_depth | Dispatch stays in code; per-option consequences are your cost model; confidence-gate the selection | **Empirical recipe** | +| 3 | **Gate** | if Jev(x): apply F — Jev decides *whether* F runs, or whether F's result is admitted | Pre-action gates (destructive .90/exfil .70); winnow context sieve; pi-heed side-effect check | A gate is a filter, not authorization — validate operation+target in code. Error paths fail **per action**, not always open: an advisory guard fails open *because* a hard interlock or sandbox sits underneath; a gate that selects or authorizes a side effect fails closed (`mixed-architecture.md` prefilter table; `mappings.md` §18) | **Empirical recipe** | +| 4 | **Selector (of F or its parameters)** | Jev picks which F runs: Choice over functions/models/effort levels | jev-router (cheapest model), jev-codex-router (model+effort), DiffJury review_depth, jeffrey next-tool | Dispatch stays in code; per-option consequences are your cost model; confidence-gate the selection. **Pick ≠ fill:** the LLM may write args; Jev does not | **Empirical recipe** | | 5 | **Comparator** | Replace a semantic comparator inside sort/rank: "more relevant / more severe" as a key | Rerank; skillranker; order statistics over semantic keys | Comparability needs a shared rubric; measure recall separately from rerank quality | **Empirical recipe** | | 6 | **Prior / initializer** | Jev distribution seeds a deterministic method that refines it | MCTS PUCT priors; beam-search branch priority | It's a heuristic prior, not a posterior; refine with real observations | **Empirical recipe** | | 7 | **State estimator, F = controller** | Jev estimates named probabilities; deterministic policy with hysteresis acts | foreman (progress/stuck/complete → continue/stop/retry/verify) | The model never commands; interventions enumerated in code | **Empirical recipe** | @@ -40,6 +40,34 @@ mappings.md conventions. never the model. - **→ (implication) / chains**: decompose into gate → act → post-judge; never encode multi-hop logic in one question (indirection costs accuracy). +- **Named combinators** ([jev-combinators](https://github.com/voidning/jev-combinators); + renamed from decision-combinators, same repo): + Then / Gate / Vote / Cascade / Weighted plus **extended** + Router / Loop / Retry / Fallback / Memory are + **control-plane** wiring, not a license to treat parallel + Nouls as independent. Digital-design slogan (transistors / + logic gates / chip) is a *metaphor* for soft classifiers; + ∧/∨ aggregation still follows the rule above. Fallback is + the fail-closed node; Memory gates what to remember. No + measurements. `notes.md` §66, §69. +- **Decider ≠ executor** ([jeffrey](https://github.com/thomasbrueggemann/jeffrey)): + position 4 (Selector of next F) stays Jev; arg fill is + generation, not a Jev position. The loop is Jev→tool→Jev. + Pick ≠ fill. Mapping §9 still rejects the fused + planner-writer. `notes.md` §70. +- **Selector of the next word** ([jev-gpt](https://github.com/florian-hoenicke/jev-gpt)): + position 4 applied to generation itself. Each token is a + Choice over a closed lexicon; the model never + free-generates. Architecture demo (~400 calls / 75 s / + 2¢ *theirs*). Distinct from jeffrey (selector of next + *tool*). `notes.md` §71. +- **Hand no-text steps** ([jev-use](https://github.com/shitianfang/jev-use)): + selector + gate; writing stays generation. Vercel drops + confidence so margin is a different statistic. + `notes.md` §71. +- **TLA+ kernel around votes** ([jev-labs](https://github.com/copyleftdev/jev-labs)): + aggregation (quorum, stability, escalate) is the spec, + not a multiplied joint of five Nouls. `notes.md` §67. ## Rules that hold across every position @@ -143,10 +171,22 @@ Reusable shapes when generating applications: meeting action items ~150 ms after each utterance. 7. **Formula embedding**: JUDGE/SCORE/CHOOSE as first-class spreadsheet formulas. 8. **Pixel-free computer use**: accessibility tree → compact actionable-JSON → one - batched question set per step → execute via AX actions. + batched question set per step → execute via AX actions. Encoder-backend + cousin: gliner2-ultrafast scores observed a11y/DOM controls with + local GLiNER2; code clicks; `DONE` ≠ success (`notes.md` §52). + Specialist-form cousin: Cua-S1 option-attention (fill/check/click/skip); + not TypeSafe Jev; plan ≠ execute; source-only (`notes.md` §54). 9. **Shadow-mode harness** (jev-harness): policy + confidence gate + shadow mode + offline eval CLI replaying fixtures, asserting on actions; 24-row filter 48.9 s - (Claude CLI) vs 1.3 s Jev at concurrency 8. + (Claude CLI) vs 1.3 s Jev at concurrency 8. Compaction rollout: + gliner25-compaction public default `shadowMode: true` (log proposed + reduction; do not replace history) (`notes.md` §50). + Stdout-prune cousin: [jev-pruner](https://github.com/tamaratran/jev-pruner) + archives full stdout before scoring; fail-safe keep original + (`notes.md` §53). Marketplace id still `fast-jev-output`. + Recovery cousin: [jevons](https://github.com/LilDojd/jevons) default + recovery **shadow** (record, do not interrupt); steering never + generates commands (`notes.md` §51). Calibration warning (calibre): routing thresholds and ROI do **not** transfer across datasets — every gate is a per-dataset measurement (see validation.md). @@ -155,10 +195,124 @@ datasets — every gate is a per-dataset measurement (see validation.md). Jev as gate/selector/verifier *around* a generator, never instead of one. Cost-sensitive prefilter (drop chunks/lines/hunks before the LLM); tool/skill routing (Choice + fits-Noul, code dispatches); preference lint - (project-defined rules as criteria). Fail-open vs fail-closed is per + (project-defined *soft* rules as criteria; linter owns hard rules; + Abide is the productized path, `notes.md` §47). Fail-open vs fail-closed is per action — LlamaIndex Jev rerank fails open (keep retrieval order), select fails closed. Full card: `references/mixed-architecture.md`. -11. **Structural prove ∩ remainder judge** (jevgate, doc-router): code - (allowlist, text layer) decides the easy cases; typed questions only +11. **Structural prove ∩ remainder judge** (jevgate, doc-router, Abide): code + (allowlist, text layer, linter) decides the easy cases; typed questions only on leftovers; fail-open unless a real sandbox sits under. Full card: `mappings.md` §18. +12. **1-token selector / tree of Choices** (chakuho, jev-gpt): the + generator is reduced to a next-label or next-word Choice. Softmax + over declared labels is not a Noul. Numeric rules and writing stay + exact. Full cards: `judgment-class.md`, `mixed-architecture.md`. +13. **Productized System One HTTP** (classifier-dev): the public + contract is label + calibrated confidence, not a paragraph. Batch + state, escalate-under-threshold, and a `FALLBACK` marker are + *code*. Full cards: `mixed-architecture.md`, `validation.md`. +14. **Evidence-synthesis pointer** (choxos/jev-reviewer, ≠ egma-ai): + Jev picks line ids; code copies verbatim; a second absolute Noul + checks "does this line itself answer?"; *Not found* is an answer; + the human tick is the product. Full cards: `applied-mappings.md` + §2, `mixed-architecture.md`. +15. **Prompted-JSON wire** (githubnext/localjev, ≠ kunchenguid/local-jev): + the TypeSafe SDK talks to a local `/v1/systemone`; the model + *writes* probabilities rather than exposing logits. Entropy + confidence is computed in code from that vector. Full cards: + `judgment-class.md`, `mixed-architecture.md`, `validation.md`. +16. **Script-before-p router** (NandhaKishorM/laya packaging of Hub + Laya): pick the checkpoint from script/lang/task *before* the + forward pass, because confidence will not drop on OOD (Khmer + 0.000@0.952). Post-T ECE is not raw ECE; 0.85 is still soft. + Full cards: `judgment-class.md`, `mixed-architecture.md`, + `faq.md`, `validation.md`. +17. **External class census** (@airesearch12 / Benchmark + Heaven): a named list is not a rank; GLiNER2 and + routers on the list are a class-boundary, not + identity; incompleteness is lag. Watch the board + URL; do not paste live scores into the census card. + Full cards: `mixed-architecture.md`, `faq.md`, + `validation.md`, `toolbox-mapping.md`. +18. **Geometric-mean product** (JevBench v1.2): four + axes at 25% each; a weak axis cannot be bought + back; weighting is a product design, not a law; + calibration on the rank is a choice (v1.1 kept it + off); instruction models in the same table as NAR + rebuilds; ×2 latency and est. costs are assumptions + to name. Full cards: `mixed-architecture.md`, + `validation.md`, `faq.md`, `mental-models.md`. +19. **Already-folded class as a recipe** (hourly 0842): + when the named HIGHs are already on the branch, + extract how-to-apply instead of re-carding — + wire-compat ≠ logit-equiv; productize label+p and + mark `FALLBACK`; packaging ≠ new species / script- + before-p; pointer-not-generator (two-pass; *Not + found*; human tick); external census ≠ scored + bake-off / geo-mean weights are a design. Skip + thin noise. Hard-gating a Noul as a PR/quality + gate is soundness theater. Full cards: + `mixed-architecture.md`, `faq.md`, + `mental-models.md`, `validation.md`. +20. **S1 keeps flying / S2 one-use** (khordoo/jev-reflex-autonomy-lab + delta of §46): position 4 (Selector of next + *action*) stays on the reflex every tick; + position 7 (state estimator) is the optional + planner — advice, not a command. Escalate- + under-threshold **without stalling**. Log + consumption, not arrival. Local rule-based vs + Live API is an A/B of backends, not a scored + bake-off; Local controller **≠** githubnext/localjev. + Seed = geometry ≠ async replay. No pixels. + Confidence ≠ selected probability. 20% still + soft. S2 never grants. Full cards: + `mixed-architecture.md`, `faq.md`, + `mental-models.md`, `agent-self-assessment.md`, + `validation.md`. +21. **OCR+AX observe→score→act** (awlevin/typesafe-computer-use): + position 10 (Discretizer: screen → numbered items) + then position 4 (Selector of next action). Writer + is generation, not a Jev position. Split kind/item/site + is width-is-cheap. Overlap is concentration theater. + Perception in code rebuilds pixel-free reasoning. + Decision never ships screenshots; the answer reader + may. 155× is one screenshot *theirs*. Full cards: + `mixed-architecture.md`, `faq.md`, + `applied-mappings.md` §9, `validation.md`. +22. **ASR observe→score→act** (moritzkremb/jev-voice-browser): + position 10 (Discretizer: waveform → transcript + + numbered elements) then position 4 (Selector). + Width-is-cheap: 9–11 questions on one request. + Partial-speech wait is VOI (closed-set vs free-text). + Spoken confirm is not a grant. Overlay numbers are + exact, not a second model. Compose with item 21 + (OCR). Full cards: `mixed-architecture.md`, `faq.md`, + `applied-mappings.md` §9, `validation.md`. +23. **Wrap-as-execution ALLOW/ASK/DENY** (reddpy/AgentGhost): + position 3 (constraint on the actuator path) then + position 4 (Selector on leftovers). Rules prove; + Jev remainder; ASK is an error not a log line. + Fail-closed on judge error. Distinct from actiongate + (evidence ≠ authority) and jev-use (fail-open). + Full cards: `mixed-architecture.md`, `faq.md`, + `applied-mappings.md` §7, `mappings.md` §8/§18. +24. **Application genre atlas** (@studio_yebisu): + position 11 (catalog of holes, not a score). + Stars are research-time. Same discipline as item 17 + (class census ≠ bake-off). Full cards: + `mixed-architecture.md`, `faq.md`, `validation.md`. +25. **External pedagogy / how-to-apply** (@akshay_pachaar): + position 11 (explainer of already-owned placements, + not a new construct). LLM hammer; code owns + branches; schema-safe ≠ correct; shadow + + questions-as-code. 200×/400× are TypeSafe ceiling. + Full cards: `mixed-architecture.md`, `faq.md`, + `mental-models.md`. +26. **Boolean composition of soft Nouls** (uehaj/jev-semgrep): + position 5 (Comparator over lines) then code ∧/∨/¬ + on *thresholded* bits — never multiply parallel p + (the logical-operator caveat above). Proposition ≠ + embedding (contrast-set refund). Not a Gate (position + 3). **≠** jev-combinators digital-design metaphor + **≠** semgrep.dev. Full cards: `mixed-architecture.md`, + `faq.md`, `applied-mappings.md` §4, `mappings.md` §4. diff --git a/.agents/skills/augustus/references/faq.md b/.agents/skills/augustus/references/faq.md index e2ffc7c..6d208ca 100644 --- a/.agents/skills/augustus/references/faq.md +++ b/.agents/skills/augustus/references/faq.md @@ -60,10 +60,11 @@ needs a paragraph, the generator re-enters. No. Augustus designs for the whole class of fast/cheap categorization-classification-scoring models. TypeSafe Jev is the documented exemplar (typed Choice / Score / Noul, live docs). Neighbors -in the class — open System-1 / decision-model heads (Laya, openjev-lm, -encoder DeBERTa, LoRA distill; Hume's 27B drop is Watch), constrained-AR -(TypeAR), GLiNER/GLiClass encoder family (locate vs categorize vs local -multi-head), listwise/pairwise rankers, vision scorers — are substitutes +in the class — open System-1 / decision-model heads (Laya, kev, +blackwood-rlcd, openjev-lm, encoder DeBERTa, LoRA distill; Hume's 27B drop is Watch), +constrained-AR (TypeAR), GLiNER/GLiClass encoder family (locate vs +categorize vs local multi-head), listwise/pairwise rankers, vision +scorers — are substitutes or cousins. Pick the family from the hole, then the vendor (`judgment-class.md` species map and when-to-use table). `typesafe-ai` still owns *Jev* contracts; other families own their own @@ -85,19 +86,32 @@ on fresh rows measures *agreement with the teacher*, not gold. Self-eval on your own independent labels before you treat it as a decision API (`notes.md` §25). +[jaredpalmer/kev](https://github.com/jaredpalmer/kev) is the laptop-local +System One **API drop-in** on that same open path: Qwen2.5-0.5B LoRA + +pointer, public gold not a Jev teacher, official SDK with a `base_url` +change. Hub weights: [`jaredpalmer/kev-0.5b`](https://huggingface.co/jaredpalmer/kev-0.5b) +(`notes.md` §45 delta). Use it for development and eval. Do not use 0.5B ID ECE as a +knowledge or frontier substitute (`notes.md` §45). + ## Open weights vs Jev vs constrained decoding vs encoder vs LoRA? Five surfaces, not one family (`judgment-class.md` when-to-use table). Three *open* paths sit beside proprietary Jev: **encoder** open-jev (DeBERTa, public gold), **AR constrained decode** (TypeAR; native [pcdServer](https://github.com/stephanj/pcdServer) GGUF serving), -**trained decision-only** (Laya / Nimble / Archer Watch). Proprietary Jev is the documented decision API; you do not hold the +**trained decision-only** (Laya / Nimble / **kev** / **blackwood-rlcd** / +Archer Watch). Proprietary Jev is the documented decision API; you do not hold the weights, so checks around the boundary stay black-box -(`formal-methods.md`). A trained decision-only open head (Laya, -openjev-lm, encoder DeBERTa, a LoRA student) copies the Choice / Score / +(`formal-methods.md`). A trained decision-only open head (Laya, kev, +openjev-lm, encoder DeBERTa, a LoRA student, **blackwood-rlcd**) copies the Choice / Score / Noul *shape* and moves eval onto you. Distills trained on Jev's *answers* (openjev-lm, jev-gate-student-b) are teacher-copies — read -agreement separately from gold. Encoder open-jev +agreement separately from gold. **Domain-jev-maker** is also not that +distill: independent CLINC gold, soft targets; pick it when downstream +*reads* p, few-shot hosted when only argmax (`notes.md` §60). **kev** is not that distill: CE on +public labelled outcomes, pointer readout, isolation probes, ID ECE +0.065 (0.031 after temperature scaling) on 1,350 questions — still +self-eval, still not OOD (`notes.md` §45). Encoder open-jev ([DeBERTa-v3-large](https://huggingface.co/com-kotobalabs/open-jev-deberta-v3-large)) was trained on public gold, not Jev; in-domain ECE 0.022, OOD acc 0.854→0.690. Constrained autoregressive decoding (TypeAR; README names @@ -109,7 +123,23 @@ brittleness; compose with abstention and an allowlist gate (`mappings.md` §2, §17, §18). Hume's announced open **decision-model** (27B dense, multimodal, AU healthcare residency — not anti-TypeSafe) is **WATCH** until weights, license, and evals exist (`notes.md` §31, -§33). Constrained decoding is §32; native serving is §42. Public logit dump for the +§33). Open multimodal *decide* that already shipped: +[blackwood-rlcd](https://huggingface.co/BlackwoodAI/blackwood-rlcd) +(CC BY-NC; not that drop; `notes.md` §46). A local `POST /v1/systemone` +drop-in ([jev-local](https://github.com/us/jev-local)) is a **surface**, +not a fourth path — default scorer is a stub until `hf` (`notes.md` §48). +[jeff](https://github.com/logan-markewich/jeff) is a GLiFormer encoder +behind the same wire (not a Jev replica; `notes.md` §60). +[sysone](https://github.com/hraness/sysone) is a loopback **router**, +not a scorer. Distinct name collision: +[sysone-help/sysone](https://github.com/sysone-help/sysone) is an +evaluation-model-first TypeScript SDK (predicate/classifier/rubric +as data; cancellable; never auto-retry; first adapter Jev via +Vercel AI Gateway). Do not merge the two (`notes.md` §60, §63). +Laya ONNX port: [laya-onnx](https://huggingface.co/Mattepiu/laya-onnx) +(do not copy the inherited vs-Jev table). Constrained decoding is §32; native serving is §42. Decision-token QLoRA on that graph: +[Foodoo1/Qwen3-14B-RLCD-Decision-LoRA](https://huggingface.co/Foodoo1/Qwen3-14B-RLCD-Decision-LoRA) +(train the decision token, not prose; synthetic fraud receipt). Public logit dump for the read-the-letter graph: mini-jev-runs. "Smarter than Jev" is a claim. He prefers "decision models" over "system one"; this skill still quotes TypeSafe's name for the exemplar. Before you pick any of those paths, @@ -125,6 +155,27 @@ Hole first, logo last. These are **species**, not aliases code. Paper: [GLiNER](https://arxiv.org/abs/2311.08526). Local multi-head GLiNER2.5 (fastino-ai) can also classify and extract relations on a laptop — discourse, not a measured 36× (`notes.md` §25). + Compaction receipt: [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) + uses that local multi-head as **categorize** (retention action) plus + **locate** (character-offset spans); code copies; not a summarizer + (`notes.md` §50). + Indexer cousin: [s1-graphify-indexer](https://github.com/GreyssonEnterprises/s1-graphify-indexer) + uses GLiNER2 on the bulk of a repo graph and escalates an LLM only + on the ambiguous tail **if the backend loaded**. GitHub one-liner + 10–50× is a **target, not a measured speedup** — table TBD + (`notes.md` §51). Not a Noul. + Computer-use receipt: + [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) + uses GLiNER2 (`fastino/gliner2-multi-v1`, not 2.5) to **score among + observed** a11y/DOM controls; code acts; not a screenshot model + (`notes.md` §52). Same observe→score→act hole as Jev Ultrafast. +- **GLiFormer (encoder serving the System One *wire*):** + [jeff](https://github.com/logan-markewich/jeff) on + gliformer-large-v1 (400M) answers choice/score/noul at + `/v1/systemone`. Related Knowledgator encoder lineage, **not** + GLiNER locate and **not** a Jev replica. Cheaper on GPU, + less accurate on reasoning-heavy; CPU is *more* expensive. + `notes.md` §60. - **GLiClass (categorize):** one forward pass over text + *all* labels; sigmoid multi-label or softmax single-label. Use for large or changing tag sets. Scores are class affinities, not automatically a gateable @@ -134,7 +185,13 @@ Hole first, logo last. These are **species**, not aliases limits (e.g. 255-way Choice) are Jev's, not the class's. Distilled open heads (openjev-lm, jev-gate-student-b) copy the *teacher*, not independent gold. Encoder open-jev (DeBERTa) is the same *shape* on - public gold — still self-eval, especially OOD. Hume's 27B drop is + public gold — still self-eval, especially OOD. **kev** is the + causal-decoder + pointer productization of Archer's reconstruction on + public gold (API-compatible; not a teacher-copy). **Domain-jev-maker** + is a domain LoRA on independent CLINC gold (not a teacher-copy) — + train when downstream reads p; few-shot hosted when only argmax. + **blackwood-rlcd** is + the open multimodal decide head (CC BY-NC; not Archer Watch). Hume's 27B drop is Watch. When-to-use axes: `judgment-class.md`. - **Cross-encoder / listwise ranker:** order of a retrieved shortlist. Fail **open** (keep retrieval order). Translation-invariant listwise @@ -161,7 +218,11 @@ labels on a GLiNER2 encoder, not Choice / Score / Noul). Empirical open encoder next to GLiClass; not a weight clone. "like jev" is discourse. A GLiGuard score is not a proof. LLM I/O safety is not a coding-agent tool gate (rh-guard for reward-hacking; jevgate shape for allowlist -∩ remainder). `judgment-class.md`. +∩ remainder; Abide for project soft rules on diffs; +gliner25-compaction for extractive context compaction — Fastino +sibling class, not GLiGuard; +gliner2-ultrafast for scoring observed browser controls — GLiNER2, +not 2.5, not a safety schema). `judgment-class.md`. ## Can I threshold CLIP / SigLIP as a safety gate? @@ -256,8 +317,13 @@ sentences, and that is the point ([Langfuse framing, 2026-09-18](https://x.com/langfuse/status/2100980004678971491)). When you need a paragraph rationale, a trace UI, or an annotation workflow, generation and the eval platform still own those seats. -Verbal LLM scores are uncalibrated. Do not thin this skill into a -Langfuse how-to. Mixed architecture: traces stay; the judge step can +Verbal LLM scores are uncalibrated. The Harbor/jevals-adjacent +practice is: shadow mode + fixtures that assert on the **action**, +not on prose (`jev-harness`, `validation.md`; `notes.md` §44). Shared +bake-off this hour scores **ECE / NLL / Brier**, not an LLM paragraph +([open-jev-laya-bench](https://huggingface.co/datasets/pngwn/open-jev-laya-bench); +`notes.md` §46). Do not +thin this skill into a Langfuse how-to. Mixed architecture: traces stay; the judge step can be a System One model. ## Allowlist first, then Jev? @@ -270,7 +336,1057 @@ first so a comment can talk it into a write is the rejected design. The same three-way test, one hour later: if a regex, a DNS lookup, or a database query already answers, **do not call a model** (`wotai-dev/typesafe-jev-tools`, `notes.md` §42). That is meta-VOI, not -a hook tutorial. +a hook tutorial. This hour's wording of the same sandwich: the allowlist +**proves** read-only verbs; Jev judges only unlisted leftovers; +fail-open (cannot block) (`notes.md` §46). Same family, different +remainder: a **linter proves** lintable rules; [Abide](https://github.com/coldteadotai/abide) +Scores residual soft AGENTS.md / CLAUDE.md rules; fail-open, banded +(`notes.md` §47). Soft judgment is never the sole hard veto. + +## Soft project rules — Jev or the linter? + +The linter owns what it can prove. Soft instruction-file rules +("no helper with one caller", "don't add what wasn't asked") are +residual judgment. Same layering as jevgate: structure first, typed +Score only on the remainder; **fail-open**. Name the observation +window (edit vs turn). Fix false positives in the rubric, not the +model. Productized path: [Abide](https://github.com/coldteadotai/abide); +earlier contract pointer: jev-pref. Complementary, not the same +product: [rh-guard](https://github.com/24601/rh-guard) (reward-hacking / +eval integrity). Plain-English PR check: +[if-ai](https://github.com/Victor-Casado/if-ai) (one condition + +required min-confidence; fail-closed on error). [jev-marshal](https://github.com/LightningK0ala/jev-marshal) +is Watch / empty this pass. Request-shape lint still sits upstream (wellposed / +`tenbin`). `mixed-architecture.md`; `question-design.md`; `notes.md` +§47, §51. + +## Can confidence gating catch a forced wrong Choice? + +No. Choice probabilities are conditional on the offered set. If coverage +is open and you omit `other`, the model must pick a listed option — and +the distribution can peak at **1.00 on the wrong label**. Downstream +confidence gates see a healthy answer. Lint the *request* (missing +escape hatch, broken state paths) before you trust the number. Recipe: +[wellposed](https://github.com/suraj-phanindra/wellposed) (unsubscribe +email → `"support issue"` at 1.00 without `other`; overlapping options +collapse to 0.19 — that failure is loud). `tenbin` still owns the +design-time lint *skill*; Augustus owns the placement. +`question-design.md`; `notes.md` §46. Putting `"other"` on the request +is necessary and not sufficient for an open head you train: the +residual option must also appear as a **wrong** alternative, with +varied wording, or the hatch becomes a shortcut +([kev](https://github.com/jaredpalmer/kev) first-run lesson; +`none_of_the_above` eval; `notes.md` §45 delta). + +## Should Jev live inside the database? + +The *hole* is a semantic index over a structured store: cheap exact +predicates first, typed questions on the remainder. Two serving +choices, same hole (`mappings.md` §4; `notes.md` §44): + +- **In-engine extension** (`sqlite-jev`; cousin `pg-jev`): SQL sees + `jev()` / `jev_rows`. Convenient. The database process now has an + API key, a spend guard, and a residency problem. +- **Out-of-process CLI** (`jevql`): vanilla Postgres never sees + `jev()`. The rewrite layer owns the call. + +Neither is an index. Full-scan the post-filter remainder. Row contents +leave the store. zoxide/`joxide` is the same hole over paths. Do not +copy SQL. + +## Can Jev pick the bitrate, the join order, the model? + +Yes as a **proposal inside a hard envelope**, no as the actuator. +bitrate-advisor: Jev may only match the deterministic cap or be more +conservative; missing the model returns the policy's answer. +mmalisper's JOB planner: Postgres plans first; Jev overrides only when +confident (+12% geomean author-reported; join-order Choice alone was +2× slower). routeKit / jev-claw / Higgsfield: classify requirements; +code picks the generator. Compaction cousin (encoder, not Jev): +gliner25-compaction — mutating tools / shell operators prove +`keep_full`; the model may only match that or be more conservative; +uncertain fails closed to `keep_full`. The envelope is load-bearing +(`mappings.md` §12, §15, §18; `notes.md` §50). + +## Wait for Archer to ship omni System One? + +No. Archer's 27B dense drop is still **Watch** (no Hub weights this +pass; user watch ~2026-09-19). Omni perception→decision already has an +open model: [`BlackwoodAI/blackwood-rlcd`](https://huggingface.co/BlackwoodAI/blackwood-rlcd) +(CC BY-NC; Jev-compatible shim; screenshot + marked candidates → +Choice). Soft judgment over pixel candidates inside deterministic +code. Jev still leads general *text* (0.850 vs 0.786 on their 8,456-item +table). Specialist composition (SAM / OCR → text → Jev) remains valid. +Do not wait, and do not treat screenshot-vs-Jev-text as the same input. +Pixel-free computer-use (DOM/a11y candidates → score → code acts) does +not wait either: [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) +is that hole with a local encoder (`notes.md` §52). +`judgment-class.md`; `notes.md` §46, §52. + +## When does a decision model hold? + +When the answer is **extractable from the state you feed it** +(classification, citation/paraphrase/reversed-meaning with claim + +quote both given, sarcasm whose trigger is in the text, DOM-as-text +fan-out). It **fails, often confidently**, on recall without a +supporting passage. Atlas history suite (`notes.md` §49): Case A +wrong @ 0.90 with no context; Case B near-flat 0.07 (correct by luck); +Case C — same item as B — right @ 0.97 once the passage is in +`state`. Retrieve first. Overlapping categories can still be +**dangerous-high** (DAIR Emotion 48% acc / mean conf 0.819). Not a +leaderboard; one axis. `mental-models.md` §boundary. + +## Are the internals a state machine? + +No. Distributed LM understanding, not hand-written transitions +(paraphrase vs reversed-meaning contrast). **Placement** is a +component node in *your* code — that instinct is right; "it *is* a +state machine" is not. `mappings.md` §3; `mixed-architecture.md`. + +## Decision model or a constrained LLM? + +Neither class won on quality in DMB v2 (`notes.md` §49). Decision +models win the *axes you actually buy*: p50 ~264–276 ms, ~$0.07/1k +on banking, schema-valid, flat to 255 options — and **fail at 256+**. +Constrained LLMs handle 512 and (most of them) admit ignorance on +no-good items; they are slower and cost more. Spam: two OpenAI models +sat *below* the majority baseline. Re-run on *your* labels. Do not +merge Banking77 87% / 76.3% / 79.67% across protocols. +`judgment-class.md`; `validation.md`. + +## Can many cell-wise Choices solve a combinatorial grid? + +Not in the Direct Jev ARC-AGI-1 experiment: **4/400 (1%)**, ~$2.32. +Dimensions ~90%; complete grids rarely. Combinatorial assembly ≠ +extractive keep/drop. Search or a program stays in code. +`mappings.md` §9; `notes.md` §49. + +## Is a local `/v1/systemone` the same as Jev? + +No — not until you know **which scorer** is behind the socket. +[jev-local](https://github.com/us/jev-local) is a **contract-compatible** +drop-in (`base_url`). The **default scorer is a deterministic stub** +and carries no intelligence. `JEVLOCAL_SCORER=hf` turns on a frozen-model +logprob scorer. Their README: an interface-compatible baseline, not a +reproduction of Jev's undisclosed model. kev is the other local +drop-in (trained pointer head, public gold). [von](https://github.com/wfzyx/von) +is a **tiny SAN** (14 MB needle; authored144 52.6%) at the extreme of +the speed/econ class — not the stub, not kev, **not a calibrated Jev +replica**. Do not copy its vs-Jev table. +[open-alternative-jev](https://github.com/ikermoel/open-alternative-jev) +packs one-forward logprobs on an open LLM you already have (RACE-H +92.9% @ 4.55 q/s); **not a Jev reproduction**. +[jevify](https://github.com/Mintzs/jevify) is a CUDA/PyTorch cousin +on Qwen2.5-1.5B (`ora_decision_engine`): CUDA graphs, branch kernels, +literal-label scoring. **Uncalibrated model likelihoods, not +measured correctness** — softmax over A/B/C is not a Noul. Independent +of Distillation. No LICENSE this pass. Default refund workflow is +not a validated policy. A green smoke test on +the stub is not a bake-off. `judgment-class.md`; `notes.md` §48, §49, §55. + +**Independent `/v1/decide` (not this wire; 2026-09-19 ~04:39):** +[OpenJev](https://github.com/IamBusy/OpenJev) speaks a +**different** contract. 45/60 *theirs*. Not TypeSafe. Distinct +from hraness/sysone OpenJev runners. +[semif-serve](https://github.com/dddanielliu/semif-serve) +is another `/v1/systemone` surface (SemIf runoff; 1164 vs +178 ms *theirs*; wire-compat ≠ replica). `notes.md` §69. + +**Wire-compat encoder cousin (2026-09-18 ~20:43):** +[jeff](https://github.com/logan-markewich/jeff) serves +`/v1/systemone` on GLiFormer-400M; `typesafe-sdk` drop-in via +base URL. **Not a Jev replica** — normalized sigmoids, T=3.2, +noul isolation default, DeBERTa token counts. Their card: L4 +HTTP ~$2.6 vs jev ~$15.6 per 1M requests (~6×); A10G direct +~$0.65 (~24×); AG News 75.5% vs 90.5%. CPU arm is *more* +expensive. License null this pass. A loopback **router** +([sysone](https://github.com/hraness/sysone)) is not a scorer +either — it routes hosted + local OpenJev/NanoJev/Mini-Jev +and does not run weights. `notes.md` §60. + +## Are local CUDA likelihoods a Noul? + +No. [jevify](https://github.com/Mintzs/jevify) (and packed-logprob +cousins) return **uncalibrated model likelihoods**. Do not threshold +them as P(permit) or as calibrated abstention. Temperature / ECE on +*your* labels if you use the surface. Softmax over allowed tokens ≠ +Noul. `judgment-class.md`; `notes.md` §55. + +## Should RAG stop at Top-K / a reranker? + +Not if the hole is **evidence**. +[decision-native-rag-skills](https://github.com/emergency-lee/decision-native-rag-skills): +retrieve wide → decide explicitly → evidence set → resolve conflicts +→ reason only over kept evidence. Embeddings stay candidate +generators. Provider-agnostic; no bundled harness; **no universal +benchmark**. Default migration gates are starting targets, not +promises. Do not ship because an LLM judge prefers it. +[jev-sift](https://github.com/kbhuw/jev-sift) is the same sandwich +on **agent I/O**: classify first, read selectively; content to Jev +without entering main agent context first (paths/URLs). Uncertain / +errors / truncation ≠ irrelevant. Transport tests ≠ accuracy. +`mappings.md` §4; `notes.md` §55, §56. + +## Dump files into context, or classify first? + +Classify first when the items are **not** already in the main agent +context. [jev-sift](https://github.com/kbhuw/jev-sift): batch path / +public URL / inline text → Jev; the main LLM opens survivors. +Uncertain → closer look; errors and truncation ≠ irrelevant. Inline +text the agent already read cannot recover that cost. Hard envelope +in code (50 / 60k / 2MB / public-IP). Transport tests ≠ accuracy. +Same family as decision-native RAG. Topology A MCP — not +jev-routing (host adapter). Do not copy plugin how-to. +`applied-mappings.md` §1; `notes.md` §56. + +## Is jevable.com a 342-title census? + +No. [jevable.com](https://jevable.com/) is a living **applied-mappings +atlas**: extract class patterns (intent columns, score-among-observed, +VOI gates, generative UI decide, robotics text-state, draft-gate fail +modes). The site claims **342** curated projects this pass; homepage +JSON-LD lists **36** (first page / `pageSize` 36). We did not enumerate +titles. Maker clocks stay **claims** unless already a named receipt. +Cross-link exemplars already in notes; do not dump a hit list. Not a +model. Not multimodal substrate. Archer still Watch. +`applied-mappings.md`; `notes.md` §56. + +## Did Jev beat nano as an escalation gate? + +Not in the pre-registered independent eval +[jev-baselines-eval](https://github.com/ickma2311/jev-baselines-eval) +(2026-09-18). **Both experiments AMBIGUOUS.** Cascade Δ +0.265 at a +1pp-below-frontier target; **at exact parity the sign flips** +(R_jev=1.000 vs nano 0.730) because Jev confidence is exactly 1.0 on +102/200 items including 6 wrong. AUROC error-ranking neither +direction established; **no ECE**. Encoder with labels wins Banking77 +(0.933 / 9 ms). Recorded call duration ~2.2×, **serving-path not +model-speed**, not 40–200×. Same-day errata three rounds. This is +the jevals/Harbor practice exemplar this hour (honest negative + +calibration theater). `validation.md`; `notes.md` §55. + +## Did Jev pay off for Precision PDF extraction? + +Not in +[databricks-jev-pdf-lab](https://github.com/laurentfabre/databricks-jev-pdf-lab). +**No quality-equivalent, end-to-end Jev payoff demonstrated.** +Compact requests cut tokens but changed 26/236 recommendations. +**No OSS license selected.** Typed output is not truth. `notes.md` §55. + +## Should the model write the quote / the citation / the click? + +No. Extractive keep/drop: code already holds the sentences, line ids, +character offsets, or numbered controls; the model **selects**; code +**copies or clicks**. +[testimonial-miner](https://github.com/AppitStudio/testimonial-miner) +assembles quotes from per-sentence Nouls and `redecide`s without new +calls. **[choxos/jev-reviewer](https://github.com/choxos/jev-reviewer)** +(systematic-review Jev Reviewer; **≠** egma-ai) points at ids; *Not +found* is an answer; human tick never overwritten. [solari-reflex](https://github.com/hitakshiA/solari-reflex) +never lets model output become a selector. Compaction is the same +species: [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) +copies exact source spans; a prose summary of the tool result is +generation, not keep/drop (`notes.md` §50). Command-output cousin: +[jev-pruner](https://github.com/tamaratran/jev-pruner) keeps verbatim +chunks of Bash stdout; dropped spans live in an archive, not a +summary (`notes.md` §53). Claim/evidence Stop: +[clear-head](https://github.com/VladyslavHontar/clear-head) judges +against retrieved session lines, not generated prose (`notes.md` +§51). Computer-use encoder +cousin: [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) +scores observed controls; code clicks; no generated selectors +(`notes.md` §52). Specialist-form cousin: +[Cua-S1](https://github.com/trycua/cua/tree/main/libs/cua-s1) selects +fill/check/click/skip among observed elements; does not generate +values or selectors (`notes.md` §54). Generation is only for +TYPE/prose when something must be written. `applied-mappings.md` §2; +`notes.md` §48, §50, §52, §53, §54. + +## Should compaction summarize? + +No. Pointer/extractive compaction and generator summarizers are +different species. The former is auditable (every kept byte occurs in +the input). The latter can invent. Same job as +fast-jev-compaction / pi-jev-compaction (Jev Noul/Score backends); +GLiNER2.5 is an encoder backend. Mutating tools and shell operators +stay `keep_full` in **code**. Low-confidence / invalid evidence fail +**closed to `keep_full`** — the *reduction* is the irreversible act, +unlike Abide / jevgate fail-open. Ship `shadowMode` first (default +true: log, do not replace history). Not Jev. Not multimodal. +Stdout prune is the same *family* (evidence-preserving reduce) on a +**different job**: [jev-pruner](https://github.com/tamaratran/jev-pruner) +Noul-prunes a just-run Bash result before the main LLM sees it; +fast-jev-compaction / gliner25-compaction compact completed tool +pairs already in history. Hard ≤10k / JSON-diff-whole-doc envelope +in code; fail-safe keep original; archive for recovery. Marketplace +id still `fast-jev-output`. Framework-agnostic middleware cousin: +[jev-compactor](https://github.com/edwardyen724-g/jev-compactor) +— **Jev judges relevance. Code decides structure.** Never rewrite. +Regex floor in code. Compaction fails open if Jev is down; safety +gate fails closed on pending destructive/exfil. One-session +*theirs*: later product-arm table **73%** / 350 ms / +$0.0004 / 4 of 4 vs shipped summarizers (30–250× cheaper); +earlier vs-Sonnet card 64.5% / 366 ms (`notes.md` §65). +Claude Code shorter path remains fast-jev-compaction. +`judgment-class.md`; `notes.md` §50, §53, §65, §68. + +## Is pruning Bash stdout the same as compacting session memory? + +No. Same family (pointer, not summarizer; dropped bytes recoverable). +Different job. [jev-pruner](https://github.com/tamaratran/jev-pruner) +scores chunks of a command that just ran, *before* they enter the +generative turn. [fast-jev-compaction](https://github.com/tamaratran/fast-jev-compaction) +and [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) +reduce completed tool pairs already in session history (Jev Noul vs +GLiNER2.5 encoder). Host capability shapes the product: Claude wraps +Bash automatically; Codex is an opt-in wrapper because it cannot +replace native shell output from `PostToolUse`. `applied-mappings.md` +§1; `notes.md` §50, §53. + +## Fail-open or fail-closed — which? + +Name the **irreversible act**, then pick polarity. Compaction *drop* +fails closed to `keep_full`. Stdout *prune* fails closed to original +output ([jev-pruner](https://github.com/tamaratran/jev-pruner): +archive/Jev/incomplete-score failure keeps the log; Harbor plugin-eval +cannot reach Jev and therefore cannot prune). Wake *skip* fails open (wake on error / +unsure): [wakegate](https://github.com/shitianfang/wakegate) skips +only if Jev answers and p(wake) < 0.2. Merge *PASS* on a red run +fails closed at the gate: [latch](https://github.com/CaseReed/latch) +`--gate` BLOCKs unless infra is confirmed; the Playwright reporter +stays fail-open. [if-ai](https://github.com/Victor-Casado/if-ai) +fails the Action on error / empty / low confidence. jevgate cannot +block; pi-jev-approver fails closed without a key; Abide is fail-open +on diffs. Session-memory *omit* fails open (dump the ledger): +[carryforward](https://github.com/Dharundp6/jev-carryforward). +Tool *execution* fails closed on block/timeout: +[toolgate](https://github.com/fdemir/toolgate) (Jev is not +authorization). OMP/pi +[omp-jev-extensions](https://github.com/luw2007/omp-jev-extensions) +**fail open** (`confidence: 0`) if Jev is missing — contrast +pi-jev-approver fail-closed without a key. Ruby validations in +[hunch](https://github.com/carldaws/hunch) `rescue nil` at save — +spam gates should not. Draft-gate *silence* is the same rule: +missing verdict is not a block and not a pass — fail-open / +heartbeat, do not hold forever ([jevable.com](https://jevable.com/) +class pattern). Contrast Abide `<0.5` silence (the *edit proceeds*). +Same sandwich, opposite authorized act. +`notes.md` §50, §51, §53, §55, §56, §58. + +## Type-safe or correct? + +Typed answers are not a safety case. +[interlock](https://github.com/somoore/interlock): type-safe ≠ +correct; irreversible stays behind a threshold **and** a human. +Secrets never enter the agent. Jev is a SENSOR; `policy.py` +decides BLOCK/ASK/ALLOW. A launch-week firewall that asks +"dangerous?" after the LLM already decided, with real secrets in +scope, is the named anti-pattern. Distinct from +[toolgate](https://github.com/fdemir/toolgate) (pre-exec of a +proposed call). Atlas receipts make the same split a +**measurement** claim: schema-valid (cannot emit off-list) +≠ picked-right (DAIR Emotion 48% at mean conf 0.819). +`mappings.md` §8; `applied-mappings.md` §7; +`notes.md` §59, §49, §66. + +## Is Jev the policy? + +No. Jev is the sensor. Policy (checklist, ledger, `policy.py`, +two-person rule, human confirm) is the constraint. Interlock: +the model never picks allow/ask/block. +[port-cleanup](https://github.com/epiphany-dynamics/port-cleanup): +the human is the only kill trigger; shields override; displayed +explanations are app-owned mapped text, not raw model prose. +`notes.md` §59. Same split: [omp-greenlight](https://github.com/SemetricLabs/omp-greenlight) +is permission vs probability (operator owns the bar); [skill-broker](https://github.com/adamjralph/skill-broker) +outline is judgment ≠ permission (Jev never grants access). +`notes.md` §62. + +## Is Jev authorization? + +No. Jev is a probability, not a grant. +[toolgate](https://github.com/fdemir/toolgate): pre-exec +allow/block/review; Jev is not authorization. +[omp-greenlight](https://github.com/SemetricLabs/omp-greenlight): +suppresses an OMP approval prompt when Jev says allow — +**permission vs probability**. Operator owns the bar; the +plugin never self-tunes it. Host `bash.patterns: deny` stays +the floor. Default **40.9%** prompts removed / **0 of 94** +unsafe auto-approvals on the labelled corpus (not live +traffic). Not a sandbox. Composes with waymode and +omp-jev-extensions. `applied-mappings.md` §5, §7; +`notes.md` §62. + +## Does Jev grant skill access? + +No. Judgment ≠ permission. +[skill-broker](https://github.com/adamjralph/skill-broker) +is a **project-outline** (not a production recipe): +deterministic code owns catalog, policy, limits, and +grants; Jev scores relevance/confidence and **never grants +access**. Candidates ≠ grants. Jev down → foundation-only; +never broaden access. Distinct from shipped routers +(jev-hermes, GodsBoy, omp-jev-extensions). +`applied-mappings.md` §5; `notes.md` §62. + +## Is the Jev score the eval? + +No. Check the instrument, not just the score. +[dinostomp](https://github.com/collapseindex/dinostomp) +audits data, scorer, runs, numbers, claims, and itself. +`dinostomp jev` tests a Jev question like an if-statement +(accuracy, p(yes) cut, ECE, blank lean, rewording). Demo +*theirs*: 24 examples, ECE **0.062** — not a class ranking. +FINDINGS 189; 99 against itself. Cousin of rh-guard / +egma attention≠correctness / game-coach engine-owns-truth. +Beside jevals, not a Harbor taskset. `validation.md`; +`formal-methods.md`; `notes.md` §62. + +## Is Jev the sole hard gate on the hot path? + +No. Put System One on the **feature side of a constrained +optimizer**, never as the only gate that must answer before +traffic moves. +[slo-router](https://github.com/zeeshan8281/slo-router): +Jev supplies task / exactness / external-evidence; code +picks the cheapest backend meeting quality + SLO floors. +Fail-open to local features on timeout/invalid. Measured +*theirs*: same routes/accuracy as the local path; p95 +**77.93 → 490.38 ms**. Exactness raises the quality floor; +it must not override capability/context. **Hunch:** Harbor- +style measurement of decision-model latency is mandatory +before claiming “Jev routing.” Eight-row demo is not a +benchmark. `mappings.md` §6, §15; `notes.md` §63. + +## Is privilege the verdict? + +No. Privilege changes blast radius, not whether the act is +benign. +[construct-auto-classifier](https://github.com/godspede/construct-auto-classifier): +`sudo status` can be a safe read. Fast structural rules, +then Jev Choice + independent risk Nouls. Operator owns +`minConfidence` / `riskThreshold`. Fail-closed. +Certification *theirs*: Jev **0** dangerous / 975; every +chat model leaked. **Landed-script trust** (byte-identical +to the remote default branch) is a merge-gate receipt, not +a name. **Headless ≠ auto-approve.** Pair with dinostomp and omp-greenlight. +`applied-mappings.md` §7; `mappings.md` §18; `notes.md` §63. + +## Is System One a permission gate for human review? + +No. It can be an **attention filter / VOI** that never +blocks the agent and never says green unless sure. +[jev-lens](https://github.com/rashedInt32/jev-lens): +calibrated “do I need to look / which files / strip +debris?” Never edits files. Companion +[jev-lens.nvim](https://github.com/rashedInt32/jev-lens.nvim) +is a popup, no API key. Complements skill-broker (Jev +never grants access) and omp-greenlight (operator owns the +bar). Distinct from +[jev-gates](https://github.com/rashedInt32/jev-gates) +(stops writes). **Hunch:** minimize expected human cost +under false-green risk. `notes.md` §63. + +## Are the two `jev-lens` repos the same product? + +No. Qualify the owner. +[rashedInt32/jev-lens](https://github.com/rashedInt32/jev-lens) +is a Stop-hook **human** attention filter (never blocks +the agent). [dizk/jev-lens](https://github.com/dizk/jev-lens) +is **pre-send view selection** of tool results (79% fewer +tokens on 500 SWE-rebench trajectories *theirs*). +Compress-before-first-send beat post-send prune (cache +cost +17%). `notes.md` §63, §68. + +## Is a Stop-hook risk score a merge blocker? + +No. [jev-preflight](https://github.com/muse0509/jev-preflight) +redirects the agent's attention for at most one +reinspect, then finishes. Fail-open. Uncalibrated 0.85. +Not tests, not SAST, not latch / ci-gatekeeper. +`notes.md` §68. + +## Does exposing MCP tools mean the agent will recall? + +No. **tools≠use.** +[carryforward](https://github.com/Dharundp6/jev-carryforward) +eval: `recall` **0/4** with tools + skill installed. +SessionStart hook injects rules; hoping the model reaches +for memory is not a design. `notes.md` §68. + +## Is openvons TypeSafe Jev? + +No. [openvons](https://github.com/genai-craft/openvons) +is an independent open-Jev *class* (LM/vision/voice; +NOTA; execute/confirm/reject). Apache-2.0 code; GitHub +SPDX NOASSERTION. Speaks `/v1/systemone` as +**wire-compat**, not a replica (same warning as jeff / +jev-local). JevPick is menu decode (3.2–4.8× +byte-identical *theirs*), not a Noul. Unrelated to +TypeSafe; no TypeSafe API output used. `notes.md` §68. + +## Is OpenJev (IamBusy) TypeSafe Jev, or hraness/sysone? + +No to both. [OpenJev](https://github.com/IamBusy/OpenJev) is +an independent 0.6B LoRA+scalar head. `/v1/decide` is **not** +a TypeSafe drop-in. Training is supervised CE, not RLCD. +[hraness/sysone](https://github.com/hraness/sysone) "OpenJev +runners" are a **loopback gateway**, not this model. +[semif-serve](https://github.com/dddanielliu/semif-serve) +is another `/v1/systemone` **wire** (SemIf runoff; 1164 vs +178 ms *theirs*); wire-compat ≠ replica. `notes.md` §69. + +## Are the two Winnow repos the same product? + +No. Always qualify the owner. +[kevinpita/winnow](https://github.com/kevinpita/winnow) is +the launch-week **context sieve** (hide agent artifacts). +[ThinkyMiner/Winnow](https://github.com/ThinkyMiner/Winnow) +is a Chrome **worth-your-attention** VOI filter (read / +skim / save / skip from typed answers; 80%/90% *theirs*). +`notes.md` §69. + +## Can I skip the LLM when the intent is the same? + +Yes, as a **fail-open** VOI admit — never as a silent +rewrite. [jevcache](https://github.com/kushals256/jevcache) +asks Jev `same_intent` after exact SHA-256; 0 FP / recall +0.38 on n=100 *theirs*. Stream/tools/multimodal bypass. +A cosine cache with Jaccard@0.35 had fpr 0.48 on the same +fixture. `notes.md` §69. + +## Does a Noul distinguish conflict from ignorance? + +Not by itself. [jev-typed-evaluation-collapse](https://github.com/mleyvaz/jev-typed-evaluation-collapse) +(NCML field note v0.3 *theirs*): the same evidence yields +Noul 0.50–0.57 (conflict) vs 0.46–0.48 (ignorance); +Choice with named `conflicting_evidence` / +`insufficient_evidence` separates at p=1.0; binary Choice +without an escape is lexically biased (red 0.67–0.85). +Schema-as-interface. Same family as missing-`other` → +confident wrong. `notes.md` §69. + +## Does a typed baton grant the tool call? + +No. [jev-handoff](https://github.com/shitianfang/jev-handoff): +gate `allow` **never grants** — only deny/ask. Control +returns as escalate / continue / abort. Fail-open. Inverted +loop wakes the LLM only on escalate. Vercel drops +confidence. `notes.md` §69. + +## Jev WHETHER, Python HOW, LLM WHAT? + +Yes as a split, not as a stack replacement. +[hermes-jev-router](https://github.com/rsdkrasen/hermes-jev-router) +(license null; community plugin): Jev decides whether the +next main-model call is worth it; Python keeps original +chunks and suppresses duplicate observational tools; the +LLM still writes when writing is required. Skip-next needs +a Hermes core patch. Fail-open. `notes.md` §69. + +## Can I put a Noul on a lock or a heater? + +No. [HA-Jev](https://github.com/AboveColin/HA-Jev) +explicitly: a probability with no explanation should not +hold a lock, a heater, or a smoke alarm. Typed answers +as sensors, confidence gating, Jev-gates-LLM cascades — +yes. Safety actuators stay in Home Assistant interlocks. +`notes.md` §68. + +## Does measurement own endorsement? + +Yes, for question packs (and any cookbook criteria). +[jev-packs](https://github.com/dtduc-git/jev-packs): a pack +is only `verified` after recorded accuracy / ECE / cost / +latency on a **pinned** model version. No numbers, no +endorsement; otherwise `provisional`. Abstention/`unknown` +is mandatory. Named runner +[jevassert](https://github.com/dtduc-git/jevassert) +**LANDED** this pass (Apache-2.0; was 404 in §64): +record once, `check` offline from recordings, exit 0/1/2. +Calibration and cost are first-class gates, not footnotes. +Harbor/jevals pattern for the class — never +cookbook-once-and-forget. Distinct from INSTRUCT_JEV +(docs-derived seed, no evidence gate) and dinostomp +(instrument). `validation.md`; `notes.md` §64, §70. + +## Does a positive Jev score authorize the act? + +No. **Jev supplies evidence. Code owns authority.** +[actiongate-jev](https://github.com/omkarghugarkar007/actiongate-jev): +deterministic policy / RBAC / schemas / limits own +ALLOW | REVIEW | BLOCK. A positive model score never +overrides a deterministic security failure. Six narrow +questions, never one vague "is this safe?" Financial / +destructive / credential fail closed if Jev is down. +500-case eval is label-baseline integrity, not accuracy. +Compose with construct (privilege ≠ verdict) and interlock +(SENSOR ≠ policy). `applied-mappings.md` §7; `notes.md` §64. + +## Does "calibrated" mean I can threshold p as a frequency? + +No. Ranking ≠ calibration. +[does-jev-confidence-mean-anything](https://github.com/Adilmp/does-jev-confidence-mean-anything): +8,000 human-annotated judgments; AUC **~0.91** while when +Jev said **~75%**, humans flagged **~10%**. Two-parameter +recalibration removes **~96% of ECE** without changing +rank. Vendor "Calibrated: higher confidence means higher +accuracy" is **true as rank-correlation** (0.96 *theirs*) +and **false as probability units**. Companion +[jevcal](https://github.com/Adilmp/jevcal) (~100 labelled +rows). Never `if p > 0.9` without domain recalibration. +One dataset; do not cite `threat`. `mappings.md` §7; +`notes.md` §64. + +## Does AUC mean the probabilities are honest? + +No. Ranking can be strong while units are wrong — and the +**sign of the error can flip by question type**. +[jev-ood-calibration](https://github.com/scienthoon/jev-ood-calibration): +on 900 synthetic tickets Jev cannot have seen, Choice/Score +are overconfident (refit T ~3.3) while the boolean on the +**same** tickets is underconfident (T 0.66). The priority +label is an org rule **not in the text** (44.7% acc, mean +stated p 0.74). In-domain OpenBookQA looks almost honest +(ECE 0.024). Do not threshold the TypeSafe `confidence` +field. Complements does-jev-confidence. `notes.md` §66. + +## Can I gate on Jev `confidence` alone? + +Usually that is the **rosiest** reading of the vector. +[how-sure-is-jev](https://github.com/adarc8/how-sure-is-jev): +Choice `confidence` **is** rescaled max-prob +(`(p_max − 1/n) / (1 − 1/n)`), to 3 decimals on 60 live +answers *theirs*. A 2-option 75/25 is Jev **0.5** and +entropy **0.19**. Use margin / entropy / gini as named +features; bands CERTAIN|…|CLUELESS are **policy**. Pair +with OOD (do not threshold `confidence`) and +does-jev-confidence (ranking ≠ calibration). Zero-dep. +`notes.md` §67. + +## Are combinators a new judgment model? + +No. They are **control-plane primitives** over typed +judgments. [jev-combinators](https://github.com/voidning/jev-combinators) +is the **rename** of +[decision-combinators](https://github.com/voidning/decision-combinators) +(same repo): Then / Gate / Vote / Cascade / Weighted plus +extended Router / Loop / Retry / Fallback / Memory. +Digital-design slogan (transistors / logic gates / chip) +is a metaphor for *soft* classifiers; AND/OR aggregation of +parallel Nouls still lives in code (do not multiply). +Compose with [skillranker](https://github.com/Dicklesworthstone/skillranker) +(VOI over a skill library; abstention; hook fail-open). +Not chat turns. `composition-algebra.md`; `notes.md` §66, +§69. + +## Does TLA+ replace Jev, or the reverse? + +Neither. [jev-labs](https://github.com/copyleftdev/jev-labs) +puts TLA+ on the **protocol** (quorum, stability, crash) +and Jev on the **oracle**. The invariant is never +confidently wrong: escalate is allowed. 1,080 golden +rounds 0 wrong *theirs* bounds the violation rate below +0.28% (rule of three) — it does **not** prove zero. +Hard-gating a soft judgment without an escalation path +is soundness theater's inverse. Synthetic pharmacy, not +clinical. `formal-methods.md`; `notes.md` §67. + +## Does a typed answer unlock the next step? + +No. [seal](https://github.com/Reasonofmoon/seal): **Jev +answers questions; SEAL answers whether the world may +change.** Coverage.path ∈ {auto|code|human|escalate} +must be visible. Mint ≠ product brain. Generation fills +Candidates; only a Seal advances. `notes.md` §67. + +## Is JevBench a TypeSafe leaderboard? + +No. [jevbench](https://github.com/fstandhartinger/jevbench) +is Benchmark Heaven's unofficial v1.1 bake-off. +Calibration is **reported, not scored**. Native vs +verbalized are labelled. Partial runs are not ranked. +Main Score = 0.6 Capability + 0.2 Speed + 0.2 Cost +*theirs* (Jev 1.13.0 **87.6**). Contrast atlas +(receipts, not a ranking). Harbor/jevals practice, not +a vendor eval. `validation.md`; `notes.md` §67. + +## Is Jev weaker than a 4B model? + +Only with the **thinking budget** attached, on this set. +[jev-frontier-100](https://github.com/softpudding/jev-frontier-100): +Jev **77.0%**; Qwen3.5 4B with thinking off **56.0%**; at +512 **78.3%**; at 2048 **96.7%** (+12.7 to +26.7 pp). +Exploratory, not preregistered. Similar totals ≠ similar +skills. Not a ceiling. `notes.md` §66. + +## Does a local MLX one-pass replica give Nouls? + +No. [jevmlx](https://github.com/bnsd55/jevmlx) assembles +schema-valid JSON with a probability per field in one +forward pass on Apple Silicon. Softmax over allowed tokens +≠ a calibrated Noul. No local leaderboard yet. Distinct +from system-one-benchmark's Harbor Brier table +(PCD 0.3884 vs Jev 0.1096). `judgment-class.md`; `notes.md` +§61, §66. + +## Does Jev `done` mean the browser task succeeded? + +No. Code owns observe / execute / verify / exit. +[ego-jev](https://github.com/jiangkoumo/ego-jev): `--until` (URL +substring or a `check` function) is the deterministic success +condition; Jev self-`done` is weaker. Malformed fill JSON is +`text_model_failed`, not a guessed value. Same lesson as +`DONE` ≠ verified success (gliner2-ultrafast) and waymode +`completed` ≠ server-state success. `applied-mappings.md` §2; +`notes.md` §65. + +## Should the model's own hides become training labels? + +No. A feedback loop that auto-trains on the judge's own +negatives self-reinforces errors. +[x-reply-filter](https://github.com/zhuyansen/x-reply-filter): +local rules first; remainder Nouls; auto-collapses sit in a +confirm queue until a human says hide or keep. Distinct from +distilling Jev as teacher of record (jev-triage ~68% ceiling) +— here the poison is *self-labeled hides*. `applied-mappings.md` +§4; `notes.md` §65. + +## Do Ax / DSPy own the control plane? + +No. They climb **LM-program knobs** (prompts, demos, module +graphs). A typed control plane is deterministic code around that +program: ontology validation → security override → confidence → +state machine → tool allow-list. +[jev-dspy-control-plane](https://github.com/manikanda-kumar/jev-dspy-control-plane) +lets DSPy draft **after** route+action are fixed. Offline +heuristic + contract stubs prove plumbing, not quality. Accuracy +alone is not enough; a negative result is valuable. +`optimizer-integration.md`; `notes.md` §59. + +## Native probabilities or verbalized confidence? + +Measure the native distribution. Verbalized "I'm 80% sure" is a +different object (and usually needs temperature scaling). +[jev-arena](https://github.com/meetr1912/jev-arena) scores Jev's +`noul`/`choice`/`score` on analytically-known worlds. Their live +card (`jev-1.13.0`, 145 noul, 2 requests): Brier **0.0059**, ECE +**0.0620**, overconfident in the low bins. Fan-out is measurement +economics, not a demo flourish. Siblings: sonar (heatmap-as-policy), +vickrey (Jev never bids), bracket (Brier vs Elo; live trailed Elo — +honest). `validation.md`; `notes.md` §59. + +## Train a specialist, or few-shot the hosted API? + +Depends on **what consumes the output**, not on a vs-hosted +accuracy table. +[Domain-jev-maker](https://github.com/help-er/Domain-jev-maker) +trains a domain LoRA on **independent** CLINC-150 labels (not +a Jev teacher-copy). Their RESULTS.md: local 1.5B KL 0.168 vs +hosted zero-shot 0.580 banking (r +0.933 vs +0.343). Few-shot +hosted (one example per intent in `state`) matches or beats +local determinate accuracy (McNemar p=0.134 / p=1.000); +calibration barely moves. **Argmax routing → hosted + +examples. Threshold / deferral / expected-cost that reads +p → specialist.** Matched-precision KL (both systems rounded +to two decimals, zeros → 0.0025) is the Harbor hygiene. +`mappings.md` §2; `notes.md` §60. + +## Is Noul 0.5 "maybe / medium"? + +No. Noul 0.5 is **cannot-tell** — uncertainty about a +predicate, never medium intensity, **never rounded** into an +act. [jav-email-cascade](https://github.com/skiingfalcon/jav-email-cascade) +treats the band between mirrored thresholds as review, not +auto. Score confidence 0.0 is a flat distribution and is +never acted on. Same non-negotiable as the skill card. +`applied-mappings.md` §8; `notes.md` §60. + +## Does a passing ECE mean `ORDER BY` is safe? + +No. **Calibration ≠ sortable.** ECE/Brier ask whether a +stated 0.7 is 70%; pairwise inversion / Score ordinality ask +whether sorting by the number puts rows in a defensible +order. They come apart in both directions. +[jev-orderby-bench](https://github.com/yodablocks/jev-orderby-bench): +`jev-1.13.0` passes six pre-registered gates; Score ordinal +inversion 0.143 vs 0.15 is the weak link **and the sort +key**; 53 rows tie at 0.99 so `LIMIT 20` is +engine-dependent; two-decimal quantization. recodelabs +40-row batching fails the ranking gate that one-row-per- +request passes. Vendor 67.8% agreement is not calibration. +`mappings.md` §4; `notes.md` §60. + +## Should we distill Jev as the teacher of record? + +No. Use Jev to **decide what enters the training set**, not +as the label teacher. +[jev-triage](https://github.com/ThyFriendlyFox/jev-triage) +routes high-conf accept / middling expensive teacher / +low-or-boundary human and logs **full distributions** for a +local student. Author: a ~**68% ceiling compounds errors**. +Real outcome labels remain the training targets. Soft labels +are a bootstrap — cut the cord when the local head wins on +held-out real labels. Distinct from Domain-jev-maker +(independent gold specialist) and openjev-lm (teacher-copy). +`mappings.md` §2, §6; `notes.md` §61. + +## Is local PCD a calibrated Noul? + +No. **O(1) speed ≠ calibrated probability.** +[system-one-benchmark](https://github.com/mallahyari/system-one-benchmark) +on LMSYS toxic-chat n=50: local MLX PCD (Qwen2.5-1.5B) is 1 +forward pass / p50 227.2 ms / 52% acc / Brier **0.3884**; +Jev-1.13.0 is 84.0% / Brier **0.1096** / p50 356.5 ms +HTTPS. AR JSON ~30.8 passes and 98% schema errors. Softmax +over allowed tokens is not a Noul. Same honesty as jevify +(uncalibrated CUDA likelihoods). Productized cousin: +[jevmlx](https://github.com/bnsd55/jevmlx) (28★) — schema→JSON ++ per-field probs in one MLX pass; **no local leaderboard +yet**. Small n — *their* card. +`judgment-class.md`; `notes.md` §61, §66. + +## Closed-vote computer-use, or Stagehand pick? + +Closed-vote means **code builds every option, the decision +model only picks, no planner LLM**. +[JevOnly](https://github.com/buluoray/JevOnly) is that +harness (Apache-2.0; type without generation; verify/undo). +[waymode](https://github.com/mossburgh/waymode) is the +**product** cousin: the app keeps handlers and permissions; +Jev selects among live typed actions; `completed` is Jev's +reading. Stagehand pick is a **fast path with LLM fallback**, +not this card. `applied-mappings.md` §9; `notes.md` §61. + +## Engine eval or coaching verdict? + +The engine owns truth; Jev owns judgment. +[game-coach](https://github.com/JoelLewis/game-coach) (Wave 0 PRD; +GPL-3.0): Stockfish WASM eval/lines/swing; Jev severity / error +class / interrupt; templates + a capped writing model own words. +Jev never evaluates positions or picks moves. Same anti-soundness- +theater as PR attention ≠ correctness (egma-ai). Silence is a +feature. `formal-methods.md`; `notes.md` §59. + +## Is observe→score→act Jev-only? + +No. The hole is backend-agnostic: observe controls, score among those +candidates, code acts. [jev-ultrafast](https://github.com/browser-use/jev-ultrafast) +and [solari-reflex](https://github.com/hitakshiA/solari-reflex) use +TypeSafe Jev; [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) +uses local GLiNER2 (`fastino/gliner2-multi-v1`); +[laya-mind2web](https://huggingface.co/ShaunSpark/laya-mind2web-browser-agent) +uses a Laya head over DOM element indices; +[Cua-S1](https://github.com/trycua/cua/tree/main/libs/cua-s1) uses a +byte encoder + option-attention head (fill/check/click/skip) — **not +TypeSafe Jev**, source-only this pass. +[Stagehand #2951–#2955](https://github.com/browserbase/stagehand/pull/2955) +is the same hole **inside a major harness**: Jev picks; code copies +or acts; LLM fallback; extract `"off"`/`"judge"`/`"pick"`. Their +card: 37/75 no-LLM ~0.5 s vs baseline 4.37 s; LLM-off 36/75 — **pick +is a fast path, not a replacement.** Draft stack; do not copy the +opt-in flag. Same lesson as compaction +(Jev Noul/Score vs GLiNER2.5). Screenshot multimodal (blackwood-rlcd: +letters on an image) is a **different input**, not a better version of +this hole. Hybrid local decide + remote fill is mixed-architecture +economics, not dual-process-ai. `DONE` is loop termination, not +verified success. Plan ≠ execute; dry-run default on Cua-S1. Closed-vote +extreme: [JevOnly](https://github.com/buluoray/JevOnly) has **no +planner LLM** (code builds options, Jev only picks). +[waymode](https://github.com/mossburgh/waymode) is host-owned +handlers × System One, not a harness. Not +GLiNER2.5. Not a bake-off against the Flights demo clock. +[ego-jev](https://github.com/jiangkoumo/ego-jev) is the same +hole on ego-lite: indexed viewport table → operation+target; +code owns the loop; text model only for type; `--until` beats +Jev `done`. n=3 medians ~2×, not a bench. +[awlevin/typesafe-computer-use](https://github.com/awlevin/typesafe-computer-use) +is the **macOS OCR+AX** product of the same hole (MIT; +**427★**): never ships a screenshot for the *decision*; +writer only for free text; post-type Noul 0.5 *theirs* still +soft; overlapping options = false low confidence; split +kind/item/site. The one-shot answer reader may receive the +capture — that is not the Choice. $0.0002 / 155× is +one-screenshot *theirs*, not a Harbor taskset. **≠** +jev-ultrafast **≠** cua-s1 **≠** jev-macos-loop **≠** +open-typesafe-camoufox. `notes.md` §81. +[moritzkremb/jev-voice-browser](https://github.com/moritzkremb/jev-voice-browser) +is the **ASR** product of the same hole (MIT; +**103★**): partial transcript → 9–11 questions → +Playwright. Pointer spans. Spoken confirm ≠ auth. +**≠** jev-voice-control **≠** nikolas-j. `notes.md` §82. +`judgment-class.md`; `mixed-architecture.md`; `notes.md` §52, §54, §57, §61, §65, §81, §82. + +## Should I send the screenshot to Jev for computer use? Is 155× a Harbor score? + +No, and no. [typesafe-computer-use](https://github.com/awlevin/typesafe-computer-use) +reads the screen **deterministically** (Vision OCR + AX + +dates.py) and asks TypeSafe a Choice over numbered items. +The decision never ships pixels to a frontier model. The +one-shot **answer** writer may receive the capture — a +reader packet, not the classifier. 155× / $0.0002 is +*theirs* on **one screenshot**, with the honest caveat that +frontier read dates unaided. Re-measure on *your* taskset. +`--min-confidence` 0.4 and post-type Noul 0.5 stay product +copy, not Harbor τ. Overlapping actions read as doubt; keep +the set exclusive. `notes.md` §81. + +## Is jev-voice-browser omni System One? Is 27/27 a Harbor score? + +No, and no. [jev-voice-browser](https://github.com/moritzkremb/jev-voice-browser) +is **ASR → text-state → hosted Jev → Playwright**. Web +Speech (audio to Google) is the producer; Jev never +hears the waveform. Partial transcripts get one +9–11-question request; policy waits longer on +free-text than on closed-set. Pointer-not-generator +for spans. Spoken "confirm" is a convenience, not +auth. Numbered overlays disambiguate without another +model. 27/27 / ~$0.0002/call / ~300 ms are *theirs* +on fixtures, not a Harbor taskset. 0.5 / 0.55 / 0.6 +still soft. **≠** chris-wozniczek/jev-voice-control +**≠** nikolas-j/jev-voice-browser **≠** +typesafe-computer-use (OCR). Compose, don't collapse. +Skip Archer. `notes.md` §39, §82. + +## Is AgentGhost an advisory sidecar? Can ASK be skipped? + +No, and no. [reddpy/AgentGhost](https://github.com/reddpy/AgentGhost) +**is the tool's execution function.** README *theirs*: +the model never decides whether the wrap runs. Rules +(`allow`/`ask`/`deny`, `matchArg`) fire before Jev; +`allow` skips the judge. ASK/DENY throw so HITL +cannot be silently skipped. `failMode: "closed"` +denies on judge error. Judge is a slot (Gateway / +TypeSafe / custom). `AGENTGHOST_AUTO_APPROVE=1` is +a demo hatch, not a grant. Provider-hosted tools +and MCP (planned) are out of reach. **≠** +jwen5419807/agentghost **≠** vventirozos/AgentGhost +**≠** actiongate-jev **≠** toolgate **≠** jev-use +(fail-open). rh-guard owns the gate cousin. Do not +copy `npm` / `.env`. `notes.md` §83. + +## Is the @studio_yebisu JP roundup a bake-off? Can I quote its star counts? + +No, and no. [@studio_yebisu](https://x.com/studio_yebisu/status/2101065176069886152) +is a **genre atlas** of high-star Jev *apps* plus +open replicas. Tweet *theirs*: stars are +research-time; the post is a docs roundup, not +full eval. Engagement ephemeral (SIGNAL ~120k +views; this pass 131,234 / 1,934 / 192). Star +drift is the point (typesafe-computer-use ★203→ +**427**; jev-voice-browser ★40→**103**). **≠** +@airesearch12 class census §77 **≠** JevBench +v1.2 §78. SAM 3.1 already §39. OpenRouter Jev +"no waitlist" is WATCH, not a recipe. Do not dump +the 30 repos. Skip Archer. `notes.md` §84. + +## Is Akshay’s “Jev Clearly Explained” a TypeSafe how-to? Can I quote 200× / 400×? + +No, and no. [@akshay_pachaar](https://x.com/akshay_pachaar/status/2101037514945597645) +is **independent pedagogy**. Article *theirs*: +LLM hammer for bounded decisions; code owns the +branches; parallel questions; thresholds in +code; schema-safe ≠ correct (“cannot break the +declared output schema, but it can still be +wrong”); placements = routing / tool-risk / +verify with LLM; shadow-mode; questions-as-code. +**200× / 400×** and 70–500 ms / $0.042/MTok sit +at TypeSafe’s favorable end — treat as a +**ceiling, not a promise**. Not Harbor. Text-only; +not looking at the screen. Do not copy the +Python samples. **≠** official docs **≠** Flavio +Copes **≠** LangChain harness **≠** AgentGhost +wrap. Engagement ephemeral (SIGNAL ~183k; this +pass 233,495 / 2,280 / 235). Skip Archer. +`notes.md` §85. + +## Does Stagehand extract replace the LLM? + +No. Pick-and-copy is a **fast path**. Schema leftovers, screenshot +extract, failed gates, and abstention still call the LLM. 36/75 with +the LLM disabled is the honesty number. `applied-mappings.md` §2; +`notes.md` §57. + +## Is a public yes/no wall the product? + +It is a **primitive surface**, not a chatbot. [ask-jev-ai](https://github.com/waynesutton/ask-jev-ai): +one call, six questions, policy in `convex/questions.ts`, code +decides live/blocked. Cost-to-1M from TypeSafe token counts +($32–$41), not estimates. No-key: allowlist, UI says Jev offline. +License null this pass. Do not copy Convex how-to. +`mixed-architecture.md`; `notes.md` §58. + +## Meaning-search or grep? + +Use grep when you know the string. [jevgrep](https://github.com/Bentlybro/jevgrep) +is for "where is the code that *does* X" with no embeddings: packed +parallel Jev relevance; 79% top-5 vs BM25 40% / grep 20% on +docstring-stripped repos; BM25 still wins exact wording (top-10 +96% vs 85%). Line-level AND/OR/NOT over Nouls, including +JP↔EN and FR/RU/DE/ES/ZH/KO: [jev-semgrep](https://github.com/uehaj/jev-semgrep) +(name collides with [Semgrep.dev](https://semgrep.dev) SAST; +proposition ≠ embedding; AND/OR/NOT after threshold; +not a gate; dedicated `notes.md` §86). Citable +"where is this *enforced*?" packets, index-once: +[jev-semantic-explorer](https://github.com/jimmyhealer/jev-semantic-explorer) +(jevex; 1/8→6/8 n=8 *theirs*; packet HitFile 0.233 is not +the product number). Distinct from kazuhideoki file+fzf, +superagents-lab web, and jev-sift classify-first. +`mappings.md` §4; `notes.md` §58, §61, §86. + +## Attention or correctness on a PR? + +Attention. [egma-ai/jev-reviewer](https://github.com/egma-ai/jev-reviewer) +assigns P0/P1/P2 for where to look; OpenAI writes behavior deltas. +Incomplete never becomes P2. That is **not** +[choxos/jev-reviewer](https://github.com/choxos/jev-reviewer) +(pointer-not-generator). A Noul is not a proof the PR is good. +`formal-methods.md`; `notes.md` §48, §58. + +## Can oxlint hard-gate on a Noul? + +No. [jev-oxlint](https://github.com/cephalization/jev-oxlint) AST / +precheck prove what they can; guidance lives whole-file in state; +Jev scores the remainder. Phoenix fixtures matched the human +answer key; routing was sharp; the coarse hint was not. Experiment; +`tenbin` owns the lint skill. Do not treat remainder Noul as a +discharged proof. `mappings.md` §18; `notes.md` §58. + +## Session-sticky routing: fail-open or fail-closed? + +Name the irreversible act: *sending a model*. +[jev-adaptive-thinking](https://github.com/jxu-dev-c/jev-adaptive-thinking) +locks a declared standard (`gpt-5.6-sol`) on timeout / no session. +That is fail-closed to fallback, not jev-gateway passthrough. +First prompt classifies; later requests never reclassify. Same +family as routeKit (Jev estimates; code picks). `applied-mappings.md` +§5; `notes.md` §58. + +## Did Jev-RAG beat full-context Spark on latency? + +No. [Jev-RAG](https://github.com/Max-sm-yc/Jev-RAG) one-run: ≥70% +cost and 72% latency **vs Muse Spark rerank**. Full-context Spark +is still **faster** (10.60 s vs 62.3 s). Costs include embeddings. +Do not overclaim vs no-RAG. `mappings.md` §4; `notes.md` §58. + +## Is Cua-S1 TypeSafe Jev? + +No. [Cua-S1](https://github.com/trycua/cua/tree/main/libs/cua-s1) +uses "System One" as a computer-use research label for a small +specialist decide head. It does not ship Choice/Score/Noul, a +`/v1/systemone` drop-in, or TypeSafe contracts. Augustus places the +*hole* (observe candidates → select among a closed option set → code +acts under a hard envelope), not the logo. Same lesson as GLiNER2 +Ultrafast vs Jev Ultrafast. Source-only; no checkpoint scores. +`judgment-class.md`; `notes.md` §54. + +## Is routing the same as memory? + +No. A cheap intent gate can skip a memory/tool *tour* on +`calendar` / `mail` / `status` without turning memory off. +[jev-hermes](https://github.com/de-niji/jev-hermes): memory still +**writes**; `complex` still searches. Route ≠ memory. +`applied-mappings.md` §5; `notes.md` §48. ## Is confidence a trained score? @@ -282,5 +1398,648 @@ the adapter revision he cites, is then ordinary arithmetic: how far the leading probability sits above a uniform `1/K`. It is not a second learned estimate that the answer is correct. A peaked distribution can be confidently wrong. Threshold a p you have checked on your labels. +Independent receipt: [how-sure-is-jev](https://github.com/adarc8/how-sure-is-jev) +found Choice `confidence == max_prob` to 3 decimals on 60 +answers — the most generous metric in their table. [Essay](https://archerhume.com/posts/jevs-architecture-unmasked/), -`notes.md` §31, `mental-models.md` calibration. +`notes.md` §31, §67, `mental-models.md` calibration. + +## Can CI call Jev live on every PR? + +Prefer not. [jevassert](https://github.com/dtduc-git/jevassert) +**LANDED**: `record` once, commit `predictions.jsonl`, +`check` **offline** (accuracy + ECE/Brier + cost/latency +gates; exit 0/1/2; McNemar compare). Calibration and cost +are first-class, not footnotes. Pairs with jev-packs SPEC +v0. Do not copy `uvx`. `notes.md` §70. + +## Is jevarena the same as jev-arena? + +No. **Always qualify the owner.** +[chenmingtang830/jevarena](https://github.com/chenmingtang830/jevarena) +is an open BYOK failure-finding playground (JevJudge-Bench +harness; **not** measured model findings). +[meetr1912/jev-arena](https://github.com/meetr1912/jev-arena) +is a native-probability calibration arena on analytic +worlds (Brier/ECE; §59). Neither is a winner-crowning +leaderboard. `notes.md` §59, §70. + +## Does 97% on BBQ mean Jev is unbiased? + +No. [jev-bbq-experiment](https://github.com/simonmesmith/jev-bbq-experiment) +*theirs*: 58,492 questions, **97.28%**, amb bias **0.04** / +inf **0.34**, **$0.3429**. 12 of 13 ambiguous errors were +stereotype-aligned. One frozen English/U.S. QA template is +not a hiring/lending/healthcare cert. Pair with the +lustig framing stub, not a substitute. `notes.md` §70. + +## Does the LLM plan in jeffrey? + +No. **Decider ≠ executor.** +[jeffrey](https://github.com/thomasbrueggemann/jeffrey): +Jev owns next-tool / progress / risk / done; the LLM +**only fills args**. Risk ≥ 0.5 pauses mutating tools. +Stuck ladder withholds the looping tool and re-asks Jev. +Distinct from jev-handoff (baton around an existing host). +`notes.md` §70. + +## Is jevlint the same as JevLint? + +No. **Always qualify the owner.** +[mizchi/jevlint](https://github.com/mizchi/jevlint): +ast-grep subjects × sentence `ask:` (matcher silent, Jev +loud; 13/15 1.00/1.00 *theirs*). +[huntedman/JevLint](https://github.com/huntedman/JevLint) +is file-level convention Nouls (§26). Independent of +eslint-plugin-jev. `notes.md` §26, §70. + +## Can I treat a local `/v1/systemone` as Jev? + +Only as a **wire**. grande (Rust/WebGPU), laya-jolt +(Clojure byte-parity Laya), JEV-CPU (SemIf on CPU; +Meanblock 404), kunchenguid/local-jev (ONNX ModernBERT; +done **30%** / shape **57%** vs Jev *theirs*), +**githubnext/localjev** (prompted JSON + entropy +confidence; wire-compat ≠ logit-equiv), OpenJev +`/v1/decide`, semif-serve runoff, jeff GLiFormer, and the +GLiNER2 spec are **class substrates**. Softmax / generated +JSON ≠ Noul until calibrated on your labels. Interface +compatibility ≠ replica. Always qualify +**githubnext/localjev** vs **kunchenguid/local-jev**. The +Eran-BA GLiNER2 document is **spec-only** (no service, no +measurements) and is not jeff. Archer still Watch. +`judgment-class.md`; `notes.md` §70, §75. + +## Do user constraints survive compaction? + +Not in the model's memory. [pi-heed](https://github.com/Nyarlathoteppppp/pi-heed) +persists them as structured state and checks +side-effecting calls before they run. Jev **never writes +policy**. Fail-open. Distinct from actiongate (RBAC/schema +authority). `notes.md` §70. + +## Is jev-judge-bench the same as jevarena or jevbench? + +No. **Always qualify the owner.** +[slavadubrov/jev-judge-bench](https://github.com/slavadubrov/jev-judge-bench) +is a frozen SLA-150 Harbor-shaped **contract**: Jev vs +cheap schema-guided LLM judges, human labels, invalid = FN, +cost/latency. **No quality headline yet** (21 offline tests; +canaries are availability). +[chenmingtang830/jevarena](https://github.com/chenmingtang830/jevarena) +is a failure-finding playground (JevJudge-Bench harness; +§70). [fstandhartinger/jevbench](https://github.com/fstandhartinger/jevbench) +is Capability/Speed/Cost Main Score (calibration reported, +not scored; §67). `notes.md` §71. + +## Is jev-use the same as jev-ultrafast? Does Vercel Jev return confidence? + +No, and often no. [jev-use](https://github.com/shitianfang/jev-use) +hands **no-text** steps (did it work / which next / safe) to +Jev and leaves writing with the LLM. Distinct from +browser-use/jev-ultrafast (observe→score-act browser loop). +Through the Vercel gateway there is **no confidence field** +— jev-use reconstructs margin and defaults that backend to +**0.4**. First loop 17/20 escalate, then 0/20. Same author +as jev-handoff. Gate is fail-open. `notes.md` §71. + +## What is pi-jev-control? + +A System-One **control plane** for Pi (router, tool gate, +failure+retry, context/skill/memory, compaction epoch, +review, GUI), not a second agent. +[pi-jev-control](https://github.com/goodruizhan/pi-jev-control) +compaction never modifies the on-disk session; GUI +confidence below threshold → `unknown`, never force-click. +License null; no live quality numbers in the README. +Distinct from omp-jev-extensions / jevons / pi-heed / pi-om. +`notes.md` §71. + +## Can Jev generate text (jev-gpt)? + +It never free-generates. +[jev-gpt](https://github.com/florian-hoenicke/jev-gpt) +asks one typed question per choice over a WordNet / +jina-embeddings tree, then ranks candidate texts. README +*theirs*: ~400 calls, 75 s, 2 cents per prompt. +Architecture demo of **decider ≠ executor** taken to the +word. Distinct from jeffrey (pick next-tool, LLM fills +args). License null. `notes.md` §71. + +## Are jev-cookbook numbers a benchmark? + +No. [jev-cookbook](https://github.com/nexibeo/jev-cookbook) +samples are 16–36 handmade items; the authors say the +scores show technique, not benches. Recipes 01–13: 425 +calls / $0.015; browser 5/6 *theirs*. Pattern: code +prepares, Jev answers narrow questions. MIT. `notes.md` §71. + +## Is jevfeed a social product? + +No. [jevfeed](https://github.com/fengyiqicoder/jevfeed) +ranks links found in **your** last 200 history pages. No +likes, follows, or accounts. History stays local; one Jev +request per batch of ten (the distribution *is* ranking). +Distinct from ThinkyMiner/Winnow (grade an existing feed) +and kevinpita/winnow (context sieve). `notes.md` §71. + +## Did openJev-verdict-2.0 beat Jev? Is it OpenJev? + +Treat it as a **claim-audit**, not an endorsement, and no +it is not IamBusy/OpenJev. +[openJev-verdict-2.0](https://github.com/Heman10x-NGU/openJev-verdict-2.0) +README *theirs*: 77.10% / Brier 0.0636 / ECE 0.0144 on +LocalLLaMA/typed-decisions. The Jev table row is a +Laya-catalogued vendor baseline, not an independent run. +**Open PR #1** already flags: throughput misread as +latency; Laya gap inside the 95% CI (parity, not SOTA); +correctness-head ECE is not distribution ECE (Jev 14.40% +slightly lower like-for-like). Distinct from +IamBusy/OpenJev `/v1/decide`. `notes.md` §71. + +## Is IPECTER/jev-context-pruner a compaction product? + +Not this pass. The repo is **empty** (409). Slogan only. +Sibling of fast-jev-compaction / jev-compactor / +dizk/jev-lens. Do not invent files. `notes.md` §71. + +## Is chakuho a local Jev? Is coverage a Noul? + +No, and no. [chakuho](https://github.com/taku-me/chakuho) +is a 1-token logprob endpoint over a generic instruct +model (`POST /v1/systemone`). Softmax over declared +labels is **not** a calibrated Noul. `coverage` is the +mass that landed on those labels — format-keeping, not +correctness. 8B stays coverage 1.00 while `__none__` +collapses (3/30 vs 27B 29/30 *theirs*). GUI 336-case: +27B 95%/92% vs Jev Gateway 89%/82%. Arithmetic stays in +code. Cousin jevify / TypeAR / pcdServer / jevmlx. +`notes.md` §72. + +## Is jevinf TypeSafe Jev? + +No. [jevinf](https://github.com/zerodegress/jevinf) is an +open replica **runtime**: segmented forwards + prefix +reuse, then the Jev wire. NanoJev / decider-2b / Laya +families; only `torch-mps` is wired. 2.57× / 2.27× at +**100% argmax agreement** *theirs* is speed, not ECE. +Python ≥3.14. `notes.md` §72. + +## Is typesafe-elixir-sdk the same as dannote/jev? + +No. [typesafe-elixir-sdk](https://github.com/phiat/typesafe-elixir-sdk) +is an unofficial **HTTP client** (Req; same env/retry as +the Python SDK). [dannote/jev](https://github.com/dannote/jev) +is an OTP **peer GenServer** (answers as messages; clause +order is routing). Code still owns the `cond`. Confidence +is concentration, not P(correct). `notes.md` §72. + +## Did jevex replace jev-semantic-explorer? New numbers? + +It **is** the rename (same `created_at`; GitHub +redirects). New card *theirs*: SWE-bench Verified +**n=16**, 160s → 69s, $8.74 → $3.13, 16/16 both arms. +90s cap: 1/16 vs 11/16 finished. Keep the older n=8 +finish 1/8 → 6/8. Claude Code n=5 6.8 → 2.2 is +unchanged. Packet HitFile stays diagnostic. +`notes.md` §61, §72. + +## Does commitjev block a commit? Can I round 0.4? + +It blocks only on a **warning**. A failed check is not +a reason to refuse. Anything between 0.35 and 0.65 +(their 0.65 bands) is reported as **"review"**, never +rounded. Nouls decide the verdict; the Choice is the +headline. Six regex checks never reach the model. +Same owner as jev-orderby-bench (one commit per call +because batch ranking fails). Five clean commits is a +small control. `notes.md` §72. + +## Is hermes-plugin-jev TypeSafe System One? + +No. The README brands Choice/Noul/Score as Jev, but +`__init__.py` calls **Agnes 3.0 Flash** via +`system_one_adapter` OpenAI chat-completions +(`AGNES_API_KEY`). Distinct from hermes-jev-router +(TypeSafe WHETHER/HOW/WHAT). `notes.md` §72. + +## Is pi-jev-compact the same as pi-jev-compaction? + +No. [pi-jev-compact](https://github.com/dev-willbird1936/pi-jev-compact) +replaces Pi's **summarizer** with verbatim keep/drop of +paired tool calls (fail-open to the LLM summary if +<25% saved). [pi-jev-compaction](https://github.com/vava-nessa/pi-jev-compaction) +is the fast-jev-compaction cousin. Pair compact with +pi-jev-control / pi-heed, not as a clone of either. +`notes.md` §72. + +## Is IPECTER/jev-runway a Codex proxy? + +Not this pass. LICENSE only; created and pushed one +second apart; README 404. Second IPECTER empty slogan +(after jev-context-pruner). `notes.md` §72. + +## What is mailordinal? Is it just email classification? + +A **decision-native inbox**: nine typed questions, then +a 100-point policy in code (SLA + account tier). It +does not ask "how urgent is this?" Humans own +ambiguity — low confidence never silently lowers +priority. Life/business, not SWE-only. Cousin of +jav-email-cascade. Independent of TypeSafe. +`notes.md` §72. + +## Is jev-cli the same as jevql? Can I install it? + +No, and not yet. [jev-cli](https://github.com/shaharia-lab/jev-cli) +is an unofficial Rust CLI (exit codes as a semantic +`if`; planned MCP). README: **commands not +implemented**; crate: **not ready**; version 0.0.0. +[jevql](https://github.com/kylemclaren/jevql) is +judgment *outside* the store (vanilla Postgres never +sees `jev()`). Do not copy `claude mcp add jev`. +`notes.md` §72. + +## Should I use English Laya on non-English text? + +No. [laya-multilingual](https://huggingface.co/convaiinnovations/laya-multilingual) +exists because the English checkpoint **stays +confident while going to chance** (Khmer 0.000 acc at +0.952 confidence *theirs*; mean conf never < 0.885). +Route by **script before the forward pass**. The +multilingual checkpoint ships uncalibrated (ECE 0.314 +→ 0.106 after T *theirs*). Weaker on English — route, +do not replace. `notes.md` §72. + +## Is the Hub schema-scorer a Jev replica? GitHub 404? + +Hub +[jev-schema-scorer-deberta-v3-large](https://huggingface.co/mobarmg/jev-schema-scorer-deberta-v3-large) +is MIT; the GitHub repo 404s this pass. One DeBERTa +scalar head; **code** groups logits into +Choice/Noul/Score. v2 Choice acc 0.841 *theirs*. +Peaked probabilities are **rankings**, not calibrated +confidences. Distinct from +com-kotobalabs/open-jev-deberta-v3-large. `notes.md` +§72. + +## Why are open-jev-laya-bench / jev-tree-choice-cap / INSTRUCT_JEV 401? + +Access change this pass (historically 200). Do not +invent new numbers; do not re-fold. jevlogs-log-triage- +benchmark is GitHub 404 **and** HF 401 — skip. +`notes.md` §72. + +## Is classifier.dev just another public Jev wall? + +No. [classifier-dev](https://github.com/mrmps/classifier-dev) +(MIT; **185★**; https://classifier.dev) is a **zero-shot +classification API**: caller-supplied labels, batch +`{id, text}[]` up to ~1000, multi-label, dimensions. +Primary backend is TypeSafe Jev (`src/jev.ts`); LLM +chains are fallback only. Distinct from +[ask-jev-ai](https://github.com/waynesutton/ask-jev-ai) +(six parallel questions on one sentence). Life/business +(spam / inbox / feedback), not an agent hook. Do not +copy wrangler / `npm i -g`. `notes.md` §73. + +## Should every low-confidence answer escalate? + +No. `tier: smart` re-asks **single-label** answers +below 0.7. Multi-label **ignores** the tier: re-judging +made it worse (23 s *theirs*). 0.7 is *their* operating +point, not a class constant. Cousin of jev-use +escalate-under-threshold (different product; Vercel +drops `confidence`). `notes.md` §73. + +## Can I quote classifier.dev F1 0.887 / 87.7% AG News? + +Only with `eval/README`. `/benchmark` is generated from +tracked `src/vs-jev.json`, not hand transcription. +Multi-label n=7 is **train-on-test**, one annotator, no +CIs; ~0.03 is a coin flip. README *theirs*: F1 **0.887** +in **230 ms** vs cascade **0.799** / 1.5 s; eval *theirs* +232 ms, AG News **87.7%** vs ling-3.0-flash **82.0%**, +emotion **60.5%** vs **57.0%**. Calibration anecdote: +≥0.9 → 82% correct; <0.5 → 29%. `notes.md` §73. + +## What is the silent-fallback anti-pattern? Does rh-guard own it? + +A delisted primary left `granite-4.0-h-micro` serving +F1 **0.546** while docs advertised ~**0.800** for weeks +(*theirs*; eval/README). Nothing in the deployed numbers +said so. Digest now marks `FALLBACK`. That is **eval +integrity / soundness theater** (lying about the +instrument). **rh-guard owns the gate card**; dinostomp +owns instrument-not-score. classifier.dev is the lived +product cousin, not a new hook. `notes.md` §73. + +## Is choxos/jev-reviewer the same as egma-ai/jev-reviewer? + +No. Always write **[choxos/jev-reviewer](https://github.com/choxos/jev-reviewer)** +or **systematic-review Jev Reviewer**. It is local-first evidence +extraction: Jev points at line ids, code copies verbatim quotes, +human checks every answer (https://jevreviewer.xera.ac; MIT; +**12★**). [egma-ai/jev-reviewer](https://github.com/egma-ai/jev-reviewer) +assigns PR **attention** P0/P1/P2; OpenAI writes behavior deltas. +Attention ≠ correctness. `notes.md` §48, §58, §74. + +## Why two passes (Choice then Noul) on a paper? + +Relative Choice answers "which line, if any?" Absolute Noul answers +"does this line **itself** answer q?" Multi-row table answers +(Mean (SD) vs Median (IQR) under Age) need both. Quotes = Noul +≥ 0.5 *theirs*, not a class constant. Same species as Stagehand +extract / jev-sift; different domain (Cochrane / PRISMA). `notes.md` +§74. + +## Can I skip the human tick if Noul ≥ 0.5? + +No. The tick **is** the product. Checked or annotated answers are +never overwritten by a reworded question. Jev is SENSOR; the +reviewer is the constraint. Spot checks on the sample study are +**not** a validation study. `notes.md` §74. + +## Is *Not found* a failure of the extractor? + +No. When no line passes, the card says *Not found* (or *Unclear*, +with the closest lines) instead of guessing. Paraphrase invent is +the rejected species. *Not applicable* is a checked n/a. `notes.md` +§74. + +## Is githubnext/localjev the same as kunchenguid/local-jev? + +No. Always write **[githubnext/localjev](https://github.com/githubnext/localjev)** +(MIT; **261★**; GitHub Next). It is a Bun `POST /v1/systemone` +bridge: DiffusionGemma through ordinary Chat Completions; TypeSafe +SDK drop-in. [kunchenguid/local-jev](https://github.com/kunchenguid/local-jev) +is an ONNX ModernBERT approximation (done **30%** / shape **57%** +vs Jev *theirs*; confidence omitted). Hyphen vs no hyphen is +load-bearing. `notes.md` §70, §75. + +## Is LocalJev OpenJev? Are the probabilities logits? + +No, and no. [razorback16/openjev](https://github.com/razorback16/openjev) +obtains probabilities with a one-step DiffusionGemma **structured +read** (unmerged vLLM `diffusion_seed_canvas` / `diffusion_read_only` +/ token logprobs). githubnext/localjev **prompts** for a JSON +probability vector, validates/retries, then normalizes + entropy +confidence. Wire-compatible, **not mathematically equivalent**. +Distinct from [IamBusy/OpenJev](https://github.com/IamBusy/OpenJev) +`/v1/decide`. Cousin djev-spark is the structured-read Spark +surface already §36. `notes.md` §75. + +## Can I threshold LocalJev confidence as a Noul? + +Not until you calibrate on *your* labels. README: evaluate before +consequential decisions. The 1,200-request bake-off tests the +prompted-JSON pipeline, **not** direct logits, and says **do not +treat these outputs as calibrated probabilities** (high conf on +wrong BoolQ → large NLL; 40 samples/task). JSON-valid ≠ +picked-right. Entropy-as-confidence is a spread statistic on a +generated vector. `notes.md` §75. + +## Did Gemma 4 26B-A4B win? Should I use LM Studio instead? + +Neither as a class claim. Short-input *theirs*: Qwen3.6 macro +**76.7%**, Gemma 4 26B-A4B **75.0%** (lowest SST-5 MAE **0.533**), +DiffusionGemma **74.2%**. Qwen vs Gemma 26B is two answers of 120 — +not a statistically clear winner. Serving default **unchanged**. +LM Studio still cannot load DiffusionGemma (18 Sep 2026 *theirs*). +Even after Chat Completions support, swapping runners is not +OpenJev parity — the runner must expose structured-read primitives. +Do not copy bun / `.env`. `notes.md` §75. + +## Is NandhaKishorM/laya a new System One species? + +No. [NandhaKishorM/laya](https://github.com/NandhaKishorM/laya) +(Apache-2.0; **710★**) is the GitHub/PyPI packaging of +Hub Laya we already watch +([convaiinnovations/laya](https://huggingface.co/convaiinnovations/laya), +[laya-multilingual](https://huggingface.co/convaiinnovations/laya-multilingual), +[laya-typed-decisions](https://huggingface.co/convaiinnovations/laya-typed-decisions)). +Choice / Score / Noul, RLCD, NAR encoder. The new face +is `Router` (script-before-p) plus an honest vs-Jev +table. Always qualify GitHub vs Hub. `notes.md` §18, +§42, §46, §72, §76. + +## Is Laya a TypeSafe `/v1/systemone` drop-in? Same as localjev? + +No, and no. Different Python API (`agent.predict` / +`Router.predict`), not the TypeSafe SDK wire. +[githubnext/localjev](https://github.com/githubnext/localjev) +is a prompted-JSON Bun bridge. Do not copy `pip +install laya`. `notes.md` §76. + +## Can I hard-act at Laya confidence 0.85? + +Not from the README snippet. 0.85 is *their* +illustration because RLCD trains against strictly +proper scoring — it is **not** a Harbor-calibrated +threshold and **not** a class constant. Khmer 0.000 +acc at 95.2% confidence already shows gating cannot +catch script OOD; route **before** p. Fit T and pick +τ on *your* labels. `notes.md` §72, §76. + +## Did Laya beat Jev? + +Not as a class slogan. README: Jev figures are +**third-party published, never measured here**; +Banking77 is 72 vs 77 labels. Where Jev leads *on +that table*: Banking77 **0.870 vs 0.425**, soft-acc +**0.580 vs 0.471**, raw ECE **0.144 vs 0.213**. Where +Laya leads *theirs*: T4 **32.8 ms** vs Jev p50 +236–276 ms; post-T ECE **0.081 vs 0.246**; multilingual +router. typed-decisions **0.766** is a fine-tune on +that split (base ckpts below majority 0.461). `notes.md` +§76. + +## Is the @airesearch12 tweet JevBench v1.1? Can I quote the live board from it? + +No, and no. The tweet is a **named census + a promise** +of a first leaderboard "today" +([status/2101259522933186879](https://x.com/airesearch12/status/2101259522933186879)). +[`fstandhartinger/jevbench`](https://github.com/fstandhartinger/jevbench) +**v1.1** is already §67 (314 decisions; Main 0.6/0.2/0.2; +calibration **reported, not scored**). The watch URL is +[benchmarkheaven.com/jev-models](https://benchmarkheaven.com/jev-models). +Do **not** paste live ranks into the census card — scored +methodology is now §78 (JevBench v1.2 board). Also **≠** jev-judge-bench +**≠** jevarena **≠** jev-arena. `notes.md` §67, §77, §78. + +## Does GLiNER2 count as an openjev? Do routers? + +As a **class-boundary claim**, not as identity. +GLiNER2 ([fastino-ai/gliner2](https://github.com/fastino-ai/gliner2)) +locates/categorizes; +Succinct Router 14M / jev-model-router / Director / Loki +**route**. The tweet lists them beside NAR replicas and +TypeSafe-shaped wires. Same *job family* (fast cheap +bounded judgment) ≠ a Jev replica and ≠ a Noul ECE row. +Needle 3 is already **not Jev-class** in §67 +(function-calling / label only). Qualify `open-jev` +Dasein vs JoshuaSP; OpenJev razorback16 vs IamBusy +`/v1/decide`. `notes.md` §77. + +## Is the census complete? Are the missing names out of class? + +No, and no. Gaps vs our watch include Laya / +[NandhaKishorM/laya](https://github.com/NandhaKishorM/laya), +[githubnext/localjev](https://github.com/githubnext/localjev), +kev, TypeAR, openvons, chakuho, jevinf, grande, +laya-jolt, blackwood, classifier-dev. Incomplete census +is **lag**, not a dunk and not evidence those heads left +the class. Completeness is a board watch item. `notes.md` +§77. + +## Can I rank openjevs from the tweet? Are 15 likes a quality signal? + +No, and no. A list is not a bake-off. Likes/views/quotes +are **ephemeral** (SIGNAL ~417/9/3; this pass 564/15/5 — +do not quote as quality). The scored sibling is §78 +(geometric-mean I/C/S/K; Jev 75.3 / SemIf 74.6 *theirs*). +Harbor/jevals still wants a frozen taskset, calibration +on or honestly off the composite, named cost/latency +assumptions, partial runs not ranked, silent fallback +marked. `notes.md` §77, §78. + +## Did Jev win JevBench v1.2? Is 75.3 vs 87.6 a drop? + +Jev 1.13.0 is **#1 at 75.3** *theirs* (I 90.4 C 82.7 S 83.3 +K 51.7; $0.041 production). SemIf is **#2 at 74.6** (−0.7). +**75.3 is not a drop from v1.1 87.6** — different tiers +(hard 220 new; 534 vs 314) and scoring (geo-mean I/C/S/K +with calibration **on** the rank vs weighted sum with +calibration **off**). Do not mix versions. Unofficial, +not TypeSafe-endorsed. `notes.md` §67, §78. + +## Is Luna better than Jev? Can I hard-rank from the geometric mean? + +Luna has the highest Intelligence (**96.8**) and ranks +**#7** because Cost is **28.2** ($0.247/1k). DeepSeek +wins Calibration (**96.7**) and ranks **#11**. The +geometric mean does not let a strong axis buy back a +weak one — *theirs*. That formula is a **product +design**. Other views: Balanced no-cal SemIf #1 75.2 / +Jev #3 73.0; Emphasis Cost system-one-open #1 69.4 / +Jev #5 63.6. Limits: "The weights are a choice." +`notes.md` §78. + +## Is Qwen3.8 27B Archer? Why is Laya missing? + +No, and not as a quality verdict. Qwen3.8 27B is a +**partial** official-Qwen run via Chutes TEE (I 74.6 +C 92.1 S 61.3 K **0** ~$2.711 est.; hard 21.4%) — **≠** +Archer Hume (still Watch). Laya is **absent from the +scored table and also not in the named exclusion list** +— a gap vs our watch (§76), never a quoted "excluded." +GLiNER2 is mapping-excluded *theirs* (normalization +would drive calibration). Apps (Director/Loki) out. +OpenJev on the board = razorback16 DiffusionGemma **≠** +IamBusy `/v1/decide`. `notes.md` §78. + +## This hour's HIGHs already folded — dump the repos again? + +No. Hourly 0842's named HIGHs are already §73–§78 +(classifier-dev, choxos/jev-reviewer, githubnext/localjev, +NandhaKishorM/laya, @airesearch12 census, JevBench v1.2). +The product is a **recipe**, not a hit list: + +1. Ask how *p* was produced (wire-compat ≠ logit-equiv; + prompted JSON + entropy ≠ structured logit read). +2. Productize `{id, label, p}`; escalate-under-threshold + in code; mark `FALLBACK` if the head swaps (granite + 0.546 vs advertised 0.800 *theirs*). +3. Packaging ≠ new species; route by script **before** *p* + (Khmer 0.000@0.952); 0.85 still soft. +4. Pointer-not-generator: Choice of line ids, copy + verbatim, *Not found* is an answer, human tick never + overwritten (choxos ≠ egma-ai). +5. A named census is not a bake-off; a scored board + needs the score function, cal on/off, and named + latency/cost assumptions (geo-mean I/C/S/K; ≠ v1.1 + 87.6). + +Skip thin noise (JEValuate / jevspeak / fable-jev; +jev-semgrep dedicated now §86). Hard-gating a Noul as a PR +merge or quality seal is **soundness theater** +([totally-tim/jev-gate](https://github.com/totally-tim/jev-gate) +≠ jev-gateway). Archer still Watch. `notes.md` §79. + +## Does System 2 fly the drone? Should S1 wait for it? + +No. [khordoo/jev-reflex-autonomy-lab](https://github.com/khordoo/jev-reflex-autonomy-lab) +is **S1 keeps control / S2 one-use advisory**. README +*theirs*: “System 2 is advisory. It does not fly the +drone directly, and Jev does not pause while waiting +for it.” Escalate-under-threshold **without stalling** +— cousin of classifier.dev smart tier, different fail +polarity (that path re-asks and waits). S2 never +grants permission. Experimental viz, not a production +flight controller. `notes.md` §46, §80. + +## Does purple mean S2 arrived or that Jev used the advice? + +Used. Green = local context. **Purple confidence** = +that Jev decision **consumed** returned S2. +**Purple S2 bar** = arrival. Red = fail. Arrival +without a later purple confidence point is unused +advice. Mixed-initiative without a consumption mark +is theater. `notes.md` §80. + +## Is Local controller githubnext/localjev? Is the 20% gate Harbor-calibrated? + +No, and no. **Local controller** is the lab’s built-in +**rule-based** reflex (no credentials). **Live API** +is hosted TypeSafe `POST /v1/systemone` `jev-latest`. +Selecting Live API sets a 20% starting gate *theirs* +and **does not start a mission**; credentials do not +auto-switch. **≠** [githubnext/localjev](https://github.com/githubnext/localjev) +(prompted JSON) **≠** [kunchenguid/local-jev](https://github.com/kunchenguid/local-jev) +(ONNX). 20% is a UI default, still soft — schema-safe +≠ correct. Seed = repeatable geometry, not a +deterministic async replay. README says OpenRouter / +GLM 5.3; ARCHITECTURE proposed +`meta/muse-spark-1.3-contributor` — quote both +*theirs*; do not invent which is live. No pixels to +either provider. Do not copy npm / `.dev.vars`. +`notes.md` §80. + +## Is AgentGhost just toolgate with a new name? + +No. [AgentGhost](https://github.com/reddpy/AgentGhost) +wraps execution so the LLM cannot skip the judge. +toolgate is pre-exec of a *proposed* call; Jev is +not authorization there. actiongate's slogan is +evidence ≠ authority (policy is the hard gate). +jev-use is fail-open PreToolUse. AgentGhost is +fail-closed wrap-as-execution; ASK throws. +`notes.md` §83. + +## Is the JP genre atlas JevBench? Are the ★ live? + +No, and no. Application atlas, not a scored board. +Stars were research-time *theirs*. `notes.md` §84. + +## Is Akshay a vendor tutorial? Are 200× / 400× measured? + +No, and no. Independent pedagogy. Multiples are +TypeSafe ceiling *theirs*. schema-safe ≠ correct. +`notes.md` §85. + +## Is jev-semgrep Semgrep.dev? Are AND/OR/NOT embedding tricks? Is it a gate? + +No, no, and no. [uehaj/jev-semgrep](https://github.com/uehaj/jev-semgrep) +is **grep by meaning**: a Noul per line × meaning, +then boolean AND/OR/NOT on *thresholded* bits in +code. Proposition holds, not topical cosine. The +refund contrast-set (all six about a refund; only +customer-*asking* pass) is the pedagogy. Cosine +between angry customer and angry agent is close to +1. Do **not** multiply parallel p. Name collides +with [Semgrep.dev](https://semgrep.dev) SAST — +rename if both installed. Ranking fail-open; not a +gate (rh-guard skip). 0.94/0.98 is LLM-as-judge on +10×51 lines *theirs*, not Harbor. Stars ephemeral +(0 → ★42 SIGNAL → **51** this pass). **≠** jevgrep +**≠** jev-combinators. `notes.md` §86. diff --git a/.agents/skills/augustus/references/formal-methods.md b/.agents/skills/augustus/references/formal-methods.md index 1f20888..c468716 100644 --- a/.agents/skills/augustus/references/formal-methods.md +++ b/.agents/skills/augustus/references/formal-methods.md @@ -59,7 +59,8 @@ deployment control, not a feud with TypeSafe. Holding those weights, when they exist, still does not discharge a proof. A constrained-AR softmax (TypeAR, pcdServer) is still a sensor: it is not a discharged proof because the next token stayed in a declared set -(`notes.md` §42). +(`notes.md` §42). A local kev pointer-softmax is the same sensor on the +trained decision-only path (`notes.md` §45). Existing grammar: composition-algebra position 9 (verifier) — verdicts are evidence, not enforcement. Position 3 (gate) — a filter is not @@ -70,7 +71,15 @@ TOCTOU-of-Noul (§5), not a discharged obligation. **What transfers** into a mixed stack: triage which counterexample, property, or failing seed a human looks at first; score whether a production trace resembles a spec behavior; lint an artifact against a -*named, project-written* rule (`mixed-architecture.md` preference lint). +*named, project-written* rule (`mixed-architecture.md` preference lint; +Abide is the productized path of that hole, `notes.md` §47). +Skills→oxlint is the same ownership split on a linter runtime: AST / +precheck *prove* what they can; Jev scores the remainder; do not +hard-gate CI on an uncalibrated Noul +([jev-oxlint](https://github.com/cephalization/jev-oxlint) Phoenix +experiment, `notes.md` §58). PR **attention** is not correctness +([egma-ai/jev-reviewer](https://github.com/egma-ai/jev-reviewer) — +anti-soundness-theater; **not** choxos pointer-not-generator). **What does not:** closing a proof obligation, replacing TLC/Apalache/ GNATprove, or treating "DST hasn't failed this week" as a safety case. @@ -206,7 +215,9 @@ A Noul is still not a proof that the property holds, and a clean PBT run is not one either. A perception-to-decision handoff is a contract surface — the schema of objects or utterances, not the pixels or the waveform: property-test that interface, and do not pretend the Noul is -over raw pixels or raw audio. Hill-climb of that handoff: +over raw pixels or raw audio. A shared multimodal *decide* head +(blackwood-rlcd) still judges **marked candidates**, not an open click; +the act stays in code (`notes.md` §46). Hill-climb of that handoff: `validation.md`. ## 4. Deterministic simulation testing (semi-formal trio) @@ -301,6 +312,82 @@ Claiming a proof-shaped conclusion from a non-proof: - PufferLib Ocean scores as a comparative baseline (authors forbid this). - A listwise or CLIP affinity as fail-closed authorize (`judgment-class.md`). +- Engine eval / attention score sold as the *verdict* + ([game-coach](https://github.com/JoelLewis/game-coach) Wave 0: + Stockfish owns truth, Jev owns judgment; [egma-ai/jev-reviewer](https://github.com/egma-ai/jev-reviewer): + attention ≠ correctness). A Noul is not a proof the move was a + blunder or the PR is good (`notes.md` §58, §59). +- A Jev (or any judge) **score sold as eval truth** without + auditing data, scorer, runs, or claims + ([dinostomp](https://github.com/collapseindex/dinostomp): + checks the instrument, not just the score; `dinostomp jev` + tests a question like an if-statement; 99 of 189 findings + against itself; `notes.md` §62). +- Keyword privilege sold as a safety case + ([construct-auto-classifier](https://github.com/godspede/construct-auto-classifier): + `sudo status` can be a safe read; contracts on effects, not + tokens; `notes.md` §63). +- A stop-hook “green” sold as permission to skip review + ([jev-lens](https://github.com/rashedInt32/jev-lens): never + says green unless sure; never blocks the agent; + `notes.md` §63). +- Sync “Jev routing” sold without measuring decision-model + latency ([slo-router](https://github.com/zeeshan8281/slo-router): + same routes, p95 77.93→490.38 ms *theirs*; `notes.md` §63). +- A cookbook question set sold as verified without numbers + ([jev-packs](https://github.com/dtduc-git/jev-packs): + measurement owns endorsement; `notes.md` §64). +- A positive decision-model score sold as authorization + ([actiongate-jev](https://github.com/omkarghugarkar007/actiongate-jev): + Jev supplies evidence, code owns authority; `notes.md` §64). +- Vendor "calibrated" sold as frequency units you can + hard-threshold + ([does-jev-confidence-mean-anything](https://github.com/Adilmp/does-jev-confidence-mean-anything): + ranking ≠ calibration; stated ~75% vs human ~10%; + `notes.md` §64). +- Jev `done` sold as the browser task succeeded + ([ego-jev](https://github.com/jiangkoumo/ego-jev): `--until` + in code; `notes.md` §65). +- LLM summary sold as compaction + ([jev-compactor](https://github.com/edwardyen724-g/jev-compactor): + summarizer invented a path; pointer cannot; `notes.md` §65). +- The model's own hides sold as training labels + ([x-reply-filter](https://github.com/zhuyansen/x-reply-filter): + confirm-queue; never self-reinforce; `notes.md` §65). +- A leaderboard sold as a capability map + ([jev-capability-atlas](https://github.com/Zaious/jev-capability-atlas): + receipts, not a ranking; type-safe ≠ correct; `notes.md` §66). +- "Jev is weaker than 4B" sold without the thinking budget + ([jev-frontier-100](https://github.com/softpudding/jev-frontier-100): + 4B off 56.0% vs 2048 96.7%; `notes.md` §66). +- In-domain ECE sold as OOD honesty + ([jev-ood-calibration](https://github.com/scienthoon/jev-ood-calibration): + sign flips by type; unknowable policy still gets mean p 0.74; + `notes.md` §66). +- A local one-pass softmax sold as a Noul + ([jevmlx](https://github.com/bnsd55/jevmlx): schema-valid ≠ + calibrated; `notes.md` §66). +- A TLA+ run sold as "Jev is never wrong" + ([jev-labs](https://github.com/copyleftdev/jev-labs): the + invariant is never *confidently* wrong; escalate is + allowed; 0 of 1,080 golden is not a proof of zero; + synthetic, not clinical; `notes.md` §67). +- A Main Score sold as calibration, or a partial run sold + as a rank + ([jevbench](https://github.com/fstandhartinger/jevbench): + Brier/ECE reported **not scored**; native ≠ verbalized; + `notes.md` §67). +- A typed answer sold as permission to advance + ([seal](https://github.com/Reasonofmoon/seal): no seal, no + advance; coverage.path visible; mint ≠ product brain; + `notes.md` §67). +- Jev `confidence` sold as the strictest reading of the + vector + ([how-sure-is-jev](https://github.com/adarc8/how-sure-is-jev): + Choice confidence = max_prob, the most generous metric; + `notes.md` §67). +- "Type-safe" sold as "correct" ([interlock](https://github.com/somoore/interlock): + irreversible stays behind a threshold **and** a human). If the artifact would still say "verified" after you delete the checker, it was theater. @@ -421,6 +508,58 @@ because the model was confident. STPA asks what happens when the sensor is wrong, delayed, spoofed, or TOCTOU. Org/safety placement: judgment informs operators and cheap gates; it does not replace the constraint in the control structure. Mapping card: `mappings.md` §8. +**Capability kernel receipt (Empirical as architecture):** +[interlock](https://github.com/somoore/interlock) — LLM ring 3; +kernel ring 0; Jev SENSOR; `policy.py` constraint; secrets never +in the agent (`notes.md` §59). **Engine ∩ judgment (Empirical as +PRD):** [game-coach](https://github.com/JoelLewis/game-coach) — +Stockfish is the probe; Jev is the coaching sensor; Wave 0. +**Eval-instrument (Empirical as FINDINGS ledger):** +[dinostomp](https://github.com/collapseindex/dinostomp) — +the score is not the evidence; `dinostomp jev` is question +hygiene beside jevals, not a Harbor taskset (`notes.md` §62). +**Permission vs probability (Empirical as measured +suppression):** [omp-greenlight](https://github.com/SemetricLabs/omp-greenlight) +— host deny is the constraint; Jev is the sensor; operator +owns the criterion (`notes.md` §62). +**Contracts on effects, not tokens (Empirical as +certification; hunch as FM angle):** +[construct-auto-classifier](https://github.com/godspede/construct-auto-classifier) +— the named constraint is blast radius / reversibility, not +a privilege keyword. Independent risk Nouls are sensors; +`minConfidence` ∩ `riskThreshold` ∩ fast-deny is policy. +Fail-closed when the sensor is missing. Privilege ≠ verdict +(`notes.md` §63). +**Attention filter ≠ permission (Empirical as README):** +[jev-lens](https://github.com/rashedInt32/jev-lens) — never +blocks the agent; never authorizes a write. Complements +skill-broker and omp-greenlight (`notes.md` §63). +**Jev supplies evidence, code owns authority (Empirical +as slogan):** +[actiongate-jev](https://github.com/omkarghugarkar007/actiongate-jev) +— deterministic policy is the hard gate; Jev is soft +evidence. A positive score never overrides RBAC/schema/limit +(`notes.md` §64). +**Turnstile clone (Empirical as README):** +[turnstile](https://github.com/zyphr-labs/turnstile) — +policy first; Jev remainder; receipts + replay; missing +Jev → Review (`notes.md` §66). +**TLA+ compose with a Jev-class oracle (Empirical as spec ++ chaos table; 2026-09-19 ~02:38):** +[jev-labs](https://github.com/copyleftdev/jev-labs) +— TLC owns the protocol invariant; Jev is the noisy +sensor; **escalate** is the actuator when quorum is +unstable. Inverse of sensor-as-constraint: the kernel +**must not** return a confident wrong, and **may** hand +off to a human. 1,080 golden 0 wrong *theirs*; +underdetermined 34/120 still decided both ways. +Synthetic, not clinical (`notes.md` §67). +**Advance/coverage ledger (Empirical as README + +BEYOND-JEV.md):** +[seal](https://github.com/Reasonofmoon/seal) +— Strike/Jev is the sensor; Seal + coverage.path is the +constraint; Effects are the actuator. Exception queue +visible. Mint ≠ product brain (`notes.md` §67). **Kent — Data and Reality.** Models are approximations; **naming is load-bearing**. Question text, Choice sets, and Score rubrics *are* the diff --git a/.agents/skills/augustus/references/judgment-class.md b/.agents/skills/augustus/references/judgment-class.md index 567ea88..b4e9723 100644 --- a/.agents/skills/augustus/references/judgment-class.md +++ b/.agents/skills/augustus/references/judgment-class.md @@ -35,10 +35,11 @@ taxonomy with enough of *your* data (XGBoost still wins there — | Family | What it optimizes | Typical output | Use when | Watch | |---|---|---|---|---| -| **Closed decision API** (TypeSafe Jev) | Calibrated decision (proper-scoring / RLCD lineage) | Choice / Score / Noul + distributions | Default when you need act/abstain, fan-out, documented envelope | Cloud, pin version, re-measure on your data. AU health data-residency is a reason *not* to pick this family (`notes.md` §33) | -| **Open System-1 / decision-model head** (Laya, openjev, LightJev, openjev-lm, Nimble, Hume **Watch**) | Same *shape* as Jev, you host it | Same primitives or logits-as-options | Air-gap, $0/token, inspectable weights, deployment control | Self-eval duty; Laya text-only, 512 tok; vendor vs-Jev tables are claims (`notes.md` §18). A distill learns the *teacher's* answers: openjev-lm and jev-gate-student-b (`notes.md` §25, §33). Nimble is an open LoRA recipe on hard labels, not a Jev distill (`notes.md` §35). Hume's 27B dense drop is **Watch**, not a Hub checkpoint. He prefers the class name **decision models** over "system one" | +| **Closed decision API** (TypeSafe Jev) | Calibrated decision (proper-scoring / RLCD lineage) | Choice / Score / Noul + distributions | Default when you need act/abstain, fan-out, documented envelope. **Productized public HTTP** (classifier.dev): label + calibrated confidence as the contract; batch `{id,text}[]`; LLM chains fallback only (`notes.md` §73) | Cloud, pin version, re-measure on your data. AU health data-residency is a reason *not* to pick this family (`notes.md` §33). Silent fallback is a lie about the instrument — mark `FALLBACK` | +| **Open System-1 / decision-model head** (Laya, openjev, LightJev, openjev-lm, Nimble, **kev**, **blackwood-rlcd**, Hume **Watch**) | Same *shape* as Jev, you host it | Same primitives or logits-as-options | Air-gap, $0/token, inspectable weights, deployment control; **image-in now** (blackwood) without waiting for Archer | Self-eval duty; Laya text-only, 512 tok; vendor vs-Jev tables are claims (`notes.md` §18). A distill learns the *teacher's* answers: openjev-lm and jev-gate-student-b (`notes.md` §25, §33). Nimble is an open LoRA recipe on hard labels, not a Jev distill (`notes.md` §35). **kev** is a shipped Qwen2.5-0.5B LoRA + pointer readout of Archer's reconstruction — public gold, not a Jev teacher; ID ECE only (`notes.md` §45). **blackwood-rlcd** is open multimodal RLCD (CC BY-NC), Jev-compatible shim; Jev still leads general text (`notes.md` §46). Hume's 27B dense drop is **Watch**, not a Hub checkpoint. He prefers the class name **decision models** over "system one" | | **Encoder open-jev** (DeBERTa-v3-large) | Same *shape*, bidirectional encoder, public gold (not a Jev teacher) | Choice / Score / Noul from one pass | Self-host decide without a decoder; 512 tok | In-domain ECE 0.022; OOD acc 0.854→0.690. English / three public domains. `notes.md` §33 | -| **Constrained-AR surface** (TypeAR, **pcdServer**; not a species) | Next-token constraint on a pretrained generator | Distribution over allowed values | Typed fields without retraining; later fields must see earlier answers; local GGUF serving | Different objective from a proper-scoring head. TypeAR README enums ≤16; pcdServer 2–256 / 1–63 parallel fields. No abstention primitive. Compute-graph card below (`notes.md` §31, §32, §42). Public logit dump: mini-jev-runs | +| **GLiFormer wire-compat encoder** ([jeff](https://github.com/logan-markewich/jeff) on gliformer-large-v1 400M) | Labels-in-encoder; System One *wire*, not a Jev replica | choice / score / noul via `/v1/systemone`; typesafe-sdk `base_url` | Self-host the envelope when you own the GPU path and accept the accuracy gap | Normalized sigmoids, T=3.2; isolate nouls; DeBERTa tokens ≠ Jev billing. L4 HTTP ~6× cheaper; A10G direct ~24×; AG News 75.5% vs 90.5% *theirs*. CPU is *more* expensive. License null this pass. `notes.md` §60 | +| **Constrained-AR surface** (TypeAR, **pcdServer**; decision-token LoRA; not a species) | Next-token constraint on a pretrained generator | Distribution over allowed values | Typed fields without retraining; later fields must see earlier answers; local GGUF serving; train the *decision token* if you LoRA | Different objective from a proper-scoring head. TypeAR README enums ≤16; pcdServer 2–256 / 1–63 parallel fields. No abstention primitive. Decision-token QLoRA: `Foodoo1/Qwen3-14B-RLCD-Decision-LoRA` (`notes.md` §46). Compute-graph card below (`notes.md` §31, §32, §42). Public logit dump: mini-jev-runs | | **GLi\* encoder family** (GLiNER locate / GLiClass categorize / GLiNER2.5 local multi-head / GLiGuard safety schema) | One-pass labels-in-encoder; spans, sequence labels, a safety schema, or both | Spans + types; per-label sigmoid/softmax; optional relations/records | Laptop/local; large or changing label sets; "what's *in* the text" vs "what *is* the text" vs "which safety labels fire" | Affinities are not automatically a gateable P(permit). GLiGuard is not a Jev weight clone. Species map below. Not a Jev how-to and not a GLiNER or GLiGuard install | | **Listwise / pairwise discriminative ranker** | Order of a list (nDCG, softmax-over-list) | Relevance scores, not P(relevant) | Rerank a retrieved shortlist | Translation-invariant listwise losses are **not** calibrated for thresholds ([listwise vs pointwise](https://doi.org/10.48550/arxiv.2208.06164); [RCR](https://arxiv.org/html/2211.01494v2)). Fail **open** (keep retrieval order) | | **Vision scorer** | Image–text affinity or region Choice | Cosine/sigmoid affinity, or a closed region/label pick | Perception as classification over *candidates you extracted* | CLIP softmax = competition in the offered set; SigLIP sigmoid = pairwise affinity, not class-conditional p ([SigLIP](https://huggingface.co/docs/transformers/v4.39.2/en/model_doc/siglip)). Not a VLM captioner | @@ -54,7 +55,8 @@ code owns side effects." They are not aliases. ```text locate GLiNER (span NER) what's *in* the text categorize GLiClass / GLiGuard what the text is; which schema labels fire -decide Jev / Laya / openjev Choice / Score / Noul over a state +decide Jev / Laya / openjev / kev / blackwood / Domain-jev-maker LoRA Choice / Score / Noul over a state (blackwood: image-in) + jeff (GLiFormer) same *wire*, encoder backend, not a Jev replica rank listwise / cross-encoder order a retrieved shortlist perceive CLIP / SigLIP / region Choice score candidates you extracted ``` @@ -81,13 +83,155 @@ below, next to the when-to-use table. records + relations in one schema. Discourse (2026-09-18): same *agentic decision* jobs as Jev, local / free / laptop; a 36× Browser Use cost claim is a tweet, not a re-run (`notes.md` §25). + **Extractive compaction (Empirical as README behavior, 2026-09-18 + ~16:22):** + [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) + (Apache-2.0) is the named *job* on this family, not a new species. + Default checkpoint `fastino/gliner2.5-base-v1`. One encoder does + **categorize** (retention Choice: `keep_full` / `keep_evidence` / + `keep_call_only` / `drop`) and **locate** (character-offset spans); + code copies verbatim. **Not Jev, not a Noul, not a prose + summarizer, not multimodal.** Same compaction hole as + fast-jev-compaction / pi-jev-compaction (Jev Noul/Score backends); + Augustus stays backend-agnostic. Mutating tools and shell operators + are a hard `keep_full` envelope; low-confidence / invalid evidence + fail closed to `keep_full` (the *reduction* is the irreversible + act — contrast many fail-open Jev preference gates). Public default + `shadowMode: true`. Reduction is measured in characters, not + tokens; no published retention-quality rates (`notes.md` §50). + Fastino/GLiGuard sibling *class*, not a GLiGuard safety-schema + clone. Do not copy the plugin install. + **Stdout prune (Empirical as README / evals README, 2026-09-18 + ~17:15):** + [jev-pruner](https://github.com/tamaratran/jev-pruner) (MIT) is the + same evidence-preserving *family* on a **different job** and the + **Jev** backend: Noul-prune Bash stdout after a hard ≤10k / + JSON-diff-whole-doc envelope, before the main LLM sees it. Not a + summarizer. Fail-safe keep original. Archive for recovery. + Marketplace id still `fast-jev-output`. Codex is opt-in wrapper. + Same author as fast-jev-compaction; complementary (`notes.md` + §53). Do not copy the plugin. + **Code-graph indexer (Empirical as README behavior; 10–50× is a + target, 2026-09-18 ~16:48):** + [s1-graphify-indexer](https://github.com/GreyssonEnterprises/s1-graphify-indexer) + (+ sibling `s1-indexer`) uses GLiNER2 as the default System-1 + backend to build a semantic code graph and escalates an LLM only + on the ambiguous tail **and only if the backend loaded**. Jev / + Needle stubs `available=False` until implemented. If GLiNER2 + cannot load: degraded file-node graph, exit nonzero, do not dump + the repo to `S1_LLM_CMD`. Query does not invent edges. GitHub + one-liner 10–50× is **unfilled** — not Empirical (`notes.md` §51). + Same locate family as compaction; different hole (index vs + keep/drop). License not on GitHub this pass. + **Computer-use selection (Empirical as README / architecture + behavior, 2026-09-18 ~16:56):** + [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) + (MIT) is a *different named job* on the GLi\* encoder family, not + a new species and **not GLiNER2.5**. Checkpoint + `fastino/gliner2-multi-v1`. Adaptation of + [jev-ultrafast](https://github.com/browser-use/jev-ultrafast): + local GLiNER2 extracts requirements and **scores observed** + a11y/DOM controls; code clicks. **No screenshots. No generated + selectors.** Same observe→score-among-candidates→code-acts hole + as jev-ultrafast / solari-reflex (Jev backends) and + [laya-mind2web](https://huggingface.co/ShaunSpark/laya-mind2web-browser-agent) + (Laya over DOM element indices). Contrast blackwood-rlcd + (screenshot + marked letters → Choice). Hybrid: local decide; + remote text helper only for TYPE. `DONE` is loop termination, not + verified success. Inspector scores are not calibrated P(task + success). Their Flights demo (12.20 s / 13.785 s / ~$0.0001 API) + is a demonstration, not a bake-off (`notes.md` §52). Do not copy + `uv` / `.env`. + **Specialist computer-use S1 (Empirical as README / MODEL_CARD + behavior, 2026-09-18 ~17:21; weights Watch):** + [Cua-S1](https://github.com/trycua/cua/tree/main/libs/cua-s1) + (`trycua/cua`, MIT source) is a **parallel "System One" name**, not + TypeSafe Jev and not GLiNER. Reference `tinyx`: byte encoder + + option-attention head. Per observed element: fill (from extracted + `Label: value`) / check / click / skip. Does not generate values + or selectors. Plan ≠ execute; dry-run default; `execute` and + `submit` independent opt-ins; fail-closed on unknown checkbox + state. Profile `cua-s1-form-v0` is source-only — no weights, no + checkpoint scores (`notes.md` §54). Do not copy `uv` / MCP. + **Harness productization (Empirical as PR-body architecture + + their local eval, 2026-09-19 ~00:48; draft Watch):** + [Stagehand #2955](https://github.com/browserbase/stagehand/pull/2955) + (5/5 of #2951–#2955, all OPEN draft) puts the same + observe→score-among-candidates→code-acts hole inside + Browserbase Stagehand. Jev picks a11y elements; code copies text + or acts. Extract `"off"` | `"judge"` | `"pick"`. Schema / completion + gate / screenshot-always-LLM in **code**; LLM fallback. Their + extract card (gemini-3.8-flash, 25×3): 37/75 no-LLM ~0.5 s vs + 4.37 s; 69/75 vs 23/25 (92% both); LLM-off 36/75 — pick ≠ + replacement. Do not merge with jev-ultrafast / gliner2-ultrafast + / Cua-S1 clocks. Not multimodal pixels on the pick path + (`notes.md` §57). Do not copy `experimentalJevAct`. + **OCR+AX desktop product (Empirical as README; MIT + **427★**; 2026-09-19 ~09:51):** + [typesafe-computer-use](https://github.com/awlevin/typesafe-computer-use) + is the same observe→score-among-candidates→code-acts + hole on a **Mac**, with **hosted TypeSafe Jev**, not + GLiNER2 and **not** Cua-S1. Vision OCR (crop+tile) + + AX → numbered items → three/four Choices + (`kind`/`item`/`site`/`offscreen`) → deterministic + click/type. **The decision never ships a screenshot + to frontier.** The one-shot **answer** writer may + receive the capture because OCR misreads — a reader + packet, not the Choice. Overlapping options always + read as doubt. AX is a bonus, never a replacement + (Spotify 0 *theirs*). $0.0002 vs Opus $0.032 (155×) + is *theirs* on **one screenshot**, not a Harbor + taskset. 0.4 / 0.5 still soft. **≠** jev-ultrafast + **≠** jev-macos-loop OmniParser **≠** camoufox + **≠** blackwood-rlcd (`notes.md` §81). Do not copy + `uv` / `.env`. Skip Archer. + **ASR voice-browser product (Empirical as README; MIT + **103★**; 2026-09-19 ~10:01):** + [jev-voice-browser](https://github.com/moritzkremb/jev-voice-browser) + is the same observe→score-among-candidates→code-acts + hole with **ASR** as the producer: Web Speech + transcript + numbered Playwright elements → hosted + TypeSafe (9–11 questions) → code clicks. **Jev never + generates.** Partial-speech wait is VOI. Spoken + confirm is convenience, not auth. Numbered overlays + disambiguate without another model. 27/27 fixtures + *theirs*. **≠** jev-voice-control **≠** nikolas-j + **≠** typesafe-computer-use OCR (`notes.md` §82). + **Wrap-as-execution product (Empirical as README; MIT + **2★**; 2026-09-19 ~10:20):** + [AgentGhost](https://github.com/reddpy/AgentGhost) + puts ALLOW/ASK/DENY *on the actuator path*. The + model cannot skip the wrap. Rules first; ASK + throws; fail-closed. Judge is a slot, not a + vendor lock. **≠** actiongate (evidence ≠ + authority) **≠** toolgate **≠** jev-use + fail-open. rh-guard owns the gate cousin + (`notes.md` §83). Do not copy `npm` / `.env`. + Skip Archer. - **Decide.** Typed Choice/Score/Noul with a decision/proper-scoring objective. That is Jev's product claim. Open heads copy the *shape*; distillation copies the *teacher* (openjev-lm, jev-gate-student-b). + **Domain specialist LoRA** ([Domain-jev-maker](https://github.com/help-er/Domain-jev-maker)) + trains on **independent gold** (CLINC-150), not Jev answers — pick + it when downstream reads p; few-shot hosted when only argmax + (`notes.md` §60). + **Do not distill Jev as teacher of record.** + [jev-triage](https://github.com/ThyFriendlyFox/jev-triage) + logs full distributions for a *local* student and keeps + **real outcome labels** as targets. Author ~68% ceiling + compounds errors. Distinct from Domain-jev-maker + (independent gold) and openjev-lm (teacher-copy) + (`notes.md` §61). Encoder open-jev (DeBERTa) copies the shape on public gold and still - owes OOD self-eval. A constrained autoregressive decode can emit a - label and still not be this species — compute-graph card below. - Hume's announced 27B dense drop is Watch. + owes OOD self-eval. **kev** copies the Archer readout (block-causal + isolation, pointer head, CE vs labelled public gold) and still owes + *your* ECE — ID numbers are not OOD. A constrained autoregressive + decode can emit a label and still not be this species — compute-graph + card below. Hume's announced 27B dense drop is Watch. + **blackwood-rlcd** is the open multimodal *decide* head that ships now + (screenshot + marked candidates → Choice; CC BY-NC; Jev still leads + general text; `notes.md` §46) — not that drop, and not a vision-scorer + affinity. - **Categorize (safety schema).** [GLiGuard](https://github.com/fastino-ai/GLiGuard) ([arXiv 2605.07982](https://arxiv.org/abs/2605.07982); Zaratiana, Newhauser, Hurn-Maloney, Lewis, Fastino): a GLiNER2 encoder that @@ -113,8 +257,17 @@ below, next to the when-to-use table. **Surfaces.** GLiGuard is for LLM input/output safety. [rh-guard](https://github.com/24601/rh-guard) is a coding-agent reward-hack gate (README fetched this pass); jevgate is the - allowlist-then-judge shape (`mappings.md` §18). Different holes. - Do not point one model at both, and do not copy a hook install here. + allowlist-then-judge shape (`mappings.md` §18); + [Abide](https://github.com/coldteadotai/abide) is project-instruction + soft rules on diffs (`mixed-architecture.md`). + [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) + is extractive context compaction (locate + categorize on tool + transcripts; Fastino sibling class, not a GLiGuard clone) + (`notes.md` §50). + [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) + is browser computer-use scoring among observed controls (GLiNER2 + `gliner2-multi-v1`, not 2.5; `notes.md` §52). Different holes. + Do not point one model at every hole, and do not copy a hook install here. **Aggregation is policy-in-code, already taught.** The README's benchmark rule ORs unsafe / non-benign prompt labels and lets refusal @@ -163,7 +316,11 @@ multi-head (GLiNER2.5) can do both plus relations. Classification sigmoid/softmax *can* be a cheap multi-label sieve. It is not, without your calibration plot, a decision API. Great for "which of these 80 tags fire" or "which spans are the allergy / the amount / the verb"; -not a silent fail-closed authorize. The 255-option Choice limit is +not a silent fail-closed authorize. Named compaction job +([gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction)): +locate + categorize, then code copies exact offsets; fail-closed +`keep_full` is the *reduction* policy, not a Noul (`notes.md` §50). +The 255-option Choice limit is Jev's, not the class's — this family is why. Rule of composition (`applied-mappings.md` §4): @@ -195,7 +352,10 @@ parser. RAM / AX tree / object JSON → closed action or region set → Choice. Prices and dates stay in code. Launch-week recipes: typesafe-mario, jev-drone (classical CV → symbols, Jev advisory), lizard-agent - (visible elements only). The model never sees a screenshot. + (visible elements only). Encoder-backend cousin: + [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) + scores observed a11y/DOM controls with local GLiNER2; code clicks + (`notes.md` §52). The model never sees a screenshot. 2. **Region / label Choice over extracted boxes.** Perception (detector, grid, SAM, OCR boxes) proposes candidates; a scorer picks. `hr98w/jev-visual`: @@ -241,26 +401,49 @@ capability shift, independent of vendor: 4. **Perception is not narration.** Computer-use and robotics that caption the world then plan in prose are on the wrong side of the class. Extract candidates, score, act; generate text only when - something must be typed. + something must be typed. **Receipt (ARCHITECTURE *theirs*):** + [khordoo/jev-reflex-autonomy-lab](https://github.com/khordoo/jev-reflex-autonomy-lab) + sends **no graphical input** to either provider; sensors emit + geometry; code never labels safest; planner narrative is + excluded from Jev input (`notes.md` §80). **Receipt + (README *theirs*):** + [typesafe-computer-use](https://github.com/awlevin/typesafe-computer-use) + never ships a screenshot to frontier for the + *decision*; OCR+AX text-state → TypeSafe Choices → + code clicks; writer only for free text; the one-shot + answer reader may receive the capture (`notes.md` + §81). Overlapping options = false low confidence. + 155× is one screenshot. **≠** jev-ultrafast **≠** + cua-s1. **Receipt (README *theirs*):** + [jev-voice-browser](https://github.com/moritzkremb/jev-voice-browser) + never ships a waveform to Jev; ASR → transcript → + Choices → Playwright; pointer spans; spoken confirm + ≠ auth (`notes.md` §82). Compose with OCR, don't + collapse. Hosted TypeSafe product APIs and prompted + LocalJev are **not** a shared-prefix multimodal + species. Skip Archer. 5. **Open heads and GLi\* make the control plane local.** Air-gap / on-device / laptop (GLiNER2.5 74M–287M CPU-first; openjev-lm 0.5B LoRA overnight on 6 vCPU; encoder open-jev DeBERTa-v3-large 434M; - jev-gate-student-b 0.5B LoRA memory gate) become newly feasible *if* + jev-gate-student-b 0.5B LoRA memory gate; **kev** 0.5B LoRA + pointer, + ~160 ms / 6 questions on an M5) become newly feasible *if* you accept self-eval and envelope limits. They do not make calibration optional. Distilling a hosted teacher is not independent gold. Hume's 27B dense drop is the large-local Watch, not a third how-to. -6. **Cross-modal is still thin.** Discourse, GLiNER/GLiClass, Laya, and - the encoder open-jev are text-first. Vision is a scoring pattern - (above), not a shipped omni decision API. Hume reports that a - multimodal *base* plus text post-training generalizes to images with - little intentional multimodal training — a Watch claim, not a - recipe (`notes.md` §33). Treat "Jev but for images" as a hole to fill - with the vision-scorer family or with that drop *when it ships*. - Locate (spans on a screenshot OCR) is still locate, not perceive. - Pixel-free computer-use (jev-macos-loop, jev-mobile) keeps pixels on - the device and sends text-only decisions. SAM 3.1 or ASR then Jev - is specialist composition, not the Watch drop (`notes.md` §39). +6. **Cross-modal has an open decide head now; Archer is still Watch.** + Discourse, GLiNER/GLiClass, Laya, kev, and the encoder open-jev are + text-first. Vision scorers (CLIP/SigLIP) remain a scoring pattern, + not a decision API. [`BlackwoodAI/blackwood-rlcd`](https://huggingface.co/BlackwoodAI/blackwood-rlcd) + is the shipped omni *decide* model: screenshot or text in, typed + Choice out, Jev-compatible shim (`notes.md` §46). Hume's 27B dense + drop stays **Watch** (no Hub weights this pass; multimodal, no audio). + Prefer specialist composition (SAM / ASR / OCR → schema → System One) + when you already have the producer; prefer a shared multimodal + decision model when the joint of pixels and options matters. Pixel-free + computer-use (jev-macos-loop, jev-mobile) still keeps pixels on the + device and sends text-only decisions. Locate (spans on a screenshot + OCR) is still locate, not perceive. 7. **The agent that only has a generator is incomplete.** The missing organ is a judgment-class model plus policy in code — not another prompt. The agent that only has a ranker is also incomplete: it can @@ -376,14 +559,17 @@ program. | Need | Place | Do not | |---|---|---| | Calibrated p(y\|x) over a closed set | Trained decision-only head (Jev, or an open head you have proper-scored and measured on your labels) | Threshold a generated "90%", an affinity you have not calibrated, TypeAR constrained scores, or a LoRA student's agreement with the teacher | +| Laptop-local System One API for development / eval | **kev** — trained decision-only readout; official SDK with a `base_url` change (`notes.md` §45) | Treat 0.5B ID ECE as a knowledge or frontier substitute, or as OOD calibration | | Dependent sequential decisions | Constrained AR that conditions later steps on earlier answers (TypeAR sequential), or code-owned transitions and a new request per stage | Treat sibling questions on one request as if they attend each other | -| Open multimodal self-host / data-residency | Hume's announced **decision-model** drop **when it ships** (Qwen3.8 27B **dense**, 265k, multimodal, no audio; one forward pass locally once AR is removed; MoE next then shrink). Driver: healthcare AU residency, not anti-TypeSafe | Ship on "smarter than Jev." That is his early claim, against his own order-sensitivity and in-distribution calibration warnings. **WATCH** — no Hub weights this pass. Laya remains text-only. jev-visual is region Choice, not this drop | +| Open multimodal self-host / data-residency *now* | **blackwood-rlcd** — trained decision-only readout with image-in; Jev-compatible shim; CC BY-NC (`notes.md` §46) | Wait for Archer's 27B. Treat screenshot-vs-Jev-text as the same input. Threshold a commercial workflow on a non-commercial license. Skip self-eval because web-element acc is 0.907 | +| Open multimodal self-host / data-residency *when it ships* | Hume's announced **decision-model** drop **when it ships** (Qwen3.8 27B **dense**, 265k, multimodal, no audio; one forward pass locally once AR is removed; MoE next then shrink). Driver: healthcare AU residency, not anti-TypeSafe | Ship on "smarter than Jev." That is his early claim, against his own order-sensitivity and in-distribution calibration warnings. **WATCH** — no Hub weights this pass. Laya remains text-only. kev is text-only. jev-visual is region Choice, not this drop | | Image-in now, different graph | Diffusion structured reads that already accept images on a Jev-shaped interface ([djev-spark](https://github.com/mmastrac/djev-spark)) | Wait on the row above for image-in, or treat this graph as a proof it beats a decision head | ### When to use which decision surface -Five *surfaces*, not five species — plus an open recipe and a diffusion -graph that are not extra species either. Pick from the hole and these +Five *surfaces*, not five species — plus open recipes (Nimble; **kev** +as the runnable Archer reconstruction) and a diffusion graph that are +not extra species either. Pick from the hole and these axes; do not start from a logo. Hume prefers the class name **decision models** over "system one" ([tweet](https://x.com/4rcherhume/status/2100604161821979134)). This @@ -395,24 +581,327 @@ is the generator, not a sixth surface. | Surface | Calibration | VOI / gather | Latency / $ | Deployment control | Multimodal | Enum size | |---|---|---|---|---|---|---| | **Proprietary Jev** | Decision objective; in-dist ECE 0.0313, OOD collapse (`notes.md` §7). Choice `confidence` is arithmetic on the distribution (§31) | Independent questions cheap; sequential gather is a new request | Cloud envelope; ~$0.042/MTok input (their figure, unreproduced; `/pricing` 404 on 2026-09-18, `notes.md` §1, §42) | No weights. AU health data cannot ride this API if residency forbids it | Text. jev-visual is region Choice | ≤255 Choice | +| **blackwood-rlcd** (open multimodal RLCD; CC BY-NC) | Decision objective; temp-calibrated. Card ECE **0.037** on 300 web steps; letter-shuffle flip **0.133** vs Jev 1.13 **0.587**. Jev still leads general text **0.850** vs **0.786** (`notes.md` §46) | Screenshot/DOM candidates code already marked → Choice. Not gather-as-act | ~200 ms / decision 1×H100 (their figure) | Self-host; non-commercial license | **Image-in now.** Not Archer Watch. Not CLIP | Lettered candidates; Jev-shaped Choice | | **Archer open decision-model** | **Watch.** No Hub weights this pass. "Smarter than Jev" is a claim against *his* calibration/order warnings | Same *hole* as Jev when it ships | 27B dense for one-forward-pass local speed once AR is removed; MoE next, then shrink. Quant-friendly is a claim | Healthcare AU data-residency / deployment control, **not** anti-TypeSafe | Multimodal, no audio. Text post-training reportedly generalizes to images with little intentional multimodal training | Unknown until the drop | -| **TypeAR / pcdServer** (constrained AR) | Next-token constraint ≠ Noul. No abstention primitive. Public logit dump: [`Mikhail/mini-jev-runs`](https://huggingface.co/datasets/Mikhail/mini-jev-runs) (27.9k; scores "deliberately *not* calibrated") | TypeAR sequential conditions later fields; pcdServer batches independent fields after one prefix. Neither is gather-as-act | TypeAR 5.8× is *their* K=16 boolean example. pcdServer: native llama.cpp, Apple+Linux | Self-host the generator / GGUF | Whatever the base model has | TypeAR enums ≤16; pcdServer 2–256 strings, 1–63 fields | +| **TypeAR / pcdServer** (constrained AR) | Next-token constraint ≠ Noul. No abstention primitive. Public logit dump: [`Mikhail/mini-jev-runs`](https://huggingface.co/datasets/Mikhail/mini-jev-runs) (27.9k; scores "deliberately *not* calibrated"). Decision-token QLoRA trains *that* token under parallel constrained decode (`Foodoo1/Qwen3-14B-RLCD-Decision-LoRA`; synthetic fraud receipt, not a financial product; `notes.md` §46). **Harbor cousin:** local MLX PCD Qwen2.5-1.5B on toxic-chat n=50: O(1) / p50 227.2 ms / acc 52% / Brier **0.3884** vs Jev 84.0% / Brier **0.1096** (`system-one-benchmark`; `notes.md` §61) | TypeAR sequential conditions later fields; pcdServer batches independent fields after one prefix. Neither is gather-as-act | TypeAR 5.8× is *their* K=16 boolean example. pcdServer: native llama.cpp, Apple+Linux. Foodoo1: ~234 ms / 4-field broadcast on RTX 3090 4-bit (their figure) | Self-host the generator / GGUF / adapter | Whatever the base model has | TypeAR enums ≤16; pcdServer 2–256 strings, 1–63 fields | | **Encoder open-jev** (DeBERTa-v3-large 434M) | Public gold, CE+Brier, val temperature. In-domain ECE 0.022 / acc 0.854; OOD acc 0.690 / ECE 0.035. **Not** a Jev teacher-copy | One pass over state + all questions; 512 tok | Author: 28 ms / 10 questions H100; 1.8 s / 4q M1 Max CPU | apache-2.0, self-host | Text | Jev-shaped 255 / Score 2–10 / Noul; 512 ctx | -| **Tiny LoRA distill** (jev-gate-student-b) | Teacher-copy. P(relevant) from yes/no logits. Held-out n=60 vs vanilla 0.5B; 148,160-row corpus | Memory-gating / context sieve; **fail-open** on errors | Qwen2.5-0.5B LoRA; ~59 ms RTX 3060 | Local, apache-2.0 | Text | Binary relevance | +| **Tiny LoRA distill** (jev-gate-student-b) | Teacher-copy. P(relevant) from yes/no logits. Held-out n=60 vs vanilla 0.5B; 148,160-row corpus. HF card **unchanged** ~17:48 vs §33 (MAE 0.187 / Pearson 0.791 / 90%; ~59 ms RTX 3060; fail-open) | Memory-gating / context sieve; **fail-open** on errors | Qwen2.5-0.5B LoRA; ~59 ms RTX 3060 | Local, apache-2.0 | Text | Binary relevance | +| **Domain LoRA specialist** (Domain-jev-maker) | Independent CLINC gold, soft targets, pointer readout. **Not** a Jev teacher-copy. Calibration gap is the product: KL 0.168 vs hosted 0.580 banking; few-shot hosted matches argmax (McNemar n.s.) | Threshold / deferral / EU that *reads* p; skip when only argmax | ~0.5 s / request on 8 GB GPU *theirs*; 1.5B LoRA | Self-host; MIT | Text | Domain K + abstain; one forward pass | | **Nimble** (open LoRA recipe, not a distill) | Hard synthetic labels. They say temperature was not tuned to correctness rates. 324-row agreement is their receipt, not an ECE (`notes.md` §35) | Not a gather primitive | Their latency table, not re-run | Self-host the adapter. Model card Apache-2.0; repo license absent | Text only | Enum ≤26; 2,048 tokens | +| **kev** (Qwen2.5-0.5B LoRA + pointer; Apache-2.0) | Public gold, CE. Held-out ECE 0.065 (0.031 after T=1.47); acc 0.799 on 1,350 ID questions. Isolation exact. **Not** a Jev teacher-copy (`notes.md` §45) | Laptop-local System One drop-in for development/eval; independent questions, one prefill | ~160 ms / 6 questions; ~1h45m train on M5; 38 MB adapter | Self-host; official `typesafe-sdk` with `base_url` | Text. Not multimodal. 0.5B knowledge | noul / choice 2–255 / score | | **Diffusion structured reads** (djev-spark) | Interface claim only. **Hypothesis** it beats a decision head on your labels (`notes.md` §36) | Optional sequential chunks, text-only | Their GX10 tables, not a class benchmark | DGX Spark container. Do not copy the route | Images are an extension; think and sequential reject images | README criteria, not copied here | **Three open paths** (not three species, not extra when-to-use rows): encoder open-jev (DeBERTa, public gold); AR constrained decode (TypeAR -Python/SGLang, pcdServer native GGUF); trained decision-only (Laya / -Nimble / Archer **Watch**). Pick from the hole. A constrained softmax +Python/SGLang, pcdServer native GGUF; decision-token LoRA on that graph); +trained decision-only (Laya / Nimble / **kev** / **blackwood-rlcd** / +Archer **Watch**). **kev** is the cleanest *runnable* productization of +Archer's reconstruction on that third path (API-compatible; text-only; +`notes.md` §45). **blackwood-rlcd** is the third path with **image-in +now** (CC BY-NC; Jev-compatible shim; `notes.md` §46). Watch stays Watch. +Encoder vs decoder **replicas** of that third path: DeBERTa is public gold with an OOD +drop; openjev-lm / jev-gate LoRAs are teacher-copies with named receipts +(overnight 6-vCPU, $0/call — economics, not a new species; `notes.md` +§25, §44). kev is public gold on the same 0.5B backbone as openjev-lm, +not a teacher-copy. Pick from the hole. A constrained softmax is still not a Noul. Laya companion packaging this hour: [`laya-typed-decisions`](https://huggingface.co/convaiinnovations/laya-typed-decisions) (same 421.3M; acc 0.766 / Brier 0.066 on `LocalLLaMA/typed-decisions`, -unverified — do not overwrite `notes.md` §18). +unverified — do not overwrite `notes.md` §18). Shared bake-off this +hour: [`pngwn/open-jev-laya-bench`](https://huggingface.co/datasets/pngwn/open-jev-laya-bench) +(26 + 9 tasks, 11,959 items; ECE/NLL/Brier; **not** TypeSafe Jev vs +Laya; LLM-as-judge is not the score — `notes.md` §46, `validation.md`). +**Contract-compatible local surface, not a fourth path:** +[`us/jev-local`](https://github.com/us/jev-local) speaks `POST +/v1/systemone` so an SDK `base_url` drop-in works offline. **Default +scorer is a deterministic stub** until `JEVLOCAL_SCORER=hf`. A green +smoke test on the stub is not a local decision model (`notes.md` §48). +**Independent open-Jev class (Empirical as their docs; +2026-09-19 ~03:38):** +[openvons](https://github.com/genai-craft/openvons) +(Apache-2.0 LICENSE; GitHub SPDX NOASSERTION; **7★**) +implements finite-choice+prob for LM / vision / voice +with a mandatory none-of-the-above and +execute/confirm/reject policy. Speaks +`POST /v1/systemone` so an SDK `base_url` drop-in works — +**wire-compat, not a TypeSafe replica** (same duty as +jeff / jev-local). JevPick is menu decode (3.2–4.8× +byte-identical *theirs*), not a Noul. Flutter on-device: +audio/image stay on the phone. Unrelated to TypeSafe; no +TypeSafe API output used. Archer still Watch +(`notes.md` §68). Do not copy `uv` / APK. +**Independent local OpenJev `/v1/decide` (Empirical as +their RESULTS.md; not a TypeSafe drop-in; 2026-09-19 +~04:39):** +[OpenJev](https://github.com/IamBusy/OpenJev) (Apache-2.0) +is a 0.6B Qwen3 + LoRA + scalar head. Supervised CE + +held-out temperature — **not** RLCD, **not** a TypeSafe +replica. Hub `IamBusy/OpenJev-Branch-v0.3`. *Theirs:* +**45/60** vs v0.2 39/60; reversal 100% vs 68.75%. +`POST /v1/decide` is an OpenJev contract. Distinct from +[hraness/sysone](https://github.com/hraness/sysone) +"OpenJev runners" (loopback gateway). Do not copy `uv` +(`notes.md` §69). +**SemIf as `/v1/systemone` runoff (Empirical as their +latency table; wire-compat ≠ replica; 2026-09-19 ~04:39):** +[semif-serve](https://github.com/dddanielliu/semif-serve) +(pyproject MIT / GitHub SPDX null) serves SemIf behind +the Jev wire. No option ceiling (runoff). RTX 3080 Ti +Qwen3.5-4B **1164 ms** vs hosted Jev median **178 ms** +*theirs*. Confidence inferred for choice; runoff is a +product, not a single softmax; `output_tokens` always 0. +MiniCPM5-2B unusable. `--stub` needs no GPU. Same warning +as jeff / openvons / jev-local. Do not copy CUDA how-to +(`notes.md` §69). +**Competing NAR claims as an audit object, not a +swap (2026-09-19 ~06:43):** +[openJev-verdict-2.0](https://github.com/Heman10x-NGU/openJev-verdict-2.0) +(README Apache-2.0 / GitHub SPDX NOASSERTION) is a +149.6M ModernBERT-base NAR with Choice/Score/Noul and a +WebGPU demo. README *theirs* on +`LocalLLaMA/typed-decisions` N=2000: acc **77.10%** / +Brier **0.0636** / ECE corr **0.0144** / dist +**0.1513**. The TypeSafe Jev table row is a +Laya-catalogued vendor baseline, **not** an independent +run. Dual-channel ECE is a real design fork — and +**like-for-like is the rule**. Open PR #1 already +audits the launch write-up (throughput≠latency; Laya +gap inside CI → parity; Jev 14.40% slightly lower on +distribution ECE). **Do not endorse.** Distinct from +IamBusy/OpenJev `/v1/decide`, openvons, grande, and +openjev-lm. Do not copy train/serve how-to +(`notes.md` §71). +**1-token logprob local endpoint (constrained-AR +surface, not a trained head; 2026-09-19 ~07:49):** +[chakuho](https://github.com/taku-me/chakuho) (MIT) +serves `POST /v1/systemone` by reading one next-token +logprob over caller-enumerated labels (vLLM / Ollama +instruct). Softmax-over-labels **≠ Noul**. `coverage` +is format-mass, not correctness. GUI 336-case *theirs*: +27B **95%/92%** vs Jev Gateway **89%/82%**; `__none__` +gold 97% vs 8B 10%. Cousin jevify / TypeAR / pcdServer +/ jevmlx. Do not copy `uv run chakuho serve` +(`notes.md` §72). +**Open replica inference engine (argmax-parity +speedup, not ECE; 2026-09-19 ~07:49):** +[jevinf](https://github.com/zerodegress/jevinf) (MIT; +Python ≥3.14) runs NanoJev / decider-2b / Laya with +segmented forwards + prefix reuse, then the Jev wire. +Only `torch-mps` is wired. *Theirs:* **2.57×** / +**2.27×** at **100% argmax agreement**. Wire-compat ≠ +TypeSafe replica. Do not copy serve/port +(`notes.md` §72). +**Laya multilingual class expansion (English checkpoint +confident-wrong OOD; 2026-09-19 ~07:49):** +[`laya-multilingual`](https://huggingface.co/convaiinnovations/laya-multilingual) +(Apache-2.0; mmBERT-base **322M**; 9 Hub likes). MASSIVE +51-lang *theirs*: **0.366 / 0.387 / 45 of 51** vs +English `laya` **0.227 / 0.733 / 23 of 51**. Khmer +**0.000 acc at 0.952 confidence**; mean conf never < +0.885 — **gating cannot catch it**. Route by script +before the forward pass. Ships uncalibrated (ECE +**0.314 → 0.106** after T *theirs*). Weaker on English +(0.619 vs 0.684) — route, do not replace. Sibling of +already-folded `laya-typed-decisions`. Do not copy +`pip install laya` (`notes.md` §72). +**Laya GitHub/PyPI packaging (not a new species; +2026-09-19 ~09:07):** +[`NandhaKishorM/laya`](https://github.com/NandhaKishorM/laya) +(Apache-2.0; **710★**). SDK + `Router` over the three +Hub checkpoints already watched. T4 *theirs*: 1q +**32.8 ms**; post-T ECE **0.081** vs Jev **0.246**; +Banking77 **0.425** vs Jev **0.870** (token budget; +72 vs 77 labels). typed-decisions **0.766** is a +fine-tune (base 0.362/0.342 vs majority 0.461). Jev +rows third-party unpublished-here. 0.85 gating is +*theirs*, still soft. **≠** TypeSafe `/v1/systemone`. +**≠** githubnext/localjev. Do not copy pip / preload / +`head_max_len` (`notes.md` §76). +**External openjev census (tweet, not a new species; +2026-09-19 ~09:14):** +[@airesearch12](https://x.com/airesearch12/status/2101259522933186879) +lists ~18 named openjevs including GLiNER2 and +routers. That is a **class-boundary**, not identity. +Incomplete vs Laya / githubnext/localjev / kev / +TypeAR / openvons. **≠** jevbench v1.1. Watch +[jev-models](https://benchmarkheaven.com/jev-models); +do not paste live ranks (`notes.md` §77). +**JevBench v1.2 scored class table (live board; +2026-09-19 ~09:24):** +[jev-models](https://benchmarkheaven.com/jev-models) ++ [`RESULTS-v1.2.md`](https://github.com/fstandhartinger/jevbench/blob/main/RESULTS-v1.2.md) +(SHA `fdfab1a2`). 15 systems × 534 decisions. Geo-mean +I/C/S/K; Jev **75.3** / SemIf **74.6** *theirs*. +Instruction models (Luna/Gemini/DeepSeek/Qwen3.8) sit +in the same table as NAR rebuilds — class-boundary is +the typed task. OpenJev on the board = razorback16 +DiffusionGemma **≠** IamBusy `/v1/decide`. Laya +absent (gap). GLiNER2 mapping-excluded. Qwen3.8 27B +**≠** Archer. **≠** v1.1 87.6 (`notes.md` §78). +**Schema-conditioned DeBERTa scorer (Hub; GitHub 404; +2026-09-19 ~07:49):** +[`jev-schema-scorer-deberta-v3-large`](https://huggingface.co/mobarmg/jev-schema-scorer-deberta-v3-large) +(MIT; 4 likes). One scalar head per `(state, +question+candidate)`; **code** groups into +Choice/Noul/Score. v2 Choice **0.841** *theirs* +(chance 0.214). Peaked p = ranking, not calibration. +Distinct from com-kotobalabs/open-jev-deberta-v3-large. +Do not copy `schema_scorer.py` (`notes.md` §72). +**ONNX replica of Laya:** +[`Mattepiu/laya-onnx`](https://huggingface.co/Mattepiu/laya-onnx) +(~15 ms CPU for one Noul, their card). Do not copy the inherited +vs-Jev accuracy table (`notes.md` §18). +**Complete Laya → browser ONNX (distinct replica; 2026-09-19 +~00:39):** +[`gqgs/laya-onnx`](https://github.com/gqgs/laya-onnx) — full +Laya heads + option scorer + qtype embeddings + act/escalate +as int8 ONNX (496.8 MiB). Conversion smoke, not +task-accuracy or calibration. License null. Distinct from +Mattepiu/laya-onnx. Primary omni archive stays Jev-omni. +Do not copy npm (`notes.md` §64). +**Local ModernBERT `/v1/systemone` approximation (not +equivalence; measured 2026-09-19 ~05:46):** +[`kunchenguid/local-jev`](https://github.com/kunchenguid/local-jev) +— ONNX ModernBERT-large-zeroshot-v2.0; "API-compatible +local approximation — **not** behavioral equivalence." +`confidence` omitted. 136 checkpoints *theirs*: done +**30%** / shape **57%** / Pearson r **−0.06** vs live +Jev; gold done 26% vs 87%; 112 min vs 21 s. Distinct from +jev-local's stub and jeff's GLiFormer. MIT (`notes.md` +§64, §70). +**GitHub Next prompted-JSON `/v1/systemone` (Empirical as +README + 1,200-request eval; 2026-09-19 ~08:56):** +[`githubnext/localjev`](https://github.com/githubnext/localjev) +— MIT; **261★**. Bun bridge: DiffusionGemma through +ordinary Chat Completions; TypeSafe SDK drop-in +(`jev-latest` / `jev-preview` aliases). **Wire-compat ≠ +logit-equiv:** the model emits a JSON probability vector; +code validates/retries, normalizes, and computes +entropy-based confidence. OpenJev +([razorback16/openjev](https://github.com/razorback16/openjev)) +reads logits via structured-read vLLM extensions. +**≠** [kunchenguid/local-jev](https://github.com/kunchenguid/local-jev) +(ONNX hyphenated namesake). **≠** IamBusy/OpenJev +`/v1/decide`. Bake-off *theirs* (prompted pipeline, not +logits): Qwen3.6 short macro **76.7%**; Gemma 4 26B-A4B +**75.0%**; DiffusionGemma **74.2%**; no definitive winner; +do not treat as calibrated. LM Studio cannot load +DiffusionGemma (18 Sep 2026). Do not copy bun / `.env` +(`notes.md` §75). +**Rust/WebGPU System One (Empirical as JGLUE + isolation; +2026-09-19 ~05:46):** +[`bokuweb/grande`](https://github.com/bokuweb/grande) — +Archer/kev-shaped shared-state branches; `POST /v1/systemone`. +JGLUE *theirs*: E2B zshot JNLI 0.614 ECE 0.252→0.088 T=2.81; +JCQA 0.853; trained 270M 0.710/0.710. Isolation sibling +0.098 / state 0.996. Packed Δmax 7e-5. License null. Softmax +≠ Noul until T. Not Archer Watch. Do not copy cargo +(`notes.md` §70). +**Clojure/Jolt Laya (Empirical as golden byte parity; +was empty skip §61):** +[`jlt-commons/laya-jolt`](https://github.com/jlt-commons/laya-jolt) +— Apache-2.0; README quickstart byte-identical to Python +`RLAgent.system_one`. ~1e-7 last-digit drift (f32 vs double). +~1.7 GB f32. Do not copy `jolt` (`notes.md` §70). +**CPU SemIf (Empirical as PoC UI, not a bench):** +[`leesk212/JEV-CPU`](https://github.com/leesk212/JEV-CPU) — +MIT; Qwen3-0.6B float32; option-letter logits, no generate. +Cross-ref semif-serve §69. **Meanblock/JEV-CPU 404.** +Softmax over slots ≠ Noul (`notes.md` §70). +**GLiNER2 decide adapter (spec only):** +[`Eran-BA/Jev_from_GLiNER2`](https://github.com/Eran-BA/Jev_from_GLiNER2) +— `fastino/gliner2-base-v1` → Choice/Score/Noul `/v1/systemone`. +No service, no training, no measurements. Interface ≠ replica. +Distinct from jeff GLiFormer (`notes.md` §70). +**Packed one-forward on an open LLM (constrained-AR / logprob path, +not a Jev reproduction):** +[`ikermoel/open-alternative-jev`](https://github.com/ikermoel/open-alternative-jev) +(Apache-2.0). RACE-H packed **92.9% @ 4.55 q/s** on Qwen3.6-27B +8-bit; interference 6–9%; temperature scaling on *your* labels. +Space demo. Economics of packing a shared state, not trained +decision-only (`notes.md` §49). +**CUDA/PyTorch local replica (constrained-AR / logprob cousin, not +a Jev reproduction, not Distillation; 2026-09-18 ~17:48):** +[`Mintzs/jevify`](https://github.com/Mintzs/jevify) — Qwen2.5-1.5B, +package `ora_decision_engine` / CLI `ora-decision`. CUDA graphs, +branch kernels, literal-label scoring. Default `--answer-encoding +letters`. **Uncalibrated model likelihoods, not measured +correctness.** Default refund `workflow.json` is not a validated +policy. **No LICENSE file this pass.** Do not copy Windows CUDA/venv +(`notes.md` §55). Independent of Distillation; independent of +open-alternative-jev's RACE-H receipt — same *class*, different +repo. +**Tiny SAN local surface (extreme speed/econ class, not a replica):** +[`wfzyx/von`](https://github.com/wfzyx/von) — 14 MB Needle; `POST +/v1/systemone`; sub-15 ms CPU *claim* / ~38 ms embed in their table; +authored144 needle **52.6%**. Distinguish from jev-local's **stub** +and kev's trained pointer. **Do not copy the vs-Jev ranking table.** +**GLiFormer encoder serving the System One wire (2026-09-18 ~20:43):** +[`logan-markewich/jeff`](https://github.com/logan-markewich/jeff) — +`knowledgator/gliformer-large-v1` 400M; official SDK `base_url` +drop-in; choice/score/noul. **Not a Jev replica.** Normalized +sigmoids, T=3.2; isolate nouls; DeBERTa tokens ≠ Jev billing. +Their card: L4 HTTP ~$2.6 vs ~$15.6 (~6×); A10G direct ~$0.65 +(~24×); AG News 75.5% vs 90.5%; CPU 6–20× *more* expensive. +License null this pass. Do not copy uv / Modal (`notes.md` §60). +**Loopback gateway, not a scorer:** +[`hraness/sysone`](https://github.com/hraness/sysone) — MIT; +routes hosted Jev + local OpenJev/NanoJev/Mini-Jev; does not +install weights; credential from env never config. Early; no +auth/streaming/non-loopback. +**Evaluation-model-first TS library (name collision; MED +2026-09-18 ~23:40):** +[`sysone-help/sysone`](https://github.com/sysone-help/sysone) — +MIT. `predicate` / `classifier` / `rubric` as pure data; +`check` / `evaluate` / `filter` / `partition` / `rank`; +cancellable; never auto-retry. First adapter = Jev via Vercel +AI Gateway. Independent of TypeSafe/Vercel. **Not** the +hraness loopback gateway. Do not copy npm (`notes.md` §63). +**Local MLX PCD vs Jev (Empirical as their n=50 table, +2026-09-18 ~21:39):** +[`mallahyari/system-one-benchmark`](https://github.com/mallahyari/system-one-benchmark) +— Qwen2.5-1.5B 4-bit MLX PCD is O(1) and faster on-device +(p50 227.2 ms) than Jev HTTPS (356.5 ms), and **uncalibrated** +(Brier 0.3884 vs Jev 0.1096; acc 52% vs 84.0%). Softmax over +allowed tokens ≠ Noul. License null. Small n. Do not copy +pip (`notes.md` §61). +**Productized Apple Silicon one-pass (Empirical as README +library, not a bake-off; 2026-09-19 ~01:47):** +[`bnsd55/jevmlx`](https://github.com/bnsd55/jevmlx) — MIT; +**28★**. Schema of booleans/enums/multi-selects → JSON +valid by construction, probability per field, one batched +MLX forward pass. Not TypeSafe Jev. Softmax ≠ Noul. No +local leaderboard yet (official 67.8% cited). OpenAI-compat +backend is one request per field. Distinct from +system-one-benchmark's n=50 Brier table. Do not copy pip +(`notes.md` §66). + +**When to use a decision model vs a constrained LLM (Harbor-style +bake-off, not a quality ranking).** +[`nibzard/decision-model-benchmark`](https://github.com/nibzard/decision-model-benchmark) +v2 (`notes.md` §49): no class wins on *accuracy*. Pick from axes +you actually need: + +| Need | Lean decision-model (Jev-class) | Lean constrained LLM | +|---|---|---| +| Latency / $ at schema-valid enums | p50 ~264–276 ms; S1 ~$0.07/1k; 0% malformed this run | Thinking-mode seconds and $0.19–$2.48/1k; some malformed | +| Choice sets ≤255 | Flat latency 2→255; **fails at 256+** (`400 Too many choices`) | Handles 512 | +| Honesty / abstain on no-good items | S5 admits 49.7%; ECE 0.246 — measure, do not assume | Most 97.3–100% (gpt-5.4-mini 64.7%) | +| Position stability | S4 flip 13% this run | Up to 37% | +| Mid-pack banking accuracy | 76.3% this protocol (atlas/jev-benchmarks 87%; jevals.com 79.67% — **do not merge**) | gpt-oss-120b 81.3%; glm-5.3 80.4% | + +Quality is not the reason to skip System One. Speed, cost, schema, +and the Choice cap are. Always re-run on *your* labels. Feedstock +for recomputing named boards: [`Jevals/jevals-data`](https://github.com/Jevals/jevals-data) +(CC-BY-4.0; `validation.md`). Reject: TypeAR or pcdServer scores as fail-closed P(permit); a LoRA student's -agreement with Jev as independent gold; shipping on "smarter than Jev"; +agreement with Jev as independent gold; kev's ID ECE as a license to skip +a held-out test on *your* workflow; shipping on "smarter than Jev"; +a green `/v1/systemone` smoke test on jev-local's stub as a bake-off; +copying laya-onnx or von vs-Jev rows as independent gold; treating von's +14 MB needle (52.6% authored144) or open-alternative-jev as a Jev +reproduction; thresholding [`jp-sns-jev7-estimator`](https://huggingface.co/kokuren/jp-sns-jev7-estimator) teacher scores as P(toxic) — the card says they are **not** calibrated, and `threat` F1@0.5 is 0.0000 on their table (`notes.md` §33). Domain-local @@ -421,7 +910,24 @@ universal ranking. Diffusion beating a decision head is Hypothesis. Detail: `research/notes.md` §33 (surfaces), §34 (marginals), §35 (Nimble), §36 (diffusion), §38 (entropy allocator, Hypothesis), -§42 (pcdServer serving, meta-VOI, games). Before +§42 (pcdServer serving, meta-VOI, games), §45 (kev), §46 (blackwood, +laya-bench, decision-token LoRA), §48 (jev-local stub, laya-onnx), +§64 (gqgs complete Laya ONNX; local-jev not equivalence), +§70 (grande / laya-jolt / JEV-CPU / local-jev measured / +GLiNER2 spec), +§75 (githubnext/localjev prompted JSON ≠ structured +logit read; ≠ kunchenguid/local-jev), +§76 (NandhaKishorM/laya packaging ≠ new species; +Router script-before-p; vs-Jev unpublished-here), +§77 (@airesearch12 census ≠ scored bake-off; +GLiNER2+routers class-boundary; incomplete vs watch; +≠ jevbench v1.1), +§78 (JevBench v1.2 geo-mean I/C/S/K; cal ON rank; +Jev 75.3 / SemIf 74.6 *theirs*; instruction models +in the table; Laya absent gap; ≠ v1.1 87.6), +§71 (openJev-verdict-2.0 competing NAR as claim-audit ≠ +IamBusy/OpenJev), +§49 (boundary map; DMB vs constrained LLMs; von; open-alternative-jev). Before adopting a surface, the bake-off is a jevals-shaped suite and, for a product loop, a Harbor taskset (`validation.md`, Eval & hill-climb). The stage pipeline into that decision is the same file @@ -442,10 +948,13 @@ the route or the patches. `notes.md` §36. ### "Perception specialist then judgment specialist" vs "Shared multimodal System One" -**Hypothesis.** Basit ask, primary post not retrieved (X and web, -2026-09-18; no tweet id). Useful *application* patterns for omni-ish -products. Not a SAM tutorial, not an ASR tutorial, and not a native -omni System One. +Two placements, not a slogan. Specialist composition stays **Hypothesis** +as a general recipe (Basit ask, primary post not retrieved; `notes.md` +§39). Shared multimodal *decide* now has a named open model: +[`BlackwoodAI/blackwood-rlcd`](https://huggingface.co/BlackwoodAI/blackwood-rlcd) +(**Empirical** as that vendor receipt, not re-run; **Hypothesis** on +*your* labels; `notes.md` §46). Not a SAM tutorial, not an ASR tutorial, +and not Archer's 27B drop (**Watch**). [SAM 3.1](https://huggingface.co/facebook/sam3.1) is Meta Segment Anything 3.1 ([release](https://github.com/facebookresearch/sam3/blob/main/RELEASE_SAM3p1.md), @@ -460,23 +969,28 @@ of transcript-then-Jev, not the source of this ask: (2026-09-17). His latency and price are his receipt, not a class number. -**Information dies at the interface.** The decision call sees the -schema you serialized, not the pixels or the waveform. Prefer a shared -multimodal decision model when that joint signal matters. Archer -Hume's drop stays **Watch** (no Hub weights as of 2026-09-18; -multimodal, no audio; do not promote to Empirical). djev-spark already -accepts images (multipart or JSON; think and sequential reject -images). A future audio-capable shared model is the same hole, not -this stack. +**Information dies at the act.** Specialist composition: the decision +call sees the schema you serialized, not the pixels or the waveform. +Shared multimodal decide (blackwood): the call sees the screenshot +*and* the candidates **code already marked** (letters on the image); +it still does not invent a click. Prefer specialist composition when +you already have the producer; prefer a shared multimodal decision +model when the joint of pixels and options matters — you no longer +have to wait for Archer to place that hole. Archer Hume's drop stays +**Watch** (no Hub weights as of 2026-09-18; multimodal, no audio). +djev-spark already accepts images (multipart or JSON; think and +sequential reject images) on a *different* compute graph. A future +audio-capable shared model is the same hole, not this stack. Same cut as a low/medium-entropy allocator: specialist state, then typed decisions. Buckets are product rhetoric, not a meter, and "review this PR" is still partly generative. One Jev call is still -factorized **marginals** (Meijer); the joint of the raw signal and the -decision lives outside the call. The handoff is a contract surface — -one sentence in `formal-methods.md`. Code owns the schema and the act. -`notes.md` §39. Measure and hill-climb that composition in -`validation.md` (**Hypothesis**, `notes.md` §41). +factorized **marginals** (Meijer); blackwood's one prefill is still +not a joint over the raw waveform. The handoff is a contract surface +— one sentence in `formal-methods.md`. Code owns the schema and the +act. `notes.md` §39, §46. Measure and hill-climb in `validation.md` +(`notes.md` §41). Jev still leads blackwood on general *text* (0.850 +vs 0.786 on their 8,456-item table) — omni is not a text free lunch. ### Open recipe (Bespoke Nimble) — not a distill @@ -499,7 +1013,7 @@ The GitHub repo has no license file — do not call the repo Apache-2.0. (302/324), untuned Qwen3.8-27B 84.88% (275/324), base Qwen3.5-9B 66.36% (215/324). **Empirical** as that named receipt, not a ranking. Synthetic labels, six source families, 162 pairs. -- **Bake-off candidate** beside Laya, openjev-lm, and TypeAR. +- **Bake-off candidate** beside Laya, openjev-lm, TypeAR, and kev. Adoption still requires the eval path (`validation.md`, Eval & hill-climb). Archer stays Watch. Not a jevals how-to. @@ -511,6 +1025,81 @@ the standing Choice-conditional-on-offered-set boundary (`SKILL.md`). Add "no match" when coverage is open. A high probability is not a correctness guarantee (their sentence). +### kev — runnable Archer reconstruction (not a distill) + +[jaredpalmer/kev](https://github.com/jaredpalmer/kev) (Apache-2.0; +README, MODEL_CARD, LICENSE, and release +[v0.1.0](https://github.com/jaredpalmer/kev/releases/tag/v0.1.0) HTTP +200). LoRA + pointer readout on Qwen2.5-0.5B. Shared state, isolated +questions under a block-causal mask, one prefill, no decode. +Architecture follows [Archer Hume's reconstruction](https://archerhume.com/posts/jevs-architecture-unmasked). +Speaks TypeSafe `POST /v1/systemone`; official `typesafe-sdk` works +with a `base_url` change. Weights `kev-0.5b` (38 MB) on that release +**and** on the Hub as [`jaredpalmer/kev-0.5b`](https://huggingface.co/jaredpalmer/kev-0.5b) +(`kev.publish`; `--run` accepts Hub ids; base still downloads on first +load). PEFT `task_type=FEATURE_EXTRACTION`; publish patches legacy +adapters. Not a how-to: do not copy serve flags, ports, or train +commands. **NOTA:** training must confront Choice `"other"` as a wrong +alternative too, with varied wording (`notes.md` §45 delta; +`question-design.md`). + +**Place it on the trained decision-only open path** next to Laya / +Nimble / Archer Watch. It is the cleanest *runnable* productization of +that reconstruction (API-compatible). Watch stays Watch: 27B, +multimodal, no Hub weights this pass. Not `Kevthetech143/super-jev`. + +**Contrast** (same I/O shape, different graph / objective / duty): + +- **vs TypeAR / pcdServer.** Constrained AR decode; softmax over + allowed tokens is not a Noul. kev does not decode. Use TypeAR when a + later field must see an earlier answer. +- **vs encoder open-jev (DeBERTa).** Bidirectional encoder, public + gold, OOD drop measured (acc 0.854→0.690). kev is a causal decoder + + pointer; OOD is unmeasured for this checkpoint. +- **vs proprietary Jev.** Documented cloud decision API, ~32k envelope, + in-dist ECE 0.0313 with OOD collapse. kev is laptop-local, 0.5B + knowledge, ID calibration only. +- **vs openjev-lm / jev-gate.** Those LoRAs copy a Jev *teacher*. kev + trains CE on public labelled outcomes. Same 0.5B backbone, different + gold. +- **vs Nimble.** 9B contrastive hard labels, no measured ECE, enum + ≤26. kev publishes ECE and isolation probes. + +**Evidence (README; not re-run).** Isolation exact: packed vs separate +max Δ 3.7e-6; secret-in-sibling p=0.03 vs in-state 0.99. Held-out ECE +0.065 (0.031 after temperature scaling); overall acc 0.799 on 1,350 ID +questions. Permute argmax flips 7.4%; IIA log-odds shift mean 0.13; +boundary forgery held. Honest limits: 0.5B knowledge; ID calibration +only; not multimodal. Model card: research prototype, not production, +not Jev. + +**When to use.** Laptop-local System One API drop-in for development +and eval. Not a knowledge or frontier substitute. **Bake-off +candidate** on the jevals/Harbor path (`validation.md`); mechanism +tests (isolation, permute, IIA, boundary forgery) mirror Archer probes +— they falsify the reconstruction, they do not prove kev = Jev. +`notes.md` §45. + +### blackwood-rlcd — open multimodal RLCD (not Archer, not CLIP) + +[`BlackwoodAI/blackwood-rlcd`](https://huggingface.co/BlackwoodAI/blackwood-rlcd) +(Hub README this pass; CC BY-NC 4.0; `image-text-to-text`). Trained +decision-only readout with **image-in**. One prefill; option-letter +logits; temperature-calibrated; Jev-compatible `/v1/systemone` shim. +Screenshot + candidates **code already marked** → typed Choice; code +clicks. Not Archer's 27B drop. Not a vision-scorer affinity. Not a +commercial drop-in. Do not copy serve flags. + +**Evidence (card; paired per item; not re-run).** Web element, 300 +held-out: acc **0.907** vs Jev 1.13 text-only **0.480**; letter-shuffle +flip **0.133** vs **0.587**; ECE **0.037** vs **0.091**; ~**200 ms** +1×H100 vs 441 ms OpenRouter. General text, 85 sets / 8,456 items: +**0.786** vs Jev **0.850** — Jev still leads. Screenshot rows compare +screenshot input with Jev's text input on the same steps. Randomized +viewport crop. **Empirical** as that named receipt. **Hypothesis** on +*your* labels. Bake-off candidate beside kev / Laya (`validation.md`). +`notes.md` §46. + ### Constrained-AR surface (not `decide`) [TypeAR](https://github.com/zmtomorrow/TypeAR) ("Type-Safe Decoding for @@ -539,7 +1128,8 @@ the isolation pattern. Enum width and no abstention primitive are brittleness — compose with cost-sensitive abstention (`mappings.md` §2), paraphrase abstain (`mappings.md` §17), and a jevgate-shaped gate (`mappings.md` §18). rh-guard's README (HTTP 200) is a coding-agent -reward-hack gate, a different surface from this one and from GLiGuard. +reward-hack gate, a different surface from this one, from GLiGuard, +and from Abide's project soft-rule Scores (`notes.md` §47). Do not copy the hook install. [pcdServer](https://github.com/stephanj/pcdServer) (MIT, C++20, created @@ -556,6 +1146,17 @@ install. Same author's earlier `parallelConstraintDecoding` is the two-forward-pass cousin already in the ecosystem snapshot. `notes.md` §42. +**Decision-token LoRA (same graph, trained).** If you specialize a +pretrained generator for this surface, train **the single decision +token** under parallel constrained decoding — not generated prose. +[`Foodoo1/Qwen3-14B-RLCD-Decision-LoRA`](https://huggingface.co/Foodoo1/Qwen3-14B-RLCD-Decision-LoRA) +(Apache-2.0 adapter on Qwen3-14B): loss only on that token; one +prefill + KV broadcast across fields. Their held-out 200-case / 4-field +receipt: fraud_risk **64.0% → 95.0%**, overall **85.2% → 98.8%** at +**~234 ms** / broadcast on RTX 3090 4-bit. Synthetic fraud-triage; +do not use for real financial decisions. Softmax over allowed tokens +is still not a Noul. `notes.md` §46. + **Hypothesis, not a stack.** TypeAR's example names the same Qwen3.8-27B family as Hume's announced weights. Running that surface (TypeAR or pcdServer) on those weights versus stock Qwen is a composition to test diff --git a/.agents/skills/augustus/references/mappings.md b/.agents/skills/augustus/references/mappings.md index a2fbadc..1a73191 100644 --- a/.agents/skills/augustus/references/mappings.md +++ b/.agents/skills/augustus/references/mappings.md @@ -42,11 +42,24 @@ code: shortlist, rank, filter — weights adjustable without re-inference ``` **Example**: research-reading map — extract reusability dimensions once, let -researchers re-rank and re-filter interactively. **Beyond SWE +researchers re-rank and re-filter interactively. +**Offline re-threshold (Empirical as named receipts):** +[testimonial-miner](https://github.com/AppitStudio/testimonial-miner) +`redecide` reapplies `Thresholds` to logged answers with **no new model +calls** — judge once, explore policy in code (`notes.md` §48). Same family +as firehose sliders. **Beyond SWE (Hypothesis until labeled):** vendor bid/no-bid (fit, urgency, risk Nouls; price and deadline exact); apartment shortlist (commute/light/ noise Scores; rent exact); hiring scorecard (evidence Nouls; labor-law vetoes in policy). Full gallery: `mental-models.md` §MCDA. +**Intent columns / catalog MCDA (Empirical as *shape*; clocks are +claims, 2026-09-19 ~00:38):** +[jevable.com](https://jevable.com/) class pattern: a heading +("Urgency") scores each row. Same hole as jevpandas / jevframe +(`mappings.md` §4) and dabit3 spreadsheet JUDGE/SCORE/CHOOSE. +Snack multi-criteria at catalog scale is a **maker claim** (3,000 / +28 s / $0.11) unless independently re-run. Weights and vetoes stay +in code (`notes.md` §56). **Counterexample** (from **Contract** Score docs): levels 0,1,2 with distributions `[0,1,0]` vs `[0.5,0,0.5]` both score 1.0 with radically different extreme-outcome risk — always read probabilities beside the @@ -60,7 +73,9 @@ Links: Score docs, composite-scoring pattern, autoresearch cookbook. choosing among act / decline / gather-evidence / escalate from the distribution, with thresholds owned by each action's consequences. For a calibrated binary probability with FP/FN costs: `t = C_FP/(C_FP+C_FN)`. -**Does not transfer**: universal thresholds (no magic 0.8); model probability +**Does not transfer**: universal thresholds (no magic 0.8 — a product +band such as Abide's ≥0.8 repair is *their* operating point, still +re-measured on your labels); model probability is not auto-calibrated on YOUR population — plot confidence vs accuracy on your data (**Contract**: confidence summarizes distribution shape, nothing more); Noul 0.5 ≠ medium-anything; top-Choice probability ≠ probability the @@ -72,14 +87,62 @@ elif action low-stakes: act else: act only if confidence > high bar, else confirm ``` +**Public pedagogy receipt (Empirical as article; 2026-09-19 ~10:25):** +[@akshay_pachaar “Jev Clearly Explained”](https://x.com/akshay_pachaar/status/2101037514945597645) +— high / medium / low confidence → auto / escalate / human; +thresholds in code; schema-safe ≠ correct. **200× / 400×** +are TypeSafe ceiling claims *theirs*, not Harbor. Shadow +first; questions-as-code. Do not copy the Python samples +(`notes.md` §85). + **Example**: trading bot acts on high-confidence reads, stands down when the -book state is ambiguous (jev-trader `late → hold`). **Beyond SWE +book state is ambiguous (jev-trader `late → hold`). **Banded fail-open +(Empirical as a named product receipt):** +[Abide](https://github.com/coldteadotai/abide) on project soft rules: +≥0.8 repair in-session, 0.5–0.8 human note, <0.5 silence; hooks exit 0; +no key → the edit proceeds (`notes.md` §47). Soft judgment is never the +sole hard veto. **Beyond SWE (Hypothesis until labeled):** inbox reply/snooze/archive; "is this paper on-question?"; "call this lead / nurture / drop" — same act/abstain/ gather table, costs written in hours or dollars, threshold per *action*. VOI: pay for the full PDF or the customer call only if expected decision change beats the cost (`mental-models.md` §VOI, §decision). -**Counterexample**: a flat Choice over three fine categories may still +**Dual-process cascade (Empirical as a productized metaphor, routing +accuracy unmeasured):** +[dual-process-ai](https://github.com/taro1985/dual-process-ai) — +`confidence ≥ τ` → S1 decides; else escalate to S2 (generate). Routing +fails open; safety fails closed. Keyword fallback without a key is not +equivalent S1. Tune τ on *your* escalation log (`notes.md` §49). +**Domain specialist vs few-shot hosted (Empirical as their +RESULTS.md, 2026-09-18 ~20:43):** +[Domain-jev-maker](https://github.com/help-er/Domain-jev-maker) — +independent CLINC-150 labels, **not** a Jev teacher-copy. +Matched-precision KL (both systems rounded to two decimals, +zeros → 0.0025): local 1.5B KL 0.168 vs hosted zero-shot 0.580 +banking (r +0.933 vs +0.343). Few-shot hosted (one example per +intent in `state`) matches or beats local determinate accuracy +(McNemar p=0.134 / p=1.000); calibration barely moves (KL still +2.5–3.8× higher). **Train the specialist when downstream code +reads the probability; use few-shot hosted when only argmax +matters.** Do not copy train how-to (`notes.md` §60). +**Active-learning triage / don't distill Jev as teacher +(Empirical as README architecture, 2026-09-18 ~21:39):** +[jev-triage](https://github.com/ThyFriendlyFox/jev-triage) — +high conf accept; middling expensive teacher; low or +near-boundary human. Logs full distributions to +`soft_labels.jsonl`. **Do not distill Jev as teacher of +record** — author ~68% ceiling compounds errors. Real +outcome labels remain the training targets. Noul belief +`|p−0.5|×2`; Choice top-two within 0.15 → human. +Distinguish Domain-jev-maker (independent gold specialist) +from openjev-lm (teacher-copy). Do not copy pip how-to +(`notes.md` §61). +**Harbor-shaped decide→policy leftover (Empirical as README +architecture):** [jav-email-cascade](https://github.com/skiingfalcon/jav-email-cascade) +— 8 typed questions; policy auto/review/llm; Noul 0.5 never +rounded; Score conf 0.0 never acted on. Mock gen-json +flat-confidence is *their mock*, not a live bake-off +(`notes.md` §60). **Counterexample**: a flat Choice over three fine categories may still name a harmless best pick — low confidence need not veto a low-stakes preference. **Test**: cost/coverage curve on held-out slices; score the fallback too (escalation is not automatically correct). Links: @@ -103,11 +166,57 @@ NOT `P(A∧B)`; operation+target head pairs can be invalid (browser-use asks both heads per request but executes only the matching target after validation); relational judgments ("does passage support claim?") must stay one question, not two split classifications; exact computation stays in code. +**Internals are not a state machine; placement is a component node** +(atlas: paraphrase vs reversed-meaning are LM understanding; "a node in +your state machine" is the architectural instinct — `notes.md` §49). **Example**: game director — Jev judges whether player dialogue is conciliatory or threatening; code enforces inventory, prerequisites, chronology, reachable scenes (cf. HEIST//ONE: six guards batched, simulation -validates every proposal). **Counterexample**: decomposing tool-trace +validates every proposal). **Merge-gate circuit (Empirical as README +behavior, 2026-09-18 ~16:48):** +[latch](https://github.com/CaseReed/latch) — Jev labels a clustered +cause; a **table** maps cause × confidence × fingerprint → PASS / +BLOCK / needs_human. The judge is a sensor, not the merge act +(`notes.md` §51). **Language primitive (Empirical as README / +example suite, 2026-09-18 ~17:48):** +[hunch](https://github.com/carldaws/hunch) — Ruby `chance` / +`pick` / `rate` map to Noul / Choice / Score; English is the +configuration; `Hunch.decide` batches over one `given:`. +Validations `rescue nil` = fail-open at save; spam gates should +fail closed. Stub backend for tests. Same interface ≠ same +guarantees for a future LLM backend. Cousin of probably-lang +(a language whose loop conditions are feelings) — this is a +library, not a new language. Do not copy gem/Rails +(`notes.md` §55). **Named circuit combinators (Empirical +as README architecture, 2026-09-19 ~01:47):** +[decision-combinators](https://github.com/voidning/decision-combinators) +— Then / Gate / Vote / Cascade / Weighted over +Choice/Score/Noul. README analogizes them as logic +gates; they are **not** literal Boolean AND/OR (those +aggregations stay in code — do not multiply parallel +Nouls). Vote is majority or mean; confidence discounted +by agreement. No measurements. GitHub SPDX null; package +MIT. **Hunch:** System One as a control plane, not chat +turns. Compose with skillranker. Do not copy npm +(`notes.md` §66). **Rename + extended five (2026-09-19 +~04:39):** now +[jev-combinators](https://github.com/voidning/jev-combinators) +(same `created_at`; npm `jev-combinators` 0.1.0). +Digital-design slogan: primitives are transistors, +combinators are logic gates, you design the chip. +**Extended:** Router / Loop / Retry / Fallback / Memory. +Fallback is the fail-closed node. Still not literal +AND/OR. `notes.md` §69. **TLA+ consensus circuit (Empirical as +spec + chaos table; 2026-09-19 ~02:38):** +[jev-labs](https://github.com/copyleftdev/jev-labs) +— five paraphrased agents; stability gate; quorum 3 of 5 +stable votes; escalate when budget spent. Code/TLA+ own +transitions (`Consulting → Decided | Escalated`). The +model never is the constraint. 1,080 golden: 0 wrong +*theirs*; underdetermined records still decided 34/120 +split both ways (stability ≠ answerability). MIT. +`notes.md` §67. **Counterexample**: decomposing tool-trace verification into per-call schema nouls works; asking "is the trace correct" as one Noul hides nine judgments. **Test**: full truth table / transition cases incl. contradictory outputs, stale observations, invalid combos. @@ -142,13 +251,106 @@ is not intrinsically wrong (offline, modest corpora) but it is not an index — per-query work still scales with candidates. Low latency ≠ no retrieval. **Store as the index (Empirical as a *shape*, 2026-09-18):** -[`kylemclaren/jevql`](https://github.com/kylemclaren/jevql) judges -schema-conditioned row objects; vanilla Postgres never sees `jev()`. -Cheap SQL first; the remainder is a typed Choice/Noul/Score over rows. -Row contents leave the database (same residency warning as AU health). +Cheap exact predicates first; typed questions on the remainder. Two +forks of the same hole: **in-engine extension** +([`mgaitan/sqlite-jev`](https://github.com/mgaitan/sqlite-jev), loadable +SQLite `jev_rows`; inspired by [`realZachi/pg-jev`](https://github.com/realZachi/pg-jev)) +vs **out-of-process CLI** ([`kylemclaren/jevql`](https://github.com/kylemclaren/jevql) +— **judgment outside the store**: vanilla Postgres never +sees `jev()`; the CLI/serve/MCP/SDKs judge). sqlite-jev is a semantic full +scan, not an index; `max_rows` is a spend guard; thresholds stay in SQL. [`ant4g0nist/joxide`](https://github.com/ant4g0nist/joxide): zoxide owns the directory index; Jev scores a shortlist; destinations are existing -local paths only; fail-open. `notes.md` §42. +local paths only; fail-open. Dataframe cousin this hour: +[`yalindogusahin/jevpandas`](https://github.com/yalindogusahin/jevpandas) +— `evaluate` / `filter` / `classify` / `score` / batched `ask` over a +pandas frame; classify example includes `other`; failures never become +negative predictions; LICENSE absent this pass. Accessor sibling this +hour: [`ktaletsk/jevframe`](https://github.com/ktaletsk/jevframe) (MIT, +PyPI; pandas **and** Polars `.jev`; full `p__` columns; no silent +renormalize; one row per request). Same hole, two surfaces. Row +contents leave the store (same residency warning as AU health). Do not +copy SQL, env, or CLI flags. +`notes.md` §42, §44, §46, §48. Intent-column / snack MCDA *shape*: +`mappings.md` §1; `notes.md` §56. + +**ORDER BY over probs is a ranking job, not a calibration +certificate (Empirical as independent measurement, 2026-09-18 +~20:43):** +[jev-orderby-bench](https://github.com/yodablocks/jev-orderby-bench) +— `jev-1.13.0` passes all six pre-registered gates on 360 +human-labeled 20 Newsgroups rows. Boolean inversion 0.036; +Score ordinal inversion **0.143** vs 0.15 (weak link / the +sort key); 53 rows tie at 0.99 so `LIMIT 20` is +engine-dependent; two-decimal quantization. Calibration +(ECE 0.0453 / Brier 0.0524) ≠ sortable. recodelabs default +40-row batching **fails** the ranking gate (inversion 0.171) +that one-row-per-request passes — request shape is part of +the measurement. Not a fourth DuckDB extension. Vendor 67.8% +agreement ≠ calibration. `udf.register()` refuses SQL unless +results pass. Do not copy curl / key how-to (`notes.md` §60). + +**Decision-native evidence set (Empirical as architecture; +Hypothesis as a measured win, 2026-09-18 ~17:48):** +[decision-native-rag-skills](https://github.com/emergency-lee/decision-native-rag-skills) +promotes this card from "rerank a shortlist" to **retrieve wide → +decide → build an evidence set → resolve conflicts → generate only +over kept evidence**. Embeddings remain candidate generators; they +do not settle relevance, sufficiency, redundancy, conflict, time, +or authority. No bundled harness; no universal benchmark; default +migration gates are starting targets (`notes.md` §55). PubMed +title/abstract screening is the same *shape* on literature +([typesafe-screening-mcp](https://github.com/masa-med-ai/typesafe-screening-mcp): +include/maybe/exclude in code; 326 hits ~17 s ~$0.014 one run; +thresholds not calibrated; screening aid, not an SR replacement). +Local-file cousin: +[kazuhideoki/jev-search](https://github.com/kazuhideoki/jev-search) +(recursive files + fzf) — **not** superagents-lab/jev-search +(federated web). Max-over-chunks ≠ calibrated whole-file p. +**Classify-first MCP (Empirical as README / schema, 2026-09-19 +~00:38):** +[jev-sift](https://github.com/kbhuw/jev-sift) is the same sandwich +on agent I/O: retrieve-wide (paths / public URLs / inline text) → +decide (relevance or 1–8 typed questions) → the main LLM opens +only the evidence set. Content never enters main agent context +first (paths/URLs). Hard envelope in code. Transport tests ≠ +accuracy. No LICENSE this pass. Cousin of typesafe-screening-mcp +(abstracts never enter the LLM conversation). Not jev-routing +(host adapter). `notes.md` §56. +**Meaning-search without embeddings (Empirical as a named +stripped-repo card, 2026-09-18 ~18:46):** +[jevgrep](https://github.com/Bentlybro/jevgrep) — packed parallel +Jev relevance; two-stage outline → zoom top 30; no index. 228 +questions, docstring-stripped repos: **79% top-5** vs BM25 40% / +grep 20%. Keyword still wins exact strings (BM25 top-10 96% vs +85%). Harbor-shaped: frozen copies + labeled questions + +comparable harnesses; not a Harbor taskset. Distinct from +kazuhideoki / superagents-lab / jev-sift (`notes.md` §58). +**Meaning-grep over line Nouls (Empirical as README + their +judge test, 2026-09-18 ~21:39; dedicated 2026-09-19 ~16:30):** +[jev-semgrep](https://github.com/uehaj/jev-semgrep) — AND/OR/NOT +on *thresholded* per-line Nouls (do not multiply p); +proposition ≠ embedding; contrast-set refund; no index; +cross-lingual; Semgrep.dev collision; **not a gate**. +Distinct from jevgrep (file/chunk) and jev-combinators +(metaphor). Precision 0.94 / recall 0.98 *theirs* (not +Harbor). **51★** ephemeral. LICENSE MIT (GitHub +NOASSERTION). `notes.md` §61, §86. +**Evidence-packet explorer (Empirical as their performance.md, +author-run):** +[jev-semantic-explorer](https://github.com/jimmyhealer/jev-semantic-explorer) +— index once, BM25 shortlist, Jev ranks, citable packet. +SWE-bench Verified n=8: 1/8 → 6/8 finish (empty = miss). +Packet n=50 HitFile 0.233 vs BM25 0.159 is **not** the +product KPI. Distinct from jevgrep / jev-sift. `notes.md` +§61. +**Measured RAG rerank vs generative rerank (Empirical as one-run; +Hypothesis as a transfer):** +[Jev-RAG](https://github.com/Max-sm-yc/Jev-RAG) — ≥70% cost / 72% +latency vs Muse Spark *rerank* on ~30k tokens (costs include +embeddings). Full-context Spark is still **faster** (10.60 s). +Do not overclaim vs no-RAG. License null this pass +(`notes.md` §58). ## 5. Hierarchy → bounded heuristic search @@ -251,6 +453,210 @@ prose → frontier. Same 149 business rows: Jev 79.9% vs Haiku 4.5 83.2%; Jev 1.6× faster, not 20–200×; Jev confidence monotonic, Haiku inverts in 0.80–0.95. If you do not *branch on confidence*, use whatever you already have (`notes.md` §42). +**Training-data VOI (Empirical as README architecture, +2026-09-18 ~21:39):** +[jev-triage](https://github.com/ThyFriendlyFox/jev-triage) — +pay for an expensive teacher or a human only where +confidence says the label will change the outcome. High-conf +accept is nearly free. Soft-label full distributions for a +local student; **real outcomes** stay the training targets. +Do not distill Jev as teacher of record (~68% ceiling). +Cost sketch *theirs*: ~$21 vs ~$8,400 LLM judge for 1M × +500-tok. `notes.md` §61. +**Retrieve-then-state (Empirical as an axis proof, not a knowledge +estimate):** if the answer is not in `state`, **buy the passage first**, +then ask. Atlas history suite: wrong @ 0.90 without context → right @ +0.97 with the passage (`notes.md` §49; `mental-models.md` §boundary). +That observation is VOI with a named receipt. Do not rely on bare +recall. +**Fail-open wake/resume (Empirical as README safety table; 21/21 is +smoke, 2026-09-18 ~16:48):** +[wakegate](https://github.com/shitianfang/wakegate) — skip a sleeping +agent's LLM turn only if Jev answers **and** p(wake) < 0.2; user +message / nothing-to-judge / skip-limit / error / unsure all **wake**. +Horvitz mixed-initiative: pay for the turn iff EV(decision) beats +the token cost. Savings unmeasured. Same-author scenarios+question; +not a benchmark (`notes.md` §51). Contrast pi-jev-approver +fail-closed without a key and jevgate cannot-block. +**Selective memory / scored recall (Empirical as README behavior; +9×3 is a hint, 2026-09-18 ~17:48):** +[carryforward](https://github.com/Dharundp6/jev-carryforward) — +verbatim ledger; Jev scores which facts are still live for the +task; constraints/corrections always return (never judged). Fail- +open dump if the scorer is down. Pay for a scored brief iff it +beats dumping the whole file. No accuracy claim until a proper +test (`notes.md` §55). Eval finding *theirs*: `recall` +**0/4** with tools available — SessionStart hook > +hoping. tools≠use (`notes.md` §68). Do not copy mcp add. +**Classify-first read (Empirical as README; Hypothesis as a +measured win, 2026-09-19 ~00:38):** +[jev-sift](https://github.com/kbhuw/jev-sift) — pay for a full +agent open iff the relevance (or typed question) says it might +change the act. Uncertain → closer look. Errors and truncation are +**not** evidence of irrelevance. Webpage fetch still costs +bandwidth; this saves the *agent's* read, not the download. +Mocks ≠ accuracy (`notes.md` §56). +**Decision-model latency cost (Empirical as a *negative* +on sync Jev; 2026-09-18 ~23:40):** +[slo-router](https://github.com/zeeshan8281/slo-router) — +pay for live Jev features on the routing hot path iff +expected decision quality beats **hundreds of ms** tail. +On their fixture, same routes/accuracy as local features; +p95 **77.93 → 490.38 ms**. Author: keep Jev off the +synchronous path for this workload. **Hunch:** Harbor-style +measurement of decision-model latency is mandatory before +claiming “Jev routing.” Eight-row demo is not a benchmark +(`notes.md` §63). +**Human-review VOI (Empirical as README architecture; +hunch as a placement):** +[jev-lens](https://github.com/rashedInt32/jev-lens) — +calibrated “do I need to look / which files / strip +debris?” Never blocks the agent; never says green unless +sure (`JEV_LENS_GREEN` 0.9). Minimize expected human cost +under **false-green** risk. Attention filter, not a +permission gate. Companion +[jev-lens.nvim](https://github.com/rashedInt32/jev-lens.nvim) +is display only. Distinct from jev-gates (stops writes) +and egma attention≠correctness (PR surface) +(`notes.md` §63). Distinct from +[dizk/jev-lens](https://github.com/dizk/jev-lens) +(pre-send views; 79% fewer tokens *theirs*; +`notes.md` §68). +**Pre-send token-econ (Empirical as 500-trajectory +bench; 2026-09-19 ~03:38):** +[jev-lens](https://github.com/dizk/jev-lens) — pay to +send a line iff it changes the next edit. Compress +**before** first send; post-send prune cost 17% more +because it broke the prompt cache. Code sent in full +unless Jev is confident. Harm: 2/26 later edits missed +their block. Distinct from rashedInt32/jev-lens. +Do not copy npm (`notes.md` §68). +**Skill-library VOI (Empirical as README architecture; +2026-09-19 ~01:47):** +[skillranker](https://github.com/Dicklesworthstone/skillranker) +— pay to load a skill iff it changes the next step. +Abstention ("none of these") is first-class. Failed hook +recommendation is quiet fail-open. Distinct from +skill-broker (grants). Compose with combinators +(`notes.md` §66). +**Same-intent cache admit (Empirical as n=100 live eval; +2026-09-19 ~04:39):** +[jevcache](https://github.com/kushals256/jevcache) — pay +for the LLM iff Jev says the intent is **not** the same. +Exact SHA-256 first; fail-open to upstream. 0 FP / recall +0.38 *theirs*. Stream/tools/multimodal bypass. Do not copy +npx (`notes.md` §69). +**Human-feed VOI (Empirical as unreviewed goldens; +qualify the owner; 2026-09-19 ~04:39):** +[ThinkyMiner/Winnow](https://github.com/ThinkyMiner/Winnow) +— pay for a click iff the card says read/skim. Distinct +from kevinpita/winnow (context sieve). 80%/90% *theirs*. +Do not copy unpacked-extension how-to (`notes.md` §69). +**Second-call VOI (Empirical as README + offline pytest; +license null; 2026-09-19 ~04:39):** +[hermes-jev-router](https://github.com/rsdkrasen/hermes-jev-router) +— pay for the *next* main-model call iff Jev says +generation is still required (WHETHER/HOW/WHAT). Skip-next +needs a core patch. Fail-open. Do not copy plugin how-to +(`notes.md` §69). +**Hunk-review VOI (Empirical as 22-run cost table; +2026-09-19 ~05:46):** +[prune-review](https://github.com/shubhangi013/prune-review) +— pay for generative review of a hunk iff Jev says it is +worth looking at (and the safety escarpment does not force +keep). Target ~20%; measured 1.18% with a 305% outlier +*theirs*. Cost not quality. Do not copy pnpm +(`notes.md` §70). +**Intent-search VOI (Empirical as CLI; under +construction):** +[jev-intent-review](https://github.com/yottayoshida/jev-intent-review) +— pay to judge a place the diff did not touch iff the +stated intent applies there. UNKNOWN is cheaper than a +false VERIFIED. Empty search ≠ proof (`notes.md` §70). +**Empty compact-proxy skip (description only; 2026-09-19 +~06:43):** +[jev-context-pruner](https://github.com/IPECTER/jev-context-pruner) +— Codex compression-proxy slogan; repo empty. Not VOI +until there is a keep-set and a fail polarity +(`notes.md` §71). +**Batch ranking VOI (Empirical as README; 2026-09-19 +~06:43):** +[jevfeed](https://github.com/fengyiqicoder/jevfeed) +— one Jev request per batch of ten; the distribution *is* +the ranking. Pay per *batch*, not per item (`notes.md` +§71). +**No-text-step VOI (Empirical as 95-call card):** +[jev-use](https://github.com/shitianfang/jev-use) +— pay the LLM only when writing is the job; 12 questions +in one call 186 vs 2,672 ms *theirs* (`notes.md` §71). +**Files-to-read VOI (Empirical as n=16 SWE; 2026-09-19 +~07:49):** +[jevex](https://github.com/jimmyhealer/jevex) — pay for +Reads of cited ranges only. n=16 160s → 69s / $8.74 → +$3.13 / 16/16 both arms *theirs*. Keep n=8 finish 1/8 → +6/8. Rename of jev-semantic-explorer (`notes.md` §72). +**Commit-attention VOI (Empirical as 13 labelled):** +[commitjev](https://github.com/yodablocks/commitjev) — +pay a human iff a Noul clears 0.65 on the bad side; +middle band is review not a skip. Regex already settled +the literals (`notes.md` §72). +**Pi compact VOI (Empirical as latency table):** +[pi-jev-compact](https://github.com/dev-willbird1936/pi-jev-compact) +— pay Jev to keep/drop tool calls instead of an LLM +summary; fall back if savings <25%. 0.6 s replay vs 26 s +first spinner is host cost (`notes.md` §72). +**Empty compact-proxy skip (IPECTER runway too):** +[jev-runway](https://github.com/IPECTER/jev-runway) — +LICENSE-only Codex-proxy slogan; not VOI until there is +a keep-set (`notes.md` §72). +**Escalate-under-threshold VOI (Empirical as README; +life/business; 2026-09-19 ~08:37):** +[classifier-dev](https://github.com/mrmps/classifier-dev) +— pay a reasoning model **only** on single-label +answers below 0.7. Multi-label re-judge made it worse +(23 s) so the tier is ignored. gemini-3.8-flash helped; +other flashes did not. 0.7 is *theirs*. Cousin jev-use +(`notes.md` §73). +**Escalate-without-stall cousin (Empirical as README +delta; autonomy; 2026-09-19 ~09:50):** +[khordoo/jev-reflex-autonomy-lab](https://github.com/khordoo/jev-reflex-autonomy-lab) +— pay S2 only under the confidence threshold, but +**never pause the reflex**. Log whether the returned +strategy was **consumed** (purple confidence), not +only that it arrived. 20% starting gate *theirs* +still soft. Local vs Live is an A/B of backends, not +a scored bake-off (`notes.md` §80). +**Split-question CU VOI (Empirical as README; desktop; +2026-09-19 ~09:51):** +[typesafe-computer-use](https://github.com/awlevin/typesafe-computer-use) +— pay three/four Choices in **one** request +(`kind`/`item`/`site`/`offscreen`) instead of one +255-way soup. Perception (crop+tile OCR, dates.py, AX +walk) is the expensive gather that rebuilds what +frontier reads from pixels for free. Exclusive +actions: overlap is loud doubt, not silent 1.00. +155× / $0.0002 is *theirs* on one screenshot, not a +taskset. Writer only when free text is the job +(`notes.md` §81). +**Partial-speech VOI (Empirical as README; voice; +2026-09-19 ~10:01):** +[jev-voice-browser](https://github.com/moritzkremb/jev-voice-browser) +— pay 9–11 questions on every partial (~300 ms +*theirs*). Closed-set may act before the sentence +ends; free-text waits for final or 600 ms silence so +"search for alan" is not truncated. Numbered overlay +is cheaper than a second model. Spoken confirm is +not a gather. 27/27 fixtures *theirs*, not a taskset +(`notes.md` §82). +**Evidence-synthesis two-pass VOI (Empirical as README; +medicine/Cochrane; 2026-09-19 ~08:48):** +[choxos/jev-reviewer](https://github.com/choxos/jev-reviewer) +— fan-out every question over shared chunks (18-q +**4.6 s / $0.0101** *theirs*), then pay a second +**absolute** Noul only on the surviving lines. *Not +found* is cheaper than a paraphrase. Human tick is the +act that enters the review. **≠** egma-ai +(`notes.md` §74). **Beyond SWE (Hypothesis until you log act/outcome pairs):** full PDF vs abstract; customer call vs CRM fields that already fail a hard rule (credit limit is exact); blood test vs @@ -282,6 +688,21 @@ report hits / false alarms at the operating point, not accuracy **Example (Empirical as family shape):** firehose / Near Here moderation — judge once, re-policy in code. +**CI merge-gate (Empirical as README / demo, 2026-09-18 ~16:48):** +[latch](https://github.com/CaseReed/latch) — false PASS on a real bug +>> false BLOCK on infra; criterion lives in the policy table, not in +the cause label (`notes.md` §51). +**Physical-world criterion (Empirical as README + +measurements; 2026-09-19 ~03:38):** +[HA-Jev](https://github.com/AboveColin/HA-Jev) — +confidence gating on typed sensors; **not** for locks / +heaters / smoke. `background:` on the question triples +laundry separation *theirs*. Treat 0.9 as higher than +0.6, not as right nine times in ten (`notes.md` §68). +**Stop-hook attention (Empirical as owner-run smoke):** +[jev-preflight](https://github.com/muse0509/jev-preflight) +— 0.85 uncalibrated; fail-open; one reinspect. Criterion +for *redirect*, not for *block* (`notes.md` §68). [`jp-sns-jev7-estimator`](https://huggingface.co/kokuren/jp-sns-jev7-estimator) is the rare-class warning in one table: seven distilled teacher scores that the card says are **not** calibrated probabilities, and `threat` @@ -294,6 +715,168 @@ misses" — that was a criterion shift. **Test**: ROC/PR on held-out *your* cases; report the operating point you actually ship. **Hypothesis** for non-SWE plots. Links: `mental-models.md` §SDT; evaluator script for threshold/cost sweep. +**Operator owns the criterion (Empirical as measured OMP +suppression, 2026-09-18 ~22:38):** +[omp-greenlight](https://github.com/SemetricLabs/omp-greenlight) +— four presets, each measured for prompts removed *and* +unsafe auto-approvals. Default **40.9%** / **0 of 94** on +the 140-row corpus. The plugin **never self-tunes** the +safety bar: a self-adjusting bar cannot be audited by the +person accepting the risk. Live traffic has no labels. +Not a sandbox. `notes.md` §62. +**Exactness raises a floor, does not override capability +(Empirical as live analysis; 2026-09-18 ~23:40):** +[slo-router](https://github.com/zeeshan8281/slo-router) — +the exactness feature lifts the quality floor; it never +bypasses health/context/tool checks. 3/8 task-label +disagreements still did not change routes. **Hunch:** +signal-detection framing — a quality cue is a criterion +shift, not a capability override (`notes.md` §63). +**Privilege ≠ verdict (Empirical as certification):** +[construct-auto-classifier](https://github.com/godspede/construct-auto-classifier) +— `sudo` changes blast radius, not whether the act is +benign. Operator owns `minConfidence` / `riskThreshold`. +Jev 0 dangerous / 975; chat models leaked. **Hunch:** do +not threshold a privilege token as P(unsafe) +(`notes.md` §63). +**Ranking ≠ calibration (Empirical as 8,000-judgment +audit; 2026-09-19 ~00:39):** +[does-jev-confidence-mean-anything](https://github.com/Adilmp/does-jev-confidence-mean-anything) +— AUC **~0.91** (ranking works) while stated p is shifted +toward "yes": when Jev said **~75%**, humans flagged +**~10%**. Two-parameter recalibration removes **~96% of +ECE** without changing rank. Vendor "calibrated" here is +**rank-correlation**, not frequency units. Never +hard-threshold raw p as if it were P(event) without +**domain** recalibration (`jevcal`; ~100 labelled rows). +One dataset (`civil_comments`); do not cite `threat` (n=1). +License null. `notes.md` §64. +**OOD / AUC ≠ ECE (Empirical as 900-ticket + 3 public +benches; 2026-09-19 ~01:47):** +[jev-ood-calibration](https://github.com/scienthoon/jev-ood-calibration) +— in-domain public benches look almost honest (OpenBookQA +ECE 0.024 / T 0.96). On an **unknowable** org-policy +priority label absent from the text: 44.7% acc, mean +stated p **0.74**, ECE 0.325, refit T **3.40**. Sign +**flips by type** on the same tickets: Choice/Score +overconfident (T ~3.3), boolean underconfident (T 0.66). +Do not threshold the TypeSafe `confidence` field (worse +than max-p here). Complements does-jev-confidence +(in-domain humans) and dinostomp (instrument). +Gateway exposes no model version. `notes.md` §66. +**Sureness over the vector (Empirical as 60-q reverse- +engineer + library; 2026-09-19 ~02:38):** +[how-sure-is-jev](https://github.com/adarc8/how-sure-is-jev) +— max_prob / margin / entropy / gini / perplexity (plus +Score `spread` and Kass–Raftery `log_odds`) → one +`[0,1]` consensus and bands CERTAIN | CONFIDENT | +LEANING | TORN | CLUELESS. Choice `confidence == +max_prob` to 3 decimals *theirs*; max_prob is the +**most generous** metric (75/25 → 0.5 vs entropy 0.19). +Thresholds are opinions. Pair with this OOD card: do not +threshold TypeSafe `confidence`. Zero-dep MIT. +`notes.md` §67. +**Conflict ≠ ignorance (Empirical as NCML field note v0.3 +*theirs*; 2026-09-19 ~04:39):** +[jev-typed-evaluation-collapse](https://github.com/mleyvaz/jev-typed-evaluation-collapse) +— same evidence, three schemas. Noul/boolean collapses +conflict 0.50–0.57 vs ignorance 0.46–0.48; Choice with +named `conflicting_evidence` / `insufficient_evidence` +separates p=1.0; binary Choice without an escape is +lexically biased (red 0.67–0.85). Score exploratory +(severe conflict 2.04 vs no-evidence 3.95). Schema is +the interface. License null. `notes.md` §69. +**BBQ stereotype / uncertainty as SDT (Empirical as +full 58,492 *theirs*; 2026-09-19 ~05:46):** +[jev-bbq-experiment](https://github.com/simonmesmith/jev-bbq-experiment) +— Jev 1.13.0 **56,900 / 97.28%**; amb 99.96% / inf +94.60%; BBQ bias **0.04 / 0.34**; **$0.3429 / 7.75 min**. +12 of 13 ambiguous errors stereotype-aligned; +informative misses mostly `unknown` (1,487 / 1,579). +Order diagnostic 1/484 (0.21%). Dataset CC BY 4.0 BBQ. +License null. **Not a general bias cert** — English/U.S. +QA template does not certify hiring/lending/healthcare. +Always-unknown would score 50%; 97.28% is not abstention +theater. Pair with +[system-one-responsible-ai](https://github.com/david-j-lustig/system-one-responsible-ai) +(size-0 framing stub). `notes.md` §70. +**Cookbook moderation as a cost-sensitive dial (Empirical +as small samples, not a bench; 2026-09-19 ~06:43):** +[jev-cookbook](https://github.com/nexibeo/jev-cookbook) +recipe 12 — five hazard Nouls; act when sure, hold the +middle, escalate self-harm early. 16–36 handmade items; +authors say not benchmarks (`notes.md` §71). +**Dual-channel ECE / like-for-like (claim-audit, not +endorsement; 2026-09-19 ~06:43):** +[openJev-verdict-2.0](https://github.com/Heman10x-NGU/openJev-verdict-2.0) +— correctness-head ECE is not distribution ECE. Open PR +#1: like-for-like dist 15.13% vs Laya 21.40%; Jev 14.40% +slightly lower on that channel; do not put 1.44% beside +21.40% as a 15× win. Throughput ≠ latency. **≠** +IamBusy/OpenJev (`notes.md` §71). +**Commit middle band (Empirical as 13 labelled; +2026-09-19 ~07:49):** +[commitjev](https://github.com/yodablocks/commitjev) — +the operating point is a **three-way** criterion +(pass / review / warn), not a rounded yes. 0.65 is +theirs, not a universal t. Five clean is a small +control (`notes.md` §72). +**English-checkpoint confident-wrong OOD (Empirical as +MASSIVE; 2026-09-19 ~07:49):** +[laya-multilingual](https://huggingface.co/convaiinnovations/laya-multilingual) +— Khmer 0.000 acc at 0.952 confidence; mean conf never +< 0.885. **Gating cannot catch it.** Route by script +before the forward pass. Ships uncalibrated +(`notes.md` §72). +**Coverage ≠ correctness (Empirical as GUI 336):** +[chakuho](https://github.com/taku-me/chakuho) — 8B +coverage 1.00 while `__none__` hits 3/30. Coverage is +format-mass (`notes.md` §72). +**Peaked schema-scorer (Empirical as Hub eval):** +Hub schema-scorer v2 Choice 0.841 *theirs*; treat p as +ranking. GitHub 404 (`notes.md` §72). +**Escalate-under-threshold criterion (Empirical as +README; 2026-09-19 ~08:37):** +[classifier-dev](https://github.com/mrmps/classifier-dev) +— 0.7 is an operating point on *their* labels (emotion +≥0.9 → 82% / <0.5 → 29% *theirs*). Multi-label does +**not** share it. Do not copy 0.7. `notes.md` §73. +**Self-reported JSON ≠ calibrated Noul (Empirical as +eval caveats; 2026-09-19 ~08:56):** +[githubnext/localjev](https://github.com/githubnext/localjev) +— entropy confidence is computed from a generated +vector. Bake-off: do not treat outputs as calibrated +(wrong-BoolQ high conf → large NLL; 40 samples/task +*theirs*). Wire-compat ≠ logit-equiv. `notes.md` §75. +**Laya 0.85 still soft / post-T ≠ raw ECE (Empirical +as README; 2026-09-19 ~09:07):** +[NandhaKishorM/laya](https://github.com/NandhaKishorM/laya) +— README `conf >= 0.85` is *theirs*, not Harbor- +calibrated. Khmer 0.000@0.952. vs-Jev ECE **0.081** is +post-temperature (raw 0.213 vs Jev 0.144). Banking77 +token-budget, not a Jev loss. `notes.md` §76. +**External census ≠ scored bake-off (Empirical as +tweet; 2026-09-19 ~09:14):** +[@airesearch12](https://x.com/airesearch12/status/2101259522933186879) +— named ~18 openjevs + first leaderboard promised +"today." **≠** jevbench v1.1. Watch +[jev-models](https://benchmarkheaven.com/jev-models); +do not paste live ranks here. GLiNER2 and routers on +the list are **class-boundary**, not identity. Likes +ephemeral. Incomplete vs Laya/localjev/kev is lag. +Harbor still wants cal / cost / latency / silent +fallback named. `notes.md` §77. +**JevBench v1.2 scored board (Empirical as board + +RESULTS; 2026-09-19 ~09:24):** +[jev-models](https://benchmarkheaven.com/jev-models) +protocol `jevbench::v1.2`. Score = geometric mean of +I/C/S/K at 25% each. Jev 1.13.0 **75.3** / SemIf +**74.6** *theirs*. Calibration **on** the rank (delta +from §67). Luna I=96.8 rank #7. Self-host latency +×2 is an assumption; many costs est. Option-order +72%→21%. Laya absent (gap, not named-excluded). +Qwen3.8 27B Chutes TEE **≠** Archer. **≠** tweet +census **≠** v1.1 87.6. `notes.md` §78. ## 8. Control structure → sensor ≠ constraint (Leveson) @@ -325,6 +908,144 @@ the constraint that remains when the sensor is deleted; fill the unsafe- control-action table. Ownership split is **Contract** as a rule (`formal-methods.md`); the domain examples are **Hypothesis** until labeled. Links: `mental-models.md` §Leveson; `boundary-audit.md` TOCTOU. +**Capability kernel (Empirical as README architecture, 2026-09-18 +~19:48):** [interlock](https://github.com/somoore/interlock) — LLM +ring 3; kernel ring 0; secrets never enter the agent; closed action +space; Jev (or stand-in) is the sensor; `policy.py` decides +BLOCK/ASK/ALLOW. Type-safe ≠ correct; irreversible behind a +threshold **and** a human. Anti-pattern: launch-week firewalls that +ask "dangerous?" after the LLM already decided with real secrets in +scope. Distinct from toolgate (pre-exec of a proposed call). 38-case +set tunes the local judge, not a blind paper. `notes.md` §59. +**Host deny stays above the sensor (Empirical as README +safety model, 2026-09-18 ~22:38):** +[omp-greenlight](https://github.com/SemetricLabs/omp-greenlight) +— `bash.patterns: deny` is the constraint and fires ahead +of Jev prompt-suppression. Jev is permission-*probability*, +not permission. Agent prose withheld after 0→3 corpus +misses. Never shadows a built-in tool (would bypass deny). +`notes.md` §62. +**Spoken confirm ≠ constraint (Empirical as README; +2026-09-19 ~10:01):** +[jev-voice-browser](https://github.com/moritzkremb/jev-voice-browser) +— destructive Noul is a sensor; spoken "confirm" is +not an interlock. Control-port reach grants. rh-guard +owns the gate cousin (`notes.md` §82). +**Wrap-as-execution is the constraint (Empirical as +README; 2026-09-19 ~10:20):** +[AgentGhost](https://github.com/reddpy/AgentGhost) +— the wrap *is* the actuator path; Jev is the sensor +on leftovers after rules. ASK throws; fail-closed on +judge error. `AUTO_APPROVE` is not a constraint. +rh-guard owns the gate cousin (`notes.md` §83). +**Judgment ≠ permission (Hypothesis / outline only):** +[skill-broker](https://github.com/adamjralph/skill-broker) +— code owns grants; Jev scores relevance and **never +grants access**. Jev down never broadens the catalog. +Not a production recipe. `notes.md` §62. +**Contracts on effects, not tokens (Empirical as +certification; hunch as FM angle):** +[construct-auto-classifier](https://github.com/godspede/construct-auto-classifier) +— the constraint is reversibility / blast radius, not a +`sudo` allowlist. Independent risk Nouls are sensors; +policy (minConfidence ∩ riskThreshold ∩ fast-deny) is the +constraint. Fail-closed when the sensor is missing. +**Landed-script** is a merge-gate receipt, not a name. +**Headless** escalation is deny-and-report, not +auto-approve (`notes.md` §63, §68). +**Attention filter ≠ permission (Empirical as README; +hunch as placement):** +[jev-lens](https://github.com/rashedInt32/jev-lens) — +never blocks the agent; never grants or withholds a +write. Complements skill-broker (Jev never grants access) +and omp-greenlight (operator owns the bar) +(`notes.md` §63). +**Jev supplies evidence, code owns authority (Empirical +as README slogan; 2026-09-19 ~00:39):** +[actiongate-jev](https://github.com/omkarghugarkar007/actiongate-jev) +— deterministic policy is the constraint; Jev is the +sensor. A positive score never overrides RBAC / schema / +limit failure. Financial / destructive / credential fail +closed if Jev is down. Distinct from construct (shell +effects) and interlock (secrets never in agent) +(`notes.md` §64). +**Wrap-as-execution cousin (Empirical as README; +2026-09-19 ~10:20):** +[AgentGhost](https://github.com/reddpy/AgentGhost) +— the wrap *is* execution; rules prove allow/deny/ask +before Jev; ASK throws. Distinct from actiongate +(policy/RBAC is the hard gate, Jev only evidence). +`notes.md` §83. +**Turnstile clone (Empirical as README architecture; +2026-09-19 ~01:47):** +[turnstile](https://github.com/zyphr-labs/turnstile) — +same doctrine (policy first; Jev remainder; evidence ≠ +authority) with receipts and **threshold replay**. Missing +Jev → Review, not a silent allow. Starting 0.85/0.35 are +not calibrated. Experimental alpha. `notes.md` §66. +**Advance gate / coverage ledger (Empirical as README + +BEYOND-JEV.md; 2026-09-19 ~02:38):** +[seal](https://github.com/Reasonofmoon/seal) +— sensor (Strike / Jev / code) ≠ constraint (Seal + +coverage.path). Exception queue must be visible. Mint ≠ +product brain. Effects locked while escalations open. +`notes.md` §67. +**Never confidently wrong (Empirical as TLA+ + chaos +table):** [jev-labs](https://github.com/copyleftdev/jev-labs) +— the constraint is "may escalate; must not return a +confident wrong." Hard-gating without that path is +soundness theater's inverse. Synthetic pharmacy, not +clinical. `notes.md` §67. +**Skill-broker sibling (Hypothesis / outline; delta +§67):** grants stay in code beside turnstile (runtime) +and skillranker (advisory). Same doctrine, different +hole. `notes.md` §62, §67. +**Conversational constraint sensor (Empirical as +79-session bench; 2026-09-19 ~05:46):** +[pi-heed](https://github.com/Nyarlathoteppppp/pi-heed) +— user constraints persist as structured state across +compaction; replayed **without** calling Jev again. +Jev classifies KEEP/LIFT/…; **never writes policy**. +Side-effecting calls checked before they run. Fail-open. +Shadow default. *Theirs:* recall **98.5%** / false +block **0.0%** / lifecycle 100% / task success 98.7% / +**$0.000058**; mid-session rule change 8/13 off vs +0/13 on. Distinct from actiongate (RBAC/schema +authority) — this is *what the user meant* surviving +the context window. Do not copy `pi install` +(`notes.md` §70). +**Pi control-plane sensors (Empirical as README; license +null; 2026-09-19 ~06:43):** +[pi-jev-control](https://github.com/goodruizhan/pi-jev-control) +— router / tool gate / retry / sieve / review / GUI are +named sensors; code owns model switch, session bytes, +and click. GUI never force-clicks. Distinct from pi-heed +(constraint ledger) (`notes.md` §71). +**jev-use gate never grants (Empirical as 12/12 +fail-open):** +[jev-use](https://github.com/shitianfang/jev-use) +— PreToolUse deny/ask; missing Jev does not deny +(`notes.md` §71). +**Inbox policy is the constraint (Empirical as README; +2026-09-19 ~07:49):** +[mailordinal](https://github.com/Milo318/mailordinal) +— typed signals are sensors; the 100-point policy and +the review lane are constraints. The model never sets +queue order. Humans own ambiguity (`notes.md` §72). +**Agnes branded as Jev is not a sensor (identity lock):** +[hermes-plugin-jev](https://github.com/Mrmimee/hermes-plugin-jev) +— chat-completions path wearing Choice/Noul/Score +vocabulary. Distinct from hermes-jev-router +(`notes.md` §72). +**Silent fallback is an unsafe control action +(Empirical as eval/README; 2026-09-19 ~08:37):** +[classifier-dev](https://github.com/mrmps/classifier-dev) +— delisted primary left granite serving F1 **0.546** +vs advertised ~**0.800** for weeks (*theirs*). Digest +now marks `FALLBACK`. The constraint is honesty about +which model answered, not a better softmax. +**rh-guard owns the eval-integrity gate**; this is the +lived product cousin (`notes.md` §73). ## 9. Search / control loops → one substituted classifier step @@ -336,10 +1057,45 @@ mapping §5 is the taxonomy-beam special case). Economics inversion: per-node judgments were known and too expensive; they are now default. Control: hysteresis, continue / stop / retry / verify — the model estimates named probabilities; the controller is a table with memory. -**Does not transfer**: Jev as the planner that picks its next tool in a -loop; bandits without observed rewards; speculative depth without a -simulator; PufferLib Ocean scores as a capability claim -(`formal-methods.md` DST trio). +**Bounded Pi supervisor (Empirical as README policy, 2026-09-18 +~16:48):** [jevons](https://github.com/LilDojd/jevons) — not a second +agent; Jev interprets evidence; code owns freshness/limits; default +recovery **shadow**; steering never generates commands +(`notes.md` §51). Distinguish from pi-jev-approver / pi-jev-context. +**Does not transfer**: Jev as the planner-writer that picks its next +tool *and writes* the call; bandits without observed rewards; +speculative depth without a simulator; PufferLib Ocean scores as a +capability claim (`formal-methods.md` DST trio). **Does transfer as a +split** ([jeffrey](https://github.com/thomasbrueggemann/jeffrey)): Jev +owns next-tool / progress / risk / done; the LLM **only fills args**; +the loop is Jev→tool→Jev. Pick ≠ fill. Mapping §9 still rejects the +fused planner. +**Continuous-control cousin (Empirical as README delta; +2026-09-19 ~09:50):** +[khordoo/jev-reflex-autonomy-lab](https://github.com/khordoo/jev-reflex-autonomy-lab) +— the *algorithm* is the physics loop; the substituted +classifier step is typed flight Choice every tick. +Optional S2 is one-use strategy, not the next act. +Escalate **without stalling**. Local rule-based vs +Live `jev-latest` is an A/B of backends (**≠** +githubnext/localjev). Seed = geometry ≠ replay. No +pixels. 20% still soft. S2 never grants. Do not copy +npm / `.dev.vars` (`notes.md` §46, §80). +**OCR+AX desktop cousin (Empirical as README; +2026-09-19 ~09:51):** +[typesafe-computer-use](https://github.com/awlevin/typesafe-computer-use) +— the *algorithm* is the capture→OCR+AX→Choice→act +loop; the substituted classifier step is exclusive +kind/item/site. Perception stays in code. Writer is +leftover generation. Decision never ships pixels; the +answer reader may. Do not copy `uv` (`notes.md` §81). +**ASR voice-browser cousin (Empirical as README; +2026-09-19 ~10:01):** +[jev-voice-browser](https://github.com/moritzkremb/jev-voice-browser) +— the *algorithm* is debounce→snapshot→Choice→Playwright; +the substituted classifier is 9–11 questions on a +partial transcript. ASR stays off-model. Confirm is +not a grant. Do not copy `npm` (`notes.md` §82). ```text loop = yours (beam / funnel / stages / MCTS / incident command) @@ -349,7 +1105,34 @@ estimate ≠ measure — irreversible milestones concede only to the probe ``` **Example (Empirical):** jev-mcts grounded vs speculative fidelity in -types; probes-only concession (mapping §5). **Beyond SWE (Hypothesis):** +types; probes-only concession (mapping §5). +**Decider ≠ executor (Empirical as README; 2026-09-19 ~05:46):** +[jeffrey](https://github.com/thomasbrueggemann/jeffrey) — the +*algorithm* is the agent loop; the substituted classifier step is +next-tool / progress / risk / done. Arg fill is generation, not the +classifier. Risk Score ≥ 0.5 pauses mutating tools. Stuck ladder: +withhold the looping tool, re-ask Jev (2 Jev / 0 steps). Distinct +from jev-handoff (typed baton around an existing host) and +browser-jev (Playwright executes). Do not copy npm (`notes.md` §70). +**Tree-of-Choices writer (Empirical as README demo; +2026-09-19 ~06:43):** +[jev-gpt](https://github.com/florian-hoenicke/jev-gpt) +— the substituted classifier step is *which word next*; +the model never free-generates. ~400 calls / 75 s / 2¢ +*theirs*. Architecture demo. Distinct from jeffrey (pick +next-tool). License null (`notes.md` §71). +**Pick≠write plugin (Empirical as 95-call card):** +[jev-use](https://github.com/shitianfang/jev-use) +— judgment steps to Jev; writing stays generated +(`notes.md` §71). +**1-token selector (Empirical as GUI 336 + mario; +2026-09-19 ~07:49):** +[chakuho](https://github.com/taku-me/chakuho) +— the substituted classifier step is *which declared +label*; the generic LLM never writes. Numeric rules stay +in code (mario loss is misapplied "< 6 tiles"). Softmax +≠ Noul (`notes.md` §72). +**Beyond SWE (Hypothesis):** snowball citations ("still on-question?"); sales stages ("still a real opp?" — amount and close date stay exact); cook/rest/check ("looks done?" — thermometer is the probe). **Counterexample**: a weekly LLM summary of @@ -368,6 +1151,67 @@ immediate win missed once reversed; Fool's-mate confidence 31%/37% so a hole on the constrained-AR surface. Not a strength rating. `notes.md` §42; `validation.md`. +**Computer-use observe → score → act (Empirical as README / +architecture behavior, 2026-09-18 ~16:56):** +[gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) +— the *algorithm* is the browser loop; the substituted classifier +step is scoring among observed a11y/DOM controls. Local GLiNER2 is +one backend; Jev Ultrafast / solari-reflex are the Jev backends of +the same hole. Code owns actuators, dates, freshness. `DONE` is not +the probe — application verifiers are. Contrast blackwood-rlcd +(screenshot input). Hybrid remote TYPE is generation, not the +classifier step (`notes.md` §52). +[Cua-S1](https://github.com/trycua/cua/tree/main/libs/cua-s1) is the +same substituted-classifier *job* on a specialist form contract +(option-attention; plan ≠ execute; not TypeSafe Jev; source-only, +`notes.md` §54). +**Harness productization of the same job (Empirical as PR body, +2026-09-19 ~00:48; draft):** +[Stagehand #2951–#2955](https://github.com/browserbase/stagehand/pull/2955) +— the *algorithm* is Stagehand's act/observe/extract loop; the +substituted classifier step is Jev pick among a11y candidates, then +code copies or acts. LLM fallback when the pick/gate/schema fails. +Extract 37/75 no-LLM ~0.5 s vs 4.37 s is *their* card; pick ≠ +replacement. Cache-check errors never block replay. Do not merge +with demo-loop clocks (`notes.md` §57). + +**Robotics text-state, same job different body (Empirical as +showcase class pattern, 2026-09-19 ~00:38):** +MuJoCo robot-arm on [jevable.com](https://jevable.com/): Jev does +not accept images; simplified geometry and contacts **as text**; +two-call split (what to do, then how to move). MOSS: Jev picks the +target; the robot picks up. Cousins: [jev-drone](https://github.com/RomanSlack/jev-drone) +(code at 500/50 Hz, Jev advisory 2.5 Hz); Doom JSON, not pixels. +Drawing-pixel-parallel is a **claim** — contrast MuJoCo honesty. +Do not replace A* or a Sudoku solver with a Noul. Archer still +Watch (`notes.md` §56). + +**Structure induction over a bag (Empirical as a *shape*, 2026-09-18):** +[`Joymfl/dag-jev`](https://github.com/Joymfl/dag-jev) — unordered items +in, pairwise "does i depend on j?" judgments, DAG in `petgraph`. Code +owns topology; the model does not emit edges. Experiment; empty README; +no metrics this pass (`notes.md` §48). Same hole as taxonomy beam (§5): +judgment is a pairwise (or Choice) classifier step, not the scheduler. + +**Combinatorial grid assembly ≠ extractive keep/drop (Empirical as a +negative):** +[`simonmesmith/jev-arc-agi-v1-experiment`](https://github.com/simonmesmith/jev-arc-agi-v1-experiment) +— Direct Jev cell-wise Choice on ARC-AGI-1: **4/400 (1%)**. Dimensions +~90%; complete grids rarely. Many small extractive decisions do not +add up to a consistent transformation. Search / a program / a +simulator stay in code (`notes.md` §49). + +**Query planner as the envelope (author-reported, 2026-09-18):** +[@mmalisper](https://x.com/mmalisper/status/2101001041903009987) on the +Join Order Benchmark. Jev picking join order was **2× slower**. +Cardinality estimates helped when outside context informed the plan; +when Jev was wrong, one query was ~10× slower. Hybrid: Postgres plans +first; Jev overrides **only when confident** → **+12% geomean**, no +dramatic slowdowns. A Jev call is 100s of ms, not yet practical on +every plan. The planner is the hard envelope; confidence is the gate; +fail-open to Postgres. **Hypothesis** until reproduced on *your* +workload. `notes.md` §44. + ## 10. Spec property pipeline (Hypothesis) **Method**: NL/ADR → candidate properties → human strengthens → @@ -424,11 +1268,69 @@ monitor = RV / ptLTL / named invariant # exact, compiled act = code, only if monitor admits ``` +**Named live-stream shape (Empirical as a *shape*, 2026-09-18):** +[`affirmitv/bitrate-advisor`](https://github.com/affirmitv/bitrate-advisor) +— Jev proposes ABR rungs; deterministic policy (probe × headroom, +history percentiles, loss/queue/thermal/battery) is the monitor. Jev +may only match that envelope or be more conservative. Missing the +model returns the policy's answer. Author-measured three states +(~$0.00004, 0.25–0.39 s) are a receipt for the *shape*, not a codec +benchmark. `notes.md` §44. + **Counterexample:** "the model was confident" as the monitor. **Test:** inject a monitor-violating trace the Noul would have admitted; -the sandwich must refuse. **Hypothesis.** Links: `mental-models.md` +the sandwich must refuse. **Hypothesis** as domain-general; bitrate is +Empirical as the named envelope. Links: `mental-models.md` conformal; `formal-methods.md` help list. +**Named compaction envelope (Empirical as README behavior, 2026-09-18 +~16:22):** [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) +— the mutation monitor is **code** (mutating tools, unknown shell, +control operators / pipelines / substitutions / redirections → +`keep_full`). GLiNER2.5 may only propose a reduction on the remainder +or be more conservative; low-confidence / invalid evidence fail closed +to `keep_full`. Soft judgment inside a hard envelope, encoder backend +— not a Jev Score and not a summarizer (`notes.md` §50). Same sandwich +shape as bitrate-advisor; different family. + +**Named stdout-prune envelope (Empirical as README / evals README, +2026-09-18 ~17:15):** +[jev-pruner](https://github.com/tamaratran/jev-pruner) +— the monitor is **code** (≤10k estimated tokens; errors; +JSON/XML/YAML/diff/binary; whole-document commands). Jev Noul may +only score residual noisy chunks. Archive/Jev/incomplete-score +failure keeps the original. Soft judgment inside a hard envelope, +Jev backend — same family as gliner25-compaction, different *job* +(command output vs session memory) (`notes.md` §53). + +**Named computer-use envelope (Empirical as README / architecture, +2026-09-18 ~16:56):** +[gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) +— the monitor is **code** (resolve to an observed node; freshness / +visibility / disabled / occlusion; no generated selectors or JS). +GLiNER2 may only pick among candidates the snapshot already holds. +`DONE` is not the monitor. Soft judgment inside a hard envelope, +encoder backend — not a screenshot VLM (`notes.md` §52). + +**Named specialist-form envelope (Empirical as README / MODEL_CARD, +2026-09-18 ~17:21; weights Watch):** +[Cua-S1](https://github.com/trycua/cua/tree/main/libs/cua-s1) +— the monitor is **code** (plan ≠ execute; dry-run default; one +window; snapshot-bound tokens; reobserve; `execute`/`submit` +opt-ins; fail-closed unknown checkbox; fill execution fails closed +without advertised token `set_value`). The option-attention head may +only pick among observed elements and extracted `Label: value` +entities. Not TypeSafe Jev. No checkpoint scores (`notes.md` §54). + +**Named harness extract envelope (Empirical as PR body, 2026-09-19 +~00:48; draft Watch):** +[Stagehand #2955](https://github.com/browserbase/stagehand/pull/2955) +— the monitor is **code** (schema plan: scalars / bools-enums / +lists of flat objects else LLM; completion gate; screenshot extract +always LLM). Jev may only pick among a11y candidates; code copies +text. Invalid / abstain → LLM. Pick is a fast path, not a +replacement (`notes.md` §57). + ## 13. DST multiverse triage (Hypothesis) **Method**: Antithesis / Resonate DST artifacts → failure taxonomy → @@ -480,7 +1382,20 @@ policy = starvation/fairness rules in code ``` **Example (Hypothesis):** incident-commander assignment; grant-panel -paper allocation; GPU scheduling. **Counterexample:** Choice over +paper allocation; GPU scheduling. **Empirical as a *shape*:** +[`affirmitv/bitrate-advisor`](https://github.com/affirmitv/bitrate-advisor) +— Jev's rung is the soft affinity; probe/history/thermal caps are the +solver; the model cannot violate them (`notes.md` §44). +**Empirical as a *shape* (2026-09-18 ~23:40):** +[slo-router](https://github.com/zeeshan8281/slo-router) — +Jev's task/exactness/evidence scores are soft features; +the controller is min expected cost s.t. health, context, +tools, quality floor, and SLO-success probability. +Fail-open to local features. Measured negative for *sync* +Jev on the fixture (same routes; p95 77.93→490.38 ms). +**Hunch:** never let the decision model be the sole hard +gate on the hot path (`notes.md` §63). +**Counterexample:** Choice over assignees that ignores load. **Test:** a feasible assignment the solver finds that the Score alone would skip because it "felt" worse; hard constraints never yield. Links: `mental-models.md` §OR. @@ -537,15 +1452,25 @@ paraphrase set where the *act* must not change when the wording is synonymous; if it does, abstain. Until that set exists on *your* questions, **Hypothesis**. Links: `mental-models.md` §thresholds; `question-design.md` diagnosis; `validation.md` behavioral tests. +**Option-order cousin (Empirical as v1.2 footnote; 2026-09-19):** +[open-alternative-jev](https://github.com/ikermoel/open-alternative-jev) +scored **72% → 21%** on yes/no answer-judging when A/B were reversed +*theirs* (JevBench v1.2). Same act, swapped labels. Ranked row uses +the author's `A. yes, B. no`. Do not quote one order as the model. +`notes.md` §78. ## 18. Structural prove ∩ soft remainder (Hypothesis as domain-general; Empirical as named shapes) **Method**: code (or a recipe, a law, a text layer) **proves** the easy cases; a System One model judges only what the structure cannot decide. Composition-algebra position 3 *after* a constraint, not instead of one. -**Transfers**: allowlist / refused-in-code / unknown→judge +**Transfers**: allowlist / refused-in-code / unknown→judge. The +allowlist **proves** every verb is a listed read-only tool; the model +judges **only unlisted** leftovers; the gate **cannot block** (fail-open +unless a sandbox sits under) ([jevgate](https://github.com/thevibeworks/jevgate): Proven / Refused / -Unknown; cannot block; Jev alone leaks). Same sandwich as page OCR +Unknown; Jev alone leaks — `/bin/ls` at 0.04 is why it is the third +tier). Same sandwich as page OCR ([doc-router](https://github.com/misbahsy/doc-router): pdf-inspector first, "needs OCR?" Noul on the remainder — 155→87 pages billed, **1.74×** $ on 19 docs / 155 pages). **Does not:** putting the model first so a @@ -565,7 +1490,130 @@ doc-router 9 OCR-misses vs 28 for rules-only. [`poponline63/hermes-jev-north-star`](https://github.com/poponline63/hermes-jev-north-star): deterministic shell checks first; empty evidence refuses to judge; then one Jev call on the remainder. Empty state was self-contradictory — -that is why the refuse-empty rule exists. **Beyond SWE (Hypothesis):** +that is why the refuse-empty rule exists. +[`affirmitv/bitrate-advisor`](https://github.com/affirmitv/bitrate-advisor) +is the same sandwich on a live encoder: policy proves the cap; Jev +may only match it or be more conservative; missing the model returns +the policy's answer. Jev judges only inside it (`notes.md` §44). +Light sibling: +[`phin-tech/pi-jev-approver`](https://github.com/phin-tech/pi-jev-approver) +— regex `commandRules` prove allow/deny (a `deny` is a hard block); +typed Score/Nouls on the remainder; **fail-closed** without a key +(different polarity from jevgate). rh-guard-adjacent; light note only +(`notes.md` §48). +**Regex floor then remainder compact (Empirical as README + +one-session bench; 2026-09-19 ~00:39):** +[jev-compactor](https://github.com/edwardyen724-g/jev-compactor) +— code proves pins, dedup, and `rm -rf` / force-push / `DROP +TABLE` / `curl | sh` regardless of Jev; Jev keep/drop + +Foreman on the remainder. Compaction fail-open if Jev is +down; safety fail-closed on pending destructive/exfil +(`notes.md` §65). +**Local rules then remainder hide (Empirical as README + +small e2e):** +[x-reply-filter](https://github.com/zhuyansen/x-reply-filter) +— `rules.js` proves easy junk at zero cost; four Nouls on +the rest. Auto-hides are not examples until a human +confirms (`notes.md` §65). +**Pre-exec tool product (Empirical as README wiring, not as +accuracy; 2026-09-18 ~17:48):** +[toolgate](https://github.com/fdemir/toolgate) — `allow` / `block` / +`review` before execution; guard error or timeout **stops** (fail- +closed on the execution act). Jev is a probabilistic check, **not +authorization**. 72-case synthetic set is not independently +annotated. Distinct from the ndolinschi *vocabulary* (allow / +ask_human / deny) already in `agent-self-assessment.md`. +`onReview` must obtain authenticated human approval +(`notes.md` §55). Do not copy pnpm. +**Wrap-as-execution then remainder (Empirical as README; +fail-closed; 2026-09-19 ~10:20):** +[AgentGhost](https://github.com/reddpy/AgentGhost) +— `allow`/`ask`/`deny`/`matchArg` prove first; Jev on +leftovers; ASK/DENY throw; judge error → DENY. Distinct +from jevgate (fail-open, cannot block) and from toolgate +(proposed-call pre-exec). rh-guard owns the gate cousin. +Do not copy `npm` (`notes.md` §83). +**Tool-risk as a placement, not a wrap (Empirical as +article; 2026-09-19 ~10:25):** +[@akshay_pachaar](https://x.com/akshay_pachaar/status/2101037514945597645) +cites LangChain middleware as tool-risk gating. That is +the *placement*, not a product card. AgentGhost owns +wrap-as-execution; rh-guard owns the gate cousin. Do not +steal Flavio Copes (`notes.md` §85). +**OMP prompt suppression, host deny proves (Empirical as +measured traffic, 2026-09-18 ~22:38):** +[omp-greenlight](https://github.com/SemetricLabs/omp-greenlight) +— OMP `bash.patterns: deny` **proves** the floor; Jev may +only suppress remaining approval prompts above an +operator-owned bar. Plugin never self-tunes. Not a sandbox. +Default 40.9% / 0 of 94 *theirs*. Distinct from toolgate +(pre-exec of a proposed call) and omp-jev-extensions +(fail-open route). `notes.md` §62. +**Effect-based fast-path then remainder (Empirical as +certification; 2026-09-18 ~23:40):** +[construct-auto-classifier](https://github.com/godspede/construct-auto-classifier) +— fast-deny / fast-allow **prove** catastrophic and +read-only verbs in <1 ms; Jev judges blast radius / +reversibility on the remainder. Fail-closed on the +*execution* act (contrast jevgate cannot-block). Privilege +stripped before the allow rule, not used as the verdict. +Jev 0 dangerous / 975 *theirs*. Landed-script trust / +headless ≠ auto-approve (`notes.md` §68). **Hunch:** contracts on +effects, not tokens (`notes.md` §63). +**Capability kernel, different trust boundary (Empirical as README +architecture, 2026-09-18 ~19:48):** +[interlock](https://github.com/somoore/interlock) — the LLM never +saw the secret and cannot emit an unlisted action; Jev is SENSOR; +policy is the prove/constraint layer. Do not merge with toolgate. +**Human-confirmed kill (Empirical as README safety model):** +[port-cleanup](https://github.com/epiphany-dynamics/port-cleanup) +— Jev recommends; human confirm + identity re-check + shields are +the prove layer for SIGTERM; mapped explanations, not raw model +prose (`notes.md` §59). +[`coldteadotai/abide`](https://github.com/coldteadotai/abide) is the +same *family* on project instructions: the **linter proves** lintable +rules; Jev Scores only residual soft AGENTS.md rules; fail-open, banded +(`notes.md` §47). Different remainder from jevgate's unlisted verbs +and from rh-guard's eval-integrity hole — do not merge products. +**Skills → oxlint (Empirical as a named Phoenix experiment, +2026-09-18 ~18:46):** +[jev-oxlint](https://github.com/cephalization/jev-oxlint) — AST +facts and prechecks in **code**; guidance files copied whole into +`state`; one remaining request of atomic questions; +survey / calibrate / propose. Status: experiment, nothing +published. Phoenix: jev agrees with the human answer key on every +fixture; found a real flush-only-on-success bug (noul 0.07); +routing 0.80–0.94 vs <0.50 across 41 files; coarse hint is not; +~$0.002 fixtures / ~$0.015 41 files; second run zero requests. +Formal methods compose with soft judgment **without hard-gating** +a Noul as a proof. `tenbin` owns the lint skill. License null +this pass. MED cousin: [safe-sh](https://github.com/EpicEric/safe-sh) +(AGPL-3.0) static shell-script analysis — not pre-exec +authorization (`notes.md` §58). Do not copy pnpm. +Compaction polarity is the other way: +[gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) +— code proves mutating / dangerous shell → `keep_full`; the encoder +judges only the remainder; uncertain **fails closed to `keep_full`** +(`notes.md` §50). Same sandwich, opposite fail policy from jevgate +(cannot block) and Abide (fail-open on diffs): the authorized act is +a destructive reduction of memory. +**Stdout prune is the same polarity, different job (2026-09-18 +~17:15).** [jev-pruner](https://github.com/tamaratran/jev-pruner) — +code proves ≤10k / JSON-diff-whole-doc pass-through; Jev scores the +remainder; uncertain **fails closed to original stdout** plus an +archive (`notes.md` §53). Harbor plugin-eval cannot reach Jev and +therefore cannot prune — fail-safe, not a missing score. +**Name the irreversible act (2026-09-18 ~16:48).** Wake *skip* is +irreversible (the agent stays asleep) → +[wakegate](https://github.com/shitianfang/wakegate) authorizes skip +only at p < 0.2 and otherwise **wakes** (fail-open on the skip). +Merge *PASS* is irreversible if the bug was real → +[latch](https://github.com/CaseReed/latch) `--gate` BLOCKs unless +infra is confirmed; the Playwright reporter stays fail-open. +[if-ai](https://github.com/Victor-Casado/if-ai) fails the Action on +error / empty / low confidence (fail-closed on the check). +`notes.md` §51. +**Beyond SWE (Hypothesis):** recipe book ∩ "does this leftover look done?"; labor-law allowlist ∩ hiring-fit Noul; SPF/DKIM pass ∩ phishing Noul on the body. **Counterexample:** Jev on `/bin/ls` as the first tier. **Test:** planted writers never diff --git a/.agents/skills/augustus/references/mental-models.md b/.agents/skills/augustus/references/mental-models.md index ed99861..ea0edeb 100644 --- a/.agents/skills/augustus/references/mental-models.md +++ b/.agents/skills/augustus/references/mental-models.md @@ -52,6 +52,11 @@ until you label *your* cases. | Crossover metaphors | NATM, snap-fit, Norman, Kent, Shirky | This file §crossover | | Formal / semi-formal | Proof vs DST vs judgment | `formal-methods.md`, `formal-semi-formal.md` | | Class / family / objective | Decide vs locate vs categorize vs rank vs perceive | `judgment-class.md` species map | +| Boundary map / extractable-from-state | Self-contained in fed state vs needs outside knowledge | This file §boundary; atlas receipts `notes.md` §49 | +| Control-plane combinators | Then/Gate/Vote/Cascade/Weighted + Router/Loop/Retry/Fallback/Memory; digital-design metaphor ≠ literal AND/OR | `composition-algebra.md`; `notes.md` §66, §69 | +| Conflict ≠ ignorance | Noul collapses both; named Choice escape separates; binary Choice without escape is lexically biased | `question-design.md`; `notes.md` §69 | +| VOI cache / attention admit | Same-intent skip LLM; worth-your-attention before click; skip the narrating second call | jevcache / ThinkyMiner Winnow / hermes-jev-router; `notes.md` §69 | +| Eval integrity (receipts not leaderboard) | Hold/break map + budget-attached bake-off + OOD ECE with sign | atlas / frontier-100 / ood-calibration; `validation.md`; `notes.md` §66 | Pick the pillar from the hole, then the family, then the vendor. @@ -105,6 +110,61 @@ SWE examples are **Empirical** (git-jev-stage, OpenSmoke). The others are curve on held-out *your* cases; score the fallback (escalation is not automatically correct). `mappings.md` §2. +## Boundary map: extractable from state (placement judgment) + +Primary mental model this hour +([jev-capability-atlas](https://github.com/Zaious/jev-capability-atlas); +independent unofficial receipts, not a leaderboard; `notes.md` §49). +Before picking a family or a vendor, place the *task*: + +> Is the correct answer fully recoverable from the `state` you hand +> the model, or does it require outside knowledge that is not in +> `state`? + +| Self-contained (in the state) | Not self-contained (needs outside knowledge) | +|---|---| +| Classify / route / gate over text you already hold | Trivia / recall with no supporting passage | +| Citation / paraphrase / reversed-meaning given claim + quote | Score that needs comparison against a whole field | +| Sarcasm / entailment whose trigger is in the given text | Overlapping blurred categories (dangerous-high ECE) | +| DOM snapshot / numbered candidates → Choice | Combinatorial assembly (grid cells that must agree) | + +**Empirical as that named axis, not as a knowledge-breadth estimate.** +History suite (N=3, single annotator, Chinese history; Case A ground +truth itself contested): common-knowledge item **wrong @ 0.90** with +no context (Yongzheng; Kangxi by popular convention); obscure item +near-flat **0.07** without a passage (correct by luck; informal rerun +wrong @ 0.08) → **right @ 0.97** with the passage in `state` +(Xianfeng, 0.98 mass). Teaching: **bare memory is unreliable; reading +comprehension over supplied text is reliable.** Retrieve first; put +the passage in `state`. Do not treat the atlas 30-second slogan as +the table. + +**Placement, not internals.** The model is not a state machine under +the hood (distributed LM understanding: `paraphrase_support` and +`reversed_meaning_high_overlap` both judged correctly). It *is* +correctly used as a **component node** in *your* program — code owns +transitions (`mappings.md` §3). Confidence is a **statistic from the +distribution** (RLCD trains the distribution; Choice `confidence` is +how peaked it is), not a second trained correctness score. +Calibration is **population-level** and can fail **dangerous-high**: +DAIR Emotion via jev-benchmarks — 48% acc, mean conf **0.819**, 16% +of items p(correct)=0. Overlapping categories, overconfident. Plot +reliability on *your* labels before you threshold. + +**Browser-use is this axis, not vision.** Strength = DOM-as-text + +speculative fan-out over candidates code already numbered — a visual +task translated into extractive text. Not screenshots. Same +component-node placement as lizard-agent / solari-reflex / +gliner2-ultrafast (GLiNER2 encoder backend of the same hole; +`notes.md` §52). Contrast blackwood-rlcd (screenshot + marked +letters). (`applied-mappings.md` §2; `mixed-architecture.md`). + +**Does not:** merge Banking77 87% (atlas/jev-benchmarks) with DMB +76.3% or jevals.com 79.67% into one ranking — protocol / n / split +(`validation.md`, `notes.md` §49). Combinatorial grids are not +extractive keep/drop (ARC-AGI Direct Jev 4/400). FAQ: when-it-holds; +state-machine; retrieve-first. + ## Calibration and cost-sensitive thresholds A number you can threshold is a *decision* number only after you check @@ -116,8 +176,10 @@ LoRA students report agreement with the teacher (`notes.md` §33). Hume prefers the class name **decision models** over "system one" (`notes.md` §33); this file still says System One when quoting TypeSafe. Three open paths, not three species: encoder open-jev, AR constrained -decode (TypeAR + pcdServer), trained decision-only (Laya / Nimble / -Archer Watch). A constrained softmax is still not a Noul (`notes.md` §42). +decode (TypeAR + pcdServer; decision-token LoRA), trained decision-only (Laya / Nimble / +kev / **blackwood-rlcd** / Archer Watch). kev is the runnable Archer reconstruction on that +third path (text-only); blackwood-rlcd is that path with **image-in now** (CC BY-NC); +Watch stays Watch. A constrained softmax is still not a Noul (`notes.md` §42, §45, §46). **Readout versus a token; IIA is a property.** A direct probability and a generated "91%" are different objects; the format calibrates neither @@ -130,6 +192,11 @@ two existing options in every block, and reversing order moved a probability across a ~0.9 threshold. Property-test both (`validation.md`). Correctness is not that confidence field: report both, on held-out cases (`validation.md`, Eval & hill-climb). Stimulus design, not a proof. +Atlas receipt of the same arithmetic: DAIR Emotion mean conf 0.819 at +48% acc (`notes.md` §49) — population calibration can fail +dangerous-high on overlapping labels. DMB S5: jev admits-ignorance +49.7% vs most constrained LLMs 97.3–100% (ECE 0.246). Do not skip +the honesty suite because in-distribution ECE looked fine. For a calibrated binary p and unequal error costs, the Bayes threshold is `t = C_FP / (C_FP + C_FN)` when you act vs not @@ -184,7 +251,10 @@ Contract. The *placement* (gather as an enumerated act) is the method; the calculator is **Hypothesis** until you log act/outcome pairs. Mapping card: `mappings.md` §6. Paying *zero* because a regex already answers is also VOI — abstain from calling any model -(`typesafe-jev-tools`, `notes.md` §42). +(`typesafe-jev-tools`, `notes.md` §42). A harness that *picks which +primitive to run* (JevML's claim: PCA / MCMC / diffusion / NCA) is the +same gate one layer down: maybe none of them (`notes.md` §44). +Hypothesis until that picker has a labeled log. **Transfers:** "ask a second question" / "retrieve one more candidate" / "run the expensive LLM" only when VOI clears the cost. Cheap fan-out @@ -247,11 +317,20 @@ probe. | Business | sales stages | "is this still a real opp?" | amount, close date in CRM | | Life | cook / rest / check | "does this look done?" | thermometer (probe) | | Org | incident command | "is this still contained?" | head-count, location | +| Infra | Postgres query planner | override join/card when confident | the stock planner (fail-open) | +| Live media | ABR rung / resolution | "which ladder step?" | probe × headroom, thermal, battery | Rejected: bandits without observed rewards; Jev as the planner that picks its next tool in a loop (`boundary-audit.md`); PufferLib Ocean scores as a capability claim (`formal-methods.md` DST trio). Mapping -card for the cross-domain loop: `mappings.md` §9. +card for the cross-domain loop: `mappings.md` §9. Soft judgment +inside a hard envelope: bitrate-advisor (ABR) and mmalisper's JOB +hybrid (Postgres plans first) — `notes.md` §44. Compaction envelope +(encoder, not Jev): gliner25-compaction — mutating tools / shell +operators prove `keep_full`; the model may only match that or be more +conservative (`notes.md` §50). Stdout-prune envelope (Jev): +jev-pruner — ≤10k / JSON-diff-whole-doc prove pass-through; Noul on +the remainder; fail-safe keep original (`notes.md` §53). ## Signal detection @@ -304,6 +383,12 @@ This is hospital, aviation, kitchen, boardroom, and agent harness alike: - The sensor ("does this note mention an allergy?") may be a Noul. - Confidence does not waive the constraint. +**Capability kernel (Empirical as architecture, `notes.md` §59):** +[interlock](https://github.com/somoore/interlock) — the sensor is a +parallel Noul battery; the constraint is `policy.py` plus a closed +action space and canaries. Type-safe ≠ correct. Distinct from +asking "dangerous?" after the LLM already held the secret. + Org placement: cheap judgment over every incident step (OpenSmoke shape) so humans only autopsy flags. That is NATM instrumentation of the control structure, not a safety case. Mapping card: `mappings.md` @@ -391,12 +476,141 @@ Use these as *existence proofs of a position*. Write your own card. | Knowledge work | what to read next | on-question Noul + quality Score | library you hold | | Hiring | interview / reject / hold | evidence Nouls; veto rules in policy | labor law, scorecards you wrote | | Inbox | reply / snooze / archive | urgency Noul + aboutness Choice | send, calendar | +| Knowledge work | extract a quote / a cited fact | per-sentence or per-line-id Noul/Choice (**Empirical**: testimonial-miner, jev-reviewer) | verbatim join; place; human publish permission | +| Agent context | compact completed tool results without inventing prose | retention Choice + char-offset locate (**Empirical**: gliner25-compaction; same *job* as fast-jev-compaction / pi-jev-compaction) | mutation/shell envelope → keep_full; fail-closed keep_full; shadowMode before replace; copy exact bytes | +| Agent context | prune Bash stdout before the LLM without inventing prose | Noul per chunk after a hard size/format envelope (**Empirical**: jev-pruner) | ≤10k / JSON-diff-whole-doc untouched; fail-safe original; archive dropped spans | +| Dataframe labeling | classify / score rows | Noul/Choice/Score + full `p__` (**Empirical** as jevframe / jevpandas *shape*) | pandas/Polars, thresholds in code | +| Computer-use speed | one verified act per step | score / Choice among numbered a11y/DOM/OCR+AX/ASR-transcript controls (**Empirical**: solari-reflex Jev; gliner2-ultrafast GLiNER2; laya-mind2web Laya DOM indices; cua-s1 option-attention, source-only, not TypeSafe Jev; Stagehand experimental Jev harness, draft; **closed-vote no planner:** JevOnly; **host-owned:** waymode; **hot-click ego-lite:** ego-jev; **OCR+AX desktop:** typesafe-computer-use hosted Jev, **427★**; **ASR voice-browser:** jev-voice-browser hosted Jev, **103★**) | Guard check; deny-list absence; no screenshots **on the decision**; no waveform to Jev; `DONE` ≠ success; plan ≠ execute; LLM fallback; pick ≠ replacement; type without generation; host handlers/permissions; `--until` beats Jev `done`; exclusive action set; spoken confirm ≠ auth | +| Agent turn | skip memory tour on easy intent | intent Choice (**Empirical**: jev-hermes) | Memory still writes; complex still searches | | Document / lab routing | which pages need the expensive observation | Noul on remainder after a text layer / recipe | local extract, merge order (**Empirical** as OCR-router *shape*) | -| Shell / tool allowlist | unlisted remainder after a proof | five Nouls on unknown verbs | Proven/Refused in code (**Empirical**: jevgate) | +| Shell / tool allowlist | unlisted remainder after a **proof** | five Nouls on unknown verbs | Proven/Refused in code; cannot block (**Empirical**: jevgate) | +| SWE | residual AGENTS.md / CLAUDE.md rules | one Score per named instruction-file rule | linter owns hard rules; bands + fail-open (**Empirical**: Abide replay, `notes.md` §47) | +| Screenshot candidates → act | lettered elements code already marked | Choice over those letters | Click in code (**Empirical** as blackwood-rlcd *shape*; CC BY-NC) | +| Browser / DOM candidates → act | numbered elements from a **text** snapshot | score among those ids (**Empirical**: atlas browser-use / jev-ultrafast / gliner2-ultrafast *shape*: DOM-as-text, not vision; cua-s1 specialist form, source-only; Stagehand a11y + editable-id side channel; **Empirical as README**: typesafe-computer-use OCR+AX macOS, **427★**) | Click / copy in code; no screenshots **on the decision**; hybrid remote TYPE optional; plan ≠ execute; schema/gate else LLM; overlapping options = doubt | +| Extract from a page | values already in element text | pick elements; copy bytes (**Empirical** as Stagehand #2955: 37/75 no-LLM ~0.5s vs 4.37s *their* card) | Schema plan; completion gate; screenshot → LLM; 36/75 LLM-off honesty | +| Knowledge / recall | fact that is not in the document | **Do not ask.** Retrieve the passage first; then a self-contained Choice (**Empirical**: history suite A wrong@0.90 → C right@0.97) | Index, citation, the passage in `state` | +| Dual-process cascade | cheap classify / route vs write | S1 typed decision + τ; S2 generates only on low conf (**Empirical as a productized metaphor**; routing accuracy **unmeasured** — dual-process-ai). **Harbor-shaped cousin:** decide→policy→LLM leftover on labelled emails (**Empirical**: jav-email-cascade; Noul 0.5 never rounded; mock gen-json flat-confidence is *their mock*) | Safety still fail-closed in code | +| Domain specialist vs few-shot | when policy reads p vs when only argmax | Train local LoRA on independent gold if downstream uses the distribution; hosted+examples if argmax (**Empirical**: Domain-jev-maker KL 0.168 vs 0.580; McNemar n.s. on few-shot determinate) | Threshold / EU in code | +| Semantic `ORDER BY` | put rows in a defensible order | Measure pairwise inversion / Score ordinality / ties; do not treat ECE as the sort certificate (**Empirical**: jev-orderby-bench six gates; Score 0.143 weak link; 53-way 0.99 tie) | Secondary key; measure request shape | +| CI merge-gate | ignore infra noise without merging a real bug | cause Choice per cluster (**Empirical**: latch demo PASS vs BLOCK) | Cluster + fingerprint + `--gate` table; reporter never fails the runner | +| Sleeping-agent resume | skip a worthless LLM turn | p(wake) (**Empirical** as safety table; 21/21 smoke — wakegate) | User-message / skip-limit / error always wake | +| Claim integrity at Stop | do not ship hallucinated-done | supports/contradicts vs session evidence (**Empirical**: clear-head) | Keyword retrieve; firm-confidence floor never blocks | +| Code-graph index | cheap S1 extract, S2 only on the tail | GLiNER locate + confidence escalate (**Hypothesis** as 10–50×; **Empirical** as degraded-load / no-invent-edges) | Graph in code; do not dump repo if S1 failed to load | +| Combinatorial puzzle | whole grid / program that must be consistent | **Rejected as extractive.** Cell-wise Choice assembly is not keep/drop (ARC-AGI Direct Jev 4/400) | Search, a program, a simulator | | Moderation | hold before publish | hazard Nouls (**Empirical** as family) | block/review policy | | Phishing / fraud screen | hold vs deliver | SDT criterion on a Noul | blocklist, SPF/DKIM exact (**Hypothesis**) | | Personal ops | cook done / not | "looks done" Noul | thermometer probe | | Org safety | stop the line | sensor Noul | interlock, two-person rule | +| Knowledge / RAG | reason only over kept evidence | retrieve wide → decide → evidence set (**Empirical** as architecture: decision-native-rag-skills; classify-first MCP cousin: jev-sift; **Hypothesis** as a measured win) | Conflict/temporal/provenance in code; embeddings / file lists generate candidates; errors/truncation ≠ irrelevant | +| Agent I/O | classify first, read selectively | batch path/url/text → relevance or typed questions (**Empirical** as README: jev-sift; topology A MCP) | Hard envelope (50 / 60k / 2MB / public-IP); main LLM opens survivors | +| Spreadsheet / catalog | named semantic columns | heading scores each row (**Empirical** as *shape*: jevpandas / jevframe; jevable intent columns). Snack MCDA clocks are **claims** | Weights, vetoes, exact fields in code | +| Robotics / control | observe → decide → act on a body | Choice on **geometry-as-text**, not pixels (**Empirical** as showcase: MuJoCo / MOSS; cousins jev-drone, Doom JSON; **Empirical as README delta**: khordoo/jev-reflex-autonomy-lab — S1 keeps flying, S2 one-use, no graphical input) | Kinematics / Hz / physics in code; two-call split; do not replace A*. Drawing-pixel claim ≠ Archer. S2 never grants. 20% still soft | +| Draft quality gate | kill drafts that break rules | quality Noul/Score (**Empirical** as fail *mode*: silence treated as safer) | Fail-open / heartbeat on missing verdict; contrast Abide `<0.5` (edit proceeds) | +| Session memory | next task sees last session's facts | scored recall over a verbatim ledger (**Empirical**: carryforward; 9×3 hint; **0/4** recall) | Constraints always-keep; fail-open dump; SessionStart > hoping | +| Application control flow | `if` / `case` on a judgment | `chance`/`pick`/`rate` as language primitives (**Empirical**: hunch; English-as-config) | Fail polarity per action; stub backend | +| Healthcare huddle / recon / inbox | escalate / hold / route | S1 remainder after NEWS2/code (**Empirical** as synthetic report: explore-typesafe-ai; **not clinically validated**) | NEWS2, recon, routing in code; S2 blinded review | +| Intent cascade vs nano/encoder | escalate when unsure | pre-registered kill/go (**Empirical as practice**: jev-baselines-eval **AMBIGUOUS**; cascade sign-flip; encoder-with-labels wins) | Thresholds, serving-path honesty, ECE if you claim calibration | +| Public primitive / wall | typed answers on a sentence | six parallel questions (**Empirical** as README: ask-jev-ai; cost-to-1M from tokens) | Policy-in-code; no-key allowlist; safety threshold in code | +| Codebase meaning-search | relevant file/chunk without knowing names | packed parallel relevance (**Empirical**: jevgrep 79% top-5 vs BM25 40% / grep 20% on stripped repos). **Line meaning-grep** AND/OR/NOT after threshold (**Empirical**: jev-semgrep 0.94/0.98 *theirs*; proposition ≠ embedding; contrast-set refund; Semgrep.dev collision; not a gate; `notes.md` §86). **Evidence packets** index-once (**Empirical**: jevex 1/8→6/8 n=8 *theirs*) | Keyword still wins exact strings; packet HitFile 0.233 is diagnostic; Japanese noisier near threshold; do not multiply parallel p | +| PR review attention | where a human should look | P0/P1/P2 (**Empirical** as README: egma-ai/jev-reviewer). **Not** correctness; **not** choxos pointer-not-generator | alwaysReviewPaths P0; incomplete never P2; generator writes deltas | +| Skill-derived lint | remainder after AST/precheck | Noul/Choice on guidance in state (**Empirical** as Phoenix: jev-oxlint) | Parser/precheck in code; not a hard gate; `tenbin` owns lint skill | +| Session model route | which model for this thread | first-prompt Choice, then lock (**Empirical** as README: jev-adaptive-thinking) | Fail-closed declared fallback; never reclassify later turns | +| RAG vs generative rerank | which passages to keep | pointwise relevance (**Empirical** as one-run: Jev-RAG ≥70%/72% vs Spark rerank; full-context Spark still faster) | Embeddings generate candidates; name the no-RAG arm | +| Untrusted agent / secrets | never hold the real key | hazard Nouls as **sensor** (**Empirical** as architecture: interlock) | Closed action space; canaries; `policy.py` BLOCK/ASK/ALLOW; type-safe ≠ correct | +| Live chess coaching | speak only when it matters | severity / interrupt / error-class (**Empirical** as Wave 0 PRD: game-coach) | Stockfish owns eval; templates + capped writing model own words | +| Kill a listening port | stop stale listeners without murdering the wrong PID | Stop/Keep/Review Choice (**Empirical**: port-cleanup) | Human confirm; identity re-check; shields override; mapped explanations | +| Calibration measurement | honesty of native probabilities | Brier/ECE/reliability on analytic worlds (**Empirical**: jev-arena live Brier 0.0059 / ECE 0.0620 *theirs*) | Oracle stub; fan-out batches; not verbalized confidence | +| Ranking vs calibration | does `ORDER BY` put rows right | Pairwise inversion / Score ordinality / two-decimal ties (**Empirical**: jev-orderby-bench) | Secondary key; do not quote ECE as sortable | +| Training-data VOI | which unlabeled rows are worth an expensive label | Confidence routes accept / teacher / human; log full distributions (**Empirical**: jev-triage) | Real outcome labels stay the targets; do **not** distill Jev as teacher (~68% ceiling) | +| Constrained-AR speed vs calibration | O(1) structured decode vs a Noul | Measure Brier/ECE, not only latency (**Empirical**: system-one-benchmark n=50; PCD Brier 0.3884 vs Jev 0.1096) | Schema-valid is not calibrated | +| Closed-vote CU | task with no planner LLM | Code builds options; model only picks (**Empirical**: JevOnly; waymode host-owned) | Type without generation; `completed` ≠ server-state success | +| OMP/pi gate | done-check / subagent topology | Choice, not boolean; fail-open missing Jev (**Empirical**: omp-jev-extensions) | Contrast pi-jev-approver fail-closed | +| Permission vs probability | auto-approve a gated tool call | Operator-owned criterion; plugin never self-tunes the bar (**Empirical**: omp-greenlight 40.9% / 0 of 94 *theirs*) | Not a sandbox; host deny stays above; live traffic unlabelled | +| Judgment ≠ permission | which specialised skills to inject | Jev scores relevance; code owns grants (**Hypothesis / outline**: skill-broker) | Never broaden access on Jev failure; not a production recipe | +| Eval integrity / instrument | is this eval's score trustworthy | Audit data/scorer/runs/claims; test a Jev question like an if (**Empirical**: dinostomp; ECE 0.062 *theirs* on 24) | 99 of 189 findings against itself; not a Harbor taskset | +| Constrained optimizer + S1 features | which backend meets quality + SLO at min cost | Judgment as a *feature*; solver owns floors (**Empirical as shape / negative**: slo-router p95 77.93→490.38 same routes *theirs*) | Never the sole hot-path gate; fail-open local features; eight-row demo is not a benchmark | +| Privilege ≠ verdict | is this shell command safe | Effect semantics + independent risk Nouls (**Empirical**: construct-auto-classifier; Jev 0 dangerous / 975; chat leaked) | Fast-allow/deny prove; landed-script trust; headless ≠ auto-approve; fail-closed | +| Attention filter / human-review VOI | do I need to look at what the agent did | Per-file Nouls; never blocks the agent (**Empirical as README**: rashedInt32/jev-lens; never green unless sure) | Not a permission gate; distinct from dizk/jev-lens pre-send views | +| Measurement owns endorsement | is this question pack shippable | Evidence-gated accuracy/ECE/cost/latency on a pinned version (**Empirical**: jev-packs nine verified *theirs*; **jevassert landed** record/replay CI) | No numbers → `provisional`; `unknown` mandatory; `check` offline | +| Jev supplies evidence, code owns authority | may this tool call run | Deterministic policy ALLOW/REVIEW/BLOCK; Jev is the sensor (**Empirical as slogan**: actiongate-jev) | Positive p never overrides a hard fail; fail-closed on irreversible classes if Jev is down | +| Ranking ≠ calibration | can I threshold raw p as a frequency | AUC vs ECE/Brier vs human rates (**Empirical**: 8,000 judgments; stated ~75% vs human ~10%; ~96% ECE removed) | Recalibrate on *your* labels (`jevcal` ~100 rows); vendor "calibrated" often means rank-correlation | +| Hot-click CU | next click / type from a viewport | Indexed element table → operation+target (**Empirical**: ego-jev; ~2× vs per-step LLM, n=3, not a bench) | Code owns observe/execute/`--until`; generator only for type; Jev `done` ≠ success | +| Compact without paraphrasing | drop irrelevant history, keep bytes | Keep/drop per message; pins + regex floor in code (**Empirical**: jev-compactor later **73%** / 350 ms / 4 of 4 vs shipped summarizers; §65 vs-Sonnet 64.5%/366ms) | Never rewrite; compaction fail-open if Jev down; safety fail-closed | +| Hold-before-show social | collapse junk replies | Local rules prove easy junk; remainder Nouls (**Empirical**: x-reply-filter) | Collapse not delete; never auto-train on the model's own hides | +| Never confidently wrong | protocol verdict under noisy evidence | TLA+ quorum + stability; Jev is the oracle (**Empirical**: jev-labs 1,080 golden 0 wrong *theirs*; escalate 5%→18% under severe) | Escalate is allowed; not a proof of zero; synthetic ≠ clinical | +| Advance / coverage | whether the world may change | Seal + coverage.path ledger (**Empirical as README**: seal; Jev answers questions, SEAL answers advance) | Exception queue visible; mint ≠ product brain; code seals first | +| Sureness of a distribution | act / escalate / abstain | max_prob/margin/entropy/gini (**Empirical**: how-sure-is-jev; Choice confidence = max_prob) | Bands are policy; pair with OOD; max_prob is generous | +| Cheap review triage | auto-approve / human-review / block | Four typed questions before expensive review (**Empirical**: ci-gatekeeper 504–629 ms *theirs*) | Operator owns thresholds; distinct from latch flaky-vs-real | +| Attention redirect (agent Stop) | one more look vs finish | Eight risk Nouls (**Empirical**: jev-preflight; fail-open; 0.85 uncalibrated) | Not a merge blocker; not rashedInt32/jev-lens | +| Pre-send perception | which lines enter the prompt | Code-built views; Jev picks (**Empirical**: dizk/jev-lens 79% fewer tokens / 500 trajectories) | Compress before first send; code full unless confident | +| tools≠use | will the agent call memory? | SessionStart injects; tools sitting there are not VOI (**Empirical**: carryforward 0/4) | Hook > hoping | +| Observational memory | what to keep, what kind | Keep/kind; verbatim ledger; model-free compact (**Empirical as README**: pi-om) | Failed Jev does not drain buffer; not a summary | +| Physical-world S1 | typed house questions | Sensors + automations (**Empirical**: HA-Jev 17★) | Not for locks/heaters/smoke; arithmetic in templates | +| Open-Jev class | finite choice + prob without TypeSafe | LM/vision/voice; JevPick; `/v1/systemone` wire (**Empirical**: openvons) | NOTA; execute/confirm/reject; not a replica | +| Judgment outside the store | semantic SQL over vanilla Postgres | CLI judges; DB sees ordinary SQL (**Empirical**: jevql) | Contrast pg-jev in-engine; cheap SQL first | +| Record/replay eval | can CI gate accuracy+calibration+cost | Record once; replay offline (**Empirical**: jevassert) | Live calls belong in `record`, not in PR CI | +| Failure-finding vs leaderboard | where does the judge fail | Reviewed atlas, not a winner crown (**Empirical as README**: chenmingtang830/jevarena; harness not findings) | Qualify vs meetr1912/jev-arena | +| Stereotype / uncertainty / cost | does missing evidence leak a stereotype | Typed Choice + unknown option; report bias **and** accuracy (**Empirical**: BBQ 97.28% / 0.04 / 0.34 / $0.3429 *theirs*) | Not a general bias cert; 12/13 amb errors stereotype-aligned | +| Decider ≠ executor | who picks the next act vs who writes args | Jev next-tool/progress/risk/done; LLM fills (**Empirical as README**: jeffrey) | Risk≥0.5 pause; stuck ladder; not a planner-writer | +| Decider ≠ executor across timescales | who flies vs who advises | S1 typed action every tick; S2 one-use strategy (**Empirical as README**: khordoo/jev-reflex-autonomy-lab) | S1 never stalls; S2 never grants; jeffrey is the SWE cousin | +| Escalate without stalling | pay S2 only under threshold, keep the loop | Async planner; reflex keeps steering (**Empirical as README**: khordoo; cousin classifier.dev smart *does* wait) | 20% starting gate *theirs* still soft; not Harbor τ | +| Mixed-initiative consumption | was the advice actually used? | Purple confidence = consumed; purple S2 bar = arrival; red = fail (**Contract as telemetry**: khordoo) | Arrival ≠ used. Green = local context | +| Local controller ≠ localjev | which reflex backend? | Rule-based built-in vs hosted `jev-latest` vs prompted-JSON Bun vs ONNX (**Contract**: khordoo Local/Live; **≠** githubnext/localjev **≠** kunchenguid/local-jev) | Same physics/seed; not a scored bake-off | +| Split kind/item/site | one 255-way soup vs three questions | Parallel Choices; used-only-for-matching-kind (**Empirical as README**: typesafe-computer-use) | Off-screen is a fourth question, not mixed into items | +| Exclusive CU actions | overlapping labels as false doubt | Confidence is concentration (**Contract as README**: typesafe-computer-use; wellposed cousin) | Missing `other` is the quiet 1.00 failure | +| Perception rebuild | what frontier reads from pixels for free | OCR crop/tile, AX walk, dates.py, clock, URL (**Empirical as README**: typesafe-computer-use) | AX never sole (Spotify 0). Decision ≠ answer-reader capture | +| ASR as perception | waveform vs transcript | Web Speech producer; Jev on text-state (**Empirical as README**: jev-voice-browser; compose with OCR §81) | Audio never enters the Choice. Skip Archer | +| Partial-speech VOI | act now vs wait for the rest | `complete` Noul + silence; closed-set may fire; free-text waits (**Empirical as README**: jev-voice-browser) | Truncating "search for alan" is the cheap failure | +| Spoken confirm ≠ auth | destructive click | Second Noul path; convenience not guarantee (**Contract as README**: jev-voice-browser) | Control-port reach is the real grant | +| Overlay disambiguate | which of 2–3 targets | Numbered badges; spoken digit; no second model (**Empirical as README**: jev-voice-browser) | The id is already in code | +| Wrap-as-execution | can the model skip the judge? | The wrap *is* the tool function (**Empirical as README**: AgentGhost; ASK throws; fail-closed) | Advisory sidecar is theater. rh-guard owns the gate cousin | +| Rules first then remainder | which verbs skip the model | allow-list proves; Jev on leftovers (**Empirical as README**: AgentGhost `allow` skips judge) | Contrast fail-open allowlist that cannot block | +| ASK throws | can HITL be silently skipped? | Default errors; wire `approveWith` (**Contract as README**: AgentGhost) | `AUTO_APPROVE` is a demo hatch, not a grant | +| Genre atlas ≠ bake-off | is this a rank? | Apps by hole; stars research-time (**Empirical as tweet**: [@studio_yebisu](https://x.com/studio_yebisu/status/2101065176069886152)) | ≠ class census §77 ≠ v1.2 board. Likes ephemeral | +| LLM hammer for bounded decisions | does this call need generation? | Typed answers when code already knows the options (**Empirical as article**: [@akshay_pachaar](https://x.com/akshay_pachaar/status/2101037514945597645)) | Mixed architecture, not stack replacement | +| schema-safe ≠ correct | can it still be wrong? | Cannot invent out of schema; can pick the wrong valid option (**Empirical as article**: Akshay; safer *theirs*: schema holds, judgment can fail) | Cousin of type-safe ≠ correct / jaggedness | +| Questions-as-code / shadow rollout | may this branch go live? | Rubric first; shadow beside current; plot accuracy vs confidence; pin questions (**Empirical as article**: Akshay) | Do not rebuild the agent first. 200×/400× are TypeSafe ceiling | +| Sentence-as-rule | does this named artifact contradict itself | Structural matcher × one sentence scored (**Empirical**: mizchi/jevlint 13/15 1.00/1.00 *theirs*) | Mechanical defects stay with the compiler; qualify vs huntedman/JevLint | +| VOI admission (expensive review) | which hunks are worth a generative look | Typed per-hunk probabilities; safety keep-set in code (**Empirical as pilot**: prune-review 1.18% with 305% outlier *theirs*) | Cost ≠ quality; ~20% is a target not a result | +| Whole-repo intent | does unchanged code still violate the ask | VERIFIED/VIOLATION/UNKNOWN (**Empirical as CLI**: jev-intent-review) | Empty search ≠ proof; observation window ≠ the diff | +| Persist constraints | will "don't touch that" survive compaction | Structured policy + replay; Jev classifies meaning (**Empirical**: pi-heed 98.5%/0 false block *theirs*) | Jev never writes policy; fail-open | +| Open replica substrates | same contract, different engine | Isolation / byte-parity / agreement tests (**Empirical**: grande JGLUE; laya-jolt golden; local-jev 30%/57%; JEV-CPU PoC) | Softmax ≠ Noul; spec ≠ product; Meanblock 404; Archer Watch | +| Harbor SGR-judge contract | does evaluation need generation? | Frozen protocol vs schema-guided LLM judges; invalid = FN; cost/latency first-class (**Empirical as contract**: jev-judge-bench; **no quality headline yet**) | Canaries ≠ F1; incomplete cohort ≠ replacement claim; qualify vs jevarena/jevbench | +| Hand no-text steps | which loop steps need no writing | Plugin/control plane; writing stays generated (**Empirical**: jev-use 220 ms p50 / 12/12 gate *theirs*) | Vercel drops confidence → margin ≠ vendor head; fail-open gate never grants | +| Control plane, not a second agent | router / gate / retry / sieve / review / click | Named sensors; fail polarity per act (**Empirical as README**: pi-jev-control) | GUI never force-click; compaction never writes the session | +| Generation as a tree of Choices | next word without free generation | One typed question per choice over a closed lexicon (**Empirical as README**: jev-gpt ~400 calls / 75 s / 2¢ *theirs*) | Architecture demo; not a product writer; still pick ≠ fill | +| Recipe atlas (code prepares) | which narrow questions fit this job | Samples show technique; policy in code (**Empirical as recipes**: jev-cookbook; 16–36 not benches) | Thresholds are a dial; numbers/dates stay exact | +| Personal history without a social graph | what to show next from *your* trail | Rank outbound links; distribution *is* ranking (**Empirical as README**: jevfeed) | Generating the next look converges on a mirror; history never uploaded | +| Dual-channel ECE / claim-audit | is this NAR "better calibrated"? | Like-for-like channels; n and CI before SOTA (**Hypothesis until independent run**; openJev-verdict-2.0 + PR #1) | Throughput ≠ latency; correctness-head ≠ distribution ECE; ≠ IamBusy/OpenJev | +| 1-token logprob ≠ Noul | can a generic LLM's next-token mass be the judge? | Constrained decode over caller-enumerated labels; coverage is format-mass (**Empirical as 336-case GUI**: chakuho 27B 95%/92% vs Jev 89%/82% *theirs*) | Softmax ≠ Noul; 8B coverage 1.00 while `__none__` collapses; arithmetic in code | +| Open replica runtime | same wire, faster forwards | Prefix reuse + family adapters; argmax-parity is the honesty check (**Empirical as README**: jevinf 2.57×/2.27× 100% argmax *theirs*) | MPS only; not ECE; not TypeSafe | +| Files-to-read VOI | which ranges change the next Read | Index-once, BM25, Jev packet (**Empirical as n=16**: jevex 160s→69s / $8.74→$3.13 / 16/16 *theirs*; keep n=8 finish 1/8→6/8) | Not a patcher; HitFile diagnostic; rename of jev-semantic-explorer | +| Commit attention ≠ verdict | is this commit worth a human look? | Typed Nouls + middle "review" band; regex proves literals (**Empirical as 13 labelled**: commitjev; 0 false on 5 clean *theirs*) | Small control; never round the middle; Nouls decide, Choice headlines | +| Decision-native queue | where should limited attention go next? | Atomic signals × deterministic policy (**Empirical as README**: mailordinal 100-point; humans own ambiguity) | Do not ask "how urgent"; arrival time is a poor proxy | +| Language OOD / confident-wrong | will p drop when the checkpoint cannot read? | Route by script **before** the forward pass (**Empirical as MASSIVE**: laya-multilingual; Khmer 0.000@0.952 *theirs*) | English checkpoint mean conf never < 0.885; gating cannot catch; ships uncalibrated | +| Schema-conditioned ranking | new labels without retraining | Scalar head per candidate; code softmaxes (**Empirical as Hub eval**: schema-scorer v2 Choice 0.841 *theirs*) | Peaked one-hot training ≠ calibration; GitHub 404 this pass | +| Branding ≠ backend | is this actually System One? | Read the client, not the badge (**Contract**: hermes-plugin-jev is Agnes chat-completions) | Distinct from hermes-jev-router | +| Productized System One HTTP | label + calibrated p as a public contract | Batch `{id,text}[]`; LLM fallback only (**Empirical as README**: classifier-dev **185★**; 400 headlines 650 ms *theirs*) | Distinct from ask-jev-ai wall; policy stays in code; not omni | +| Escalate-under-threshold | pay S2 only where p might change the act | Smart re-asks single-label <0.7; multi-label ignores (**Empirical**: classifier-dev ≥0.9→82% / <0.5→29%; gemini 87.5→90.0 / 61.8→63.7 *theirs*) | 0.7 is *theirs*; re-judge that made it worse is not VOI | +| Silent-fallback honesty | which model actually answered? | Named `FALLBACK` marker; alerts on a quiet chain (**Empirical**: granite F1 **0.546** vs advertised ~**0.800** *theirs*) | rh-guard owns the gate; dinostomp owns the instrument | +| Evidence-synthesis pointer | which line in the paper is the quote? | Jev picks ids; code copies verbatim (**Empirical as README**: choxos/jev-reviewer **12★**; 18-q **4.6 s / $0.0101** *theirs*) | **≠** egma-ai attention. *Not found* is an answer. Spot check ≠ validation | +| Two-pass relative + absolute | which line, and does that line itself answer? | Choice (+ none) then per-line Noul (**Empirical**: choxos; quotes Noul ≥ 0.5 *theirs*) | Multi-row tables need both; 0.5 is *theirs* | +| Human check as productized judgment | may this quote enter the review? | Tick/edit; checked never overwritten (**Empirical as README**: choxos) | Jev SENSOR; reviewer constraint. Not optional chrome | +| Wire-compat ≠ logit-equiv | does `/v1/systemone` mean the same p? | Prompted JSON + entropy-conf vs structured logit read (**Contract**: githubnext/localjev vs razorback16/openjev) | SDK drop-in is the wire. JSON-valid ≠ picked-right. **≠** kunchenguid/local-jev | +| Institutional open-replica | who ships the interchange? | GitHub Next local Bun bridge (**Empirical as product**: githubnext/localjev **261★**) | Legitimacy ≠ quality headline. Softmax / generated JSON ≠ Noul | +| Prompted-JSON bake-off | which local backbone on this *pipeline*? | Frozen AG News/BoolQ/SST-5; 1,200 req; caveats first (**Empirical as eval**: Qwen3.6 76.7% / Gemma 26B 75.0% / DiffusionGemma 74.2% short *theirs*) | No definitive winner (2/120). Not logits. Not calibrated. Not a Harbor taskset | +| Packaging ≠ new species | is this a new head or the same Laya? | GitHub/PyPI + Router over Hub ckpts (**Empirical as README**: NandhaKishorM/laya **710★**) | Weights stay convaiinnovations/*. **≠** TypeSafe drop-in. **≠** localjev | +| Token-budget cardinality | why does Jev win >20 options? | Options share `head_max_len`; ~3–4 tok/label at Banking77 (**Empirical as README**: 0.425 vs Jev 0.870 *theirs*) | Jev 255 options. Hierarchical Choice, not a silent cfg copy | +| Post-T ECE ≠ raw ECE | which ECE is on the badge? | Temperature per (type, K) on held-out (**Empirical**: 0.466→0.081 / 0.314→0.106 *theirs*) | vs-Jev 0.081 is post-T. Raw typed-decisions 0.213 vs Jev 0.144. Multilingual ships uncalibrated | +| Soft 0.85 gate | may I auto-act? | RLCD makes p *meaningful*, not Harbor-calibrated (**Contract**: README snippet) | Khmer 0.000@0.952. 0.85 is *theirs*. Route before p | +| External census ≠ scored bake-off | is this a rank? | Named list + a promised board (**Empirical as tweet**: [@airesearch12](https://x.com/airesearch12/status/2101259522933186879); watch [jev-models](https://benchmarkheaven.com/jev-models)) | ≠ jevbench v1.1 §67. Do not paste live ranks here. Likes ephemeral | +| Class-boundary (what belongs) | is GLiNER2 / a router an openjev? | Locate/categorize encoder + catalog routers counted beside NAR wires (**Contract as their list**) | Same job family ≠ replica. Needle 3 already not Jev-class §67. Qualify namesakes | +| Incomplete census vs watch | are missing names out of class? | Laya / localjev / kev / TypeAR / openvons / chakuho / jevinf / grande / laya-jolt / blackwood / classifier-dev absent | Lag, not a dunk. Completeness is a board watch item | +| Harbor honesty watch | what must a public openjev board disclose? | Calibration on/off the rank; cost/latency assumptions; silent fallback; partial runs (**Empirical as v1.2 board**: `notes.md` §78) | Soft-score-as-hard-rank is a design. Mixed class needs a class column | +| Geometric-mean product | can accuracy buy back a weak axis? | I/C/S/K 25% each; `exp(sum 0.25 ln max(axis,1))` (**Empirical as board**: Jev **75.3** / SemIf **74.6**; Luna I=96.8 rank #7 *theirs*) | Weights are a choice. Do not mix with v1.1 87.6. ≠ tweet census | +| Weight sensitivity | does the rank survive a different product? | Same axes, other views, still geo-mean (**Empirical**: no-cal SemIf #1; cost system-one-open #1 / Jev #5 *theirs*) | Official Score is 25:25:25:25. Other buttons are not the Score | +| Option-order brittleness | does A/B order change the act? | Reverse yes/no labels (**Empirical**: open-alternative-jev 72% → 21% *theirs*) | Cousin of paraphrase-brittleness. Ranked row uses author's order | +| Instruction models in a System One table | is the class the task or the architecture? | JSON-schema instruction models beside NAR rebuilds (**Contract as their legend**) | Luna/Gemini/DeepSeek/Qwen3.8. Needle 3 label-only. OpenJev = razorback16 ≠ IamBusy | Rejected in every domain: replacing the ledger with a vibe; replacing the interlock with confidence; replacing the essay with a Noul; @@ -408,6 +622,31 @@ threshold cargo-cult; ontology capture (Choice set smuggles the conclusion); Goodhart on the judge; coverage theater; scale mismatch (Web metrics on situated problems); vacuous assurance. +## Apply this hour's class (already folded — do not re-card) + +When the hourly named HIGHs are already on the branch, extract +**how to apply**, not a dump (`notes.md` §79): + +1. **Wire-compat ≠ logit-equiv** — same `/v1/systemone` SDK can + sit on prompted JSON (entropy-as-confidence) or on structured + logit read. Calibrate before consequential use. +2. **Productize label + *p*; mark FALLBACK** — escalate-under- + threshold is policy in code; a silent head-swap is a lie + about the instrument. +3. **Packaging ≠ new species** — a Router is a face. Route by + script before *p* when gating cannot catch OOD overconfidence. + 0.85 is still soft. +4. **Pointer-not-generator** — point at ids, copy verbatim, + *Not found* is an answer, human tick never overwritten. + Works at Cochrane/PRISMA scale, not only SWE. +5. **External list ≠ scored bake-off** — census VOI is + completeness/class-boundary; a board VOI is the score + function, cal on/off, and named cost/latency assumptions. + +A Noul may attend or escalate. Hard-gating it as a PR/quality +seal is soundness theater unless an exact envelope already +proved the irreversible act. + ## Decision-design extras (any domain) ```text diff --git a/.agents/skills/augustus/references/methods-catalog.md b/.agents/skills/augustus/references/methods-catalog.md index a950f66..eee0482 100644 --- a/.agents/skills/augustus/references/methods-catalog.md +++ b/.agents/skills/augustus/references/methods-catalog.md @@ -31,9 +31,13 @@ judgment component is new). | Self-consistency / ensembling of judges | Repeated independent ratings of the same object | N repeats over one state (output tokens free); entropy/disagreement across repeats as the review signal | Aggregation, escalation policy | **Empirical recipe** (self-consistency: nouls cookbook) | | Judge qualification (interrater reliability) | A judge worth gating must be repeatable | Repeated judgments over frozen outputs before trusting either Jev or LLM as judge | Variance stats, agreement metrics | **Empirical recipe** (jev-as-a-judge: 224–279× tighter than GPT judge) | | Neyman–Pearson / selective classification | Decision threshold under error costs | One threshold per action, set on split A, reported on split B; abstention path | Loss model, ROC analysis | **Contract + empirical** (confidence-routing; evaluator script) | -| Value of information (EVPI / EVSI) | Whether another observation is worth its cost | Gather as an enumerated act; pay iff expected decision-loss drop > cost | Cost of the observation; the loss table | **Hypothesis** as a numeric calculator; **Contract** as the placement (`mappings.md` §6) | +| Value of information (EVPI / EVSI) | Whether another observation is worth its cost | Gather as an enumerated act; pay iff expected decision-loss drop > cost. Wake/resume: skip the LLM turn only if the judge answers and p is low. Selective memory: score the ledger against the task; dump on failure. Classify-first: pay for a full agent open iff relevance says it might change the act | Cost of the observation; the loss table; skip-limit; always-keep rules; hard I/O envelope | **Hypothesis** as a numeric calculator; **Contract** as the placement (`mappings.md` §6). wakegate 21/21 is smoke (`notes.md` §51). carryforward 9×3 is a hint (`notes.md` §55). **0/4** recall is the tools≠use finding (`notes.md` §68). jev-sift mocks ≠ accuracy (`notes.md` §56) | | Signal detection (Green & Swets) | Evidence variable + criterion | Noul as noisy evidence; t from costs and base rate; ROC/PR on your labels | Operating point, base-rate tracking | **Hypothesis** for non-SWE plots; **Empirical** as moderation *shape* (`mappings.md` §7) | -| Reliability calibration (Platt/temperature) | Raw scores → calibrated probabilities | Noul is natively calibrated **in-distribution only**; verify with reliability bins on your own population; re-fit a correction out-of-distribution | Calibration fitting, binning | **Empirical recipe** (ECE 0.0313 in-distribution; 32% OOD collapse — Archer Hume) | +| Reliability calibration (Platt/temperature) | Raw scores → calibrated probabilities | Noul is natively calibrated **in-distribution only**; verify with reliability bins on your own population; re-fit a correction out-of-distribution | Calibration fitting, binning | **Empirical recipe** (ECE 0.0313 in-distribution; 32% OOD collapse — Archer Hume). Atlas: DAIR Emotion dangerous-high (48% / 0.819); DMB S5 ECE 0.246 (`notes.md` §49) | +| Frozen-protocol bake-off vs constrained LLMs | Same items, accuracy + ECE + latency + cost + honesty | Decision-model as one contender class, not the score | Protocol, raw logs, baselines | **Empirical as Harbor/jevals practice** (DMB v2; jevals-data CC-BY-4.0 recompute-from-logs; `notes.md` §49) | +| Pre-registered cascade vs nano/frontier/encoder | Kill/go printed; cascade R at a stated accuracy target; error-ranking AUROC ≠ ECE | Confidence as an escalation signal is a *hypothesis to kill* | Pre-reg hash, paired CIs, margin sensitivity, serving-path controls | **Empirical as Harbor/jevals practice (honest negative)** (jev-baselines-eval: both AMBIGUOUS; cascade sign-flip at exact parity; confidence=1.0 on 102/200 incl. 6 wrong; encoder 0.933/9ms; ~2.2× serving-path; errata ×3; `notes.md` §55) | +| Healthcare S1+S2 on synthetic FHIR | Remainder after NEWS2 / recon / routing in code | Typed Noul/Score/Choice; S2 blinded review | Labels first; independent outcome policy | **Empirical as a named report, not clinical validation** (explore-typesafe-ai; 20 cases/scenario; Claude wrote labels; `notes.md` §55) | +| Harbor on/off routing | Same task, routing on vs off, hidden verifier | Tool Choice per turn; cheaper unsolved is not a saving | Fresh gateway; perft / checks the agent never sees | **Empirical as a *shape* and one-run signal** (jev-gateway-bench chess-bugfix 36/36 both; 4 vs 6 LLM req; `notes.md` §51). Not a measurement until reps ≥5 | | Survey scoring / psychometrics | Rubric level judgment with defined anchors | Score with concrete level descriptions; probabilities read beside every score | Weighted aggregation, reliability analysis | **Contract** (score docs: split composite judgments) | ## Search, planning & operations research @@ -42,11 +46,15 @@ judgment component is new). |---|---|---|---|---| | MCTS / PUCT | Prune invalid actions; priors P(s,a); leaf value V(s) | Batched Noul pruning + Choice priors + Score value — depth-capped where no simulator | Tree, budget, backprop, probes | **Empirical recipe** (jev-mcts: 24/24 vs 1/24 greedy; speculative depth 2) | | Beam search over taxonomies | Which branches deserve expansion | Choice distributions as branch priority; keep K paths where ambiguity is early | Frontier, budget, final selection | **Empirical recipe** (beam K=3 cookbook) | +| Structure induction over a bag | Pairwise "does i depend on j?" (or Choice over order) | One judgment per pair; DAG / scheduler in code | Topology, cycles, execution | **Empirical as a shape** (dag-jev experiment; empty README; no metrics, `notes.md` §48) | +| Collab-arm product loop | Scripted legal set vs LLM-propose vs unconstrained | Choice over legal actions; stop on low p rather than guess | Legality, Wilson/McNemar, ceiling flags | **Empirical as a harness shape** (jev-testbench; bake into jevals/Harbor, `notes.md` §48) | +| Computer-use observe → score → act | Which observed control matches the current requirement | Score / Choice / option-attention among a11y/DOM/OCR+AX candidates (Jev *or* GLiNER2 *or* Laya *or* Cua-S1); code clicks. Harness cousin: Stagehand pick then copy/act. **Closed-vote cousin:** code builds options, Jev only picks, no planner (JevOnly). **Host-owned cousin:** app retains handlers/permissions (waymode). **OCR+AX desktop cousin:** hosted Jev over numbered Vision-OCR + AX items (typesafe-computer-use). **ASR voice-browser cousin:** hosted Jev over transcript + Playwright snapshot (jev-voice-browser). Robotics cousin: Choice on geometry-as-text, not pixels | Observation, freshness, dates; never generate selectors or values; independent outcome check (`DONE` ≠ success); plan ≠ execute; kinematics / Hz in code; schema/gate else LLM; type without generation; exclusive action set; AX never sole | **Empirical recipe** as shipped loops (jev-ultrafast / solari-reflex Jev; gliner2-ultrafast GLiNER2; laya-mind2web DOM-index Laya); **Empirical as README/MODEL_CARD** for Cua-S1 (source-only, not TypeSafe Jev, `notes.md` §54); **Empirical as PR body** for Stagehand #2951–#2955 (draft; 37/75 no-LLM ~0.5s vs 4.37s *theirs*; pick ≠ replacement, `notes.md` §57); **Empirical as README architecture** JevOnly (no planner; 11/43/~$0.014/17s *theirs*) and waymode (24/26 + 34/36 *theirs*; not a self-driving proof, `notes.md` §61); **Empirical as README** typesafe-computer-use (MIT **427★**; $0.0002 / 155× *theirs* one screenshot; `notes.md` §81); **Empirical as README** jev-voice-browser (MIT **103★**; 27/27 fixtures *theirs*; `notes.md` §82); contrast blackwood-rlcd screenshot, `notes.md` §48, §52. **Empirical as showcase** MuJoCo / MOSS text-state (`notes.md` §56); drawing-pixel claim ≠ Archer | | Screening / Wald sequential tests | Pass / fail / keep-looking per candidate | One Noul gate per candidate in one batched request; budget in code | Sequential rule, stop boundaries | **Hypothesis** | | STPA / STAMP control structure | Sensor reading vs enforced constraint | Judgment as sensor; constraints in policy/code/interlock; STPA table if the sensor lies | The constraint, the actuator, the probe | **Contract** as ownership; **Hypothesis** as domain product (`mappings.md` §8) | | PufferLib / Ocean env contracts | Does this episode look like a known trainer-bug mode? | Cluster failing episodes; never "the policy is correct" | Seeded serial env, Ocean sanity, observed rewards | **Hypothesis** as placement; **Contract** that Ocean is not a comparative baseline (`formal-methods.md` DST trio) | -| Routing / dispatch (OR) | Which queue/agent owns this item | Choice + confidence-gated escalation; code owns capacity | Cost matrix, capacity constraints | **Empirical recipe** (intent-routing; LlamaIndex Jev selectors; skillranker) | -| Cascade / prefilter (IR) | Cheap reject before an expensive scorer or LLM | Per-candidate Noul/Score; fail-open on drop, fail-closed on dispatch | Candidate generation, always-keep set, recall keys | **Empirical recipe** (classifying RAG passages; jevprune; git-jev-stage) | +| Routing / dispatch (OR) | Which queue/agent owns this item | Choice + confidence-gated escalation; code owns capacity | Cost matrix, capacity constraints | **Empirical recipe** (intent-routing; LlamaIndex Jev selectors; skillranker). **Empirical as README** skillranker VOI / abstention / hook fail-open (`notes.md` §66). **Empirical as harness** BM25 vs Jev roster 50–500 (pi-jev-skill-bench; no live numbers this pass; `notes.md` §69) | +| Named decision combinators | Then / Gate / Vote / Cascade / Weighted + Router / Loop / Retry / Fallback / Memory | Compose judgments as a circuit; trace the graph | AND/OR aggregation in code; fail polarity per act; Fallback is the fail-closed node. Literal thresholded AND/OR/NOT over line Nouls is jev-semgrep (`notes.md` §86), not this metaphor | **Empirical as README architecture** (jev-combinators, renamed from decision-combinators; digital-design metaphor ≠ literal Boolean AND/OR; `notes.md` §66, §69) | +| Cascade / prefilter (IR) | Cheap reject before an expensive scorer or LLM | Per-candidate Noul/Score; fail-open on drop, fail-closed on dispatch. Decision-native: evidence-set after wide retrieve. Classify-first MCP: content to the judge without entering main agent context first. Meaning-search: packed parallel relevance, two-stage outline→zoom. Meaning-grep: AND/OR/NOT over line Nouls. Explorer: BM25 shortlist then Jev packet | Candidate generation, always-keep set, recall keys; conflict/provenance; hard I/O envelope; keyword baseline when the string is known; packet HitFile is diagnostic | **Empirical recipe** (classifying RAG passages; jevprune; git-jev-stage). **Empirical as architecture** (decision-native-rag-skills; Hypothesis as a measured win, `notes.md` §55). **Empirical as README** (jev-sift; mocks ≠ accuracy, `notes.md` §56). **Empirical as stripped-repo card** (jevgrep 79% top-5, `notes.md` §58). **Empirical as one-run** (Jev-RAG vs Spark rerank; full-context Spark still faster, `notes.md` §58). **Empirical as README + judge test** (jev-semgrep 0.94/0.98; proposition ≠ embedding; contrast-set; Semgrep.dev; not a gate; `notes.md` §61, §86). **Empirical as author-run** (jevex 1/8→6/8 n=8, `notes.md` §61; **n=16** 160s→69s / $8.74→$3.13 / 16/16, `notes.md` §72) | | Knapsack / portfolio selection | Per-item feature vector from text | Fan-out nouls/scores as features; optimizer in code | Constraint solver, weights | **Hypothesis** (mapping 1 shape) | ## Information theory & signals @@ -54,28 +62,87 @@ judgment component is new). | Method | Judgment-shaped component | Jev substitution | Stays in code | Status | |---|---|---|---|---| | Entropy as uncertainty signal | Measuring "how spread is this belief" | Entropy of returned distributions across repeats or options — computed in code from returned probabilities | All arithmetic | **Empirical recipe** (cookbook pattern) | -| Detector / Neyman filter (context) | Is this artifact relevant to the current task? | One relevance Noul per block before it enters context; stub + recall key | Cache, recall, safety keeps | **Empirical recipe** (winnow ≤0.22 hide; compaction 2-noul rule; pi-jev-context hide-not-delete) | +| Detector / Neyman filter (context) | Is this artifact relevant to the current task? | One relevance Noul per block before it enters context; stub + recall key. Encoder cousin: retention Choice + span locate, copy verbatim. Indexer cousin: GLiNER extract on the bulk, escalate LLM on the tail. Stdout cousin: Noul per chunk after a hard size/format envelope. Session-ledger cousin: score verbatim facts; rules never judged. Classify-first cousin: batch path/url/text to the judge before the main agent reads | Cache, recall, safety keeps; mutation/shell envelope; ≤10k/JSON-diff pass-through; archive dropped spans; do not dump the repo if S1 failed to load; fail-open dump of the ledger; 50 / 60k / 2MB / public-IP envelope | **Empirical recipe** (winnow ≤0.22 hide; compaction 2-noul rule; pi-jev-context hide-not-delete; gliner25-compaction GLiNER2.5 keep_full/keep_evidence/keep_call_only/drop, `notes.md` §50; s1-graphify-indexer degraded-load / no-invent-edges, 10–50× unfilled, `notes.md` §51; jev-pruner Bash stdout prune, `notes.md` §53; carryforward, `notes.md` §55; jev-sift classify-first, `notes.md` §56) | | Anomaly detection | Does this deviate from expected shape? | Guard nouls + harm Score over {input, output, tool trace} | Baselines, alert thresholds | **Empirical recipe** (guardrails cookbook; pi-jev output judge) | -| Allowlist ∩ remainder (code-then-model) | Unlisted / unstructured leftovers after a proof | Typed questions only on the unknown tier; admit iff every p < τ | Proven/refused in code; cannot block unless a sandbox sits under | **Empirical recipe** (jevgate 0/59 unsafe unasked held-out; doc-router 1.74× $). Domain-general: `mappings.md` §18 | -| Teacher distill of judgments | Copy a hosted decision API onto a small local head | LoRA / frozen-encoder heads trained on teacher answers | Independent gold labels; ECE on *your* cases | **Empirical recipe** as one 70-row run (openjev-lm 92.9%); **Hypothesis** as a general recipe | +| Allowlist ∩ remainder (code-then-model) | Unlisted / unstructured leftovers after a **proof** | Typed questions only on the unknown tier; admit iff every p < τ | Proven/refused in code; cannot block unless a sandbox sits under | **Empirical recipe** (jevgate 0/59 unsafe unasked held-out; allowlist *proves* read-only verbs; doc-router 1.74× $). Domain-general: `mappings.md` §18 | +| Decision-token LoRA (constrained-AR) | Specialize a generator for parallel constrained fields | Loss only on the single decision token; KV broadcast across fields. **1-token logprob cousin:** read top_logprobs over declared labels, no trained head (chakuho) | Schema, candidate tokens, policy; arithmetic in code | **Empirical recipe** as Foodoo1 200-case / 4-field receipt (fraud_risk 64→95%, overall 85.2→98.8%, ~234 ms); **Hypothesis** as a general recipe. Synthetic; not a financial product. Softmax ≠ Noul. **Empirical as 336-case GUI** (chakuho 27B 95%/92% vs Jev 89%/82%; coverage ≠ correctness; `notes.md` §72) | +| Teacher distill of judgments | Copy a hosted decision API onto a small local head | LoRA / frozen-encoder heads trained on teacher answers | Independent gold labels; ECE on *your* cases | **Empirical recipe** as one 70-row run (openjev-lm 92.9%); student-b n=60 MAE 0.187 / Pearson 0.791 / 90% vs vanilla (HF card unchanged ~17:48); **Hypothesis** as a general recipe | +| Domain specialist LoRA (independent gold) | Calibrated local head when policy *reads* p | Soft-target LoRA + pointer; matched-precision KL vs hosted few-shot | Argmax-only routing can stay hosted+examples; EU/thresholds in code | **Empirical as their RESULTS.md** (Domain-jev-maker; KL 0.168 vs 0.580 banking; few-shot determinate McNemar n.s.; not a teacher-copy; `notes.md` §60) | +| Active-learning triage (don't distill Jev) | Which unlabeled rows are worth an expensive label | Confidence routes accept / teacher / human; log full distributions | Real outcome labels as training targets; ECE on the student vs those labels | **Empirical as README architecture** (jev-triage; ~68% ceiling anti-pattern; `notes.md` §61) | +| Decide→policy→LLM leftover | Typed decide; leftover text only | Shared Answer schema; three Harbor arms (native / verbalized / logprob) | Policy auto/review/llm; Noul 0.5 never rounded; Score conf 0.0 never acted | **Empirical as README architecture** (jav-email-cascade; mock gen-json flat is *their mock*; `notes.md` §60) | +| Productized System One HTTP | Label + calibrated p as a public contract | Batch `{id,text}[]`; Jev primary; LLM fallback | Policy in the caller; `FALLBACK` honesty; read eval/README | **Empirical as README + eval** (classifier-dev **185★**; 400/650 ms; F1 0.887; granite 0.546 vs 0.800 *theirs*; `notes.md` §73) | +| Evidence-synthesis pointer (two-pass) | Which line answers the extraction question | Relative Choice (+ none) then absolute Noul; copy verbatim | Human tick; *Not found* / *Unclear*; Noul ≥ 0.5 *theirs* | **Empirical as README** (choxos/jev-reviewer **12★**; ≠ egma-ai; 18-q **4.6 s / $0.0101** *theirs*; spot check not a validation study; `notes.md` §74) | +| Prompted-JSON local `/v1/systemone` | Same wire, self-reported probs | Prompt → JSON vector → validate/retry → normalize + entropy confidence | Calibration on *your* labels; arithmetic/policy in code | **Empirical as README + eval** (githubnext/localjev **261★**; wire-compat ≠ logit-equiv; 1,200-req bake-off *theirs*; **≠** kunchenguid/local-jev; **≠** razorback16/openjev; `notes.md` §75) | +| Open NAR packaging + script router | Same Laya class; pick ckpt before p | Router: model= / lang= / script / default english; auto_task_detection off | Harbor cal on *your* labels; 0.85 is *theirs*; hierarchical Choice when K>20 | **Empirical as README** (NandhaKishorM/laya **710★**; T4 32.8 ms; post-T ECE 0.081; Banking77 0.425 vs Jev 0.870 *theirs*; 0.766 fine-tune; **≠** TypeSafe drop-in; `notes.md` §76) | +| External openjev census | What belongs in the class; is the list a rank? | Named census as a watch object; class-boundary GLiNER2 + routers | Frozen taskset / cal / cost / latency / silent fallback stay in Harbor; likes ephemeral | **Empirical as tweet, not a score** (@airesearch12; ~18 named; watch jev-models; **≠** jevbench v1.1; incomplete vs Laya/localjev/kev; scored sibling §78; `notes.md` §77) | +| Geometric-mean scored class bake-off | Weak-axis product over I/C/S/K; is the rank a design? | Four axes 25% each; cal ON rank; weight views published beside | Frozen taskset; native vs verbalized; partial not ranked; name ×2/est. assumptions; do not mix versions | **Empirical as v1.2 board** (Jev 75.3 / SemIf 74.6 *theirs*; Luna I=96.8 rank #7; option-order 72→21; Laya absent gap; Qwen3.8 27B ≠ Archer; **≠** tweet **≠** v1.1 87.6; `notes.md` §78) | ## Verification & logic | Method | Judgment-shaped component | Jev substitution | Stays in code | Status | |---|---|---|---|---| -| Claim–evidence entailment (NLI) | supports / contradicts / not-established per claim–source pair | One Choice per pair + review flag; judge against the cited source text only | Quote extraction, citation graph, audit log | **Empirical recipe** (citation_check cookbook) | -| Spec vs artifact conformance (model checking *mindset*) | Property holds / violated / unverifiable for a named requirement | One Noul per requirement, batched; violated → named rule back into context (pi-warden shape). This is **not** TLC/Apalache/GNATprove | Requirement enumeration, enforcement, logging; the real checker if you have one | **Empirical recipe** (pi-warden: 6→0 rule breaks, 150 paired runs; jev-pref: YOU define the rule). Ownership split: `formal-methods.md` | +| Claim–evidence entailment (NLI) | supports / contradicts / not-established per claim–source pair | One Choice per pair + review flag; judge against the cited source text only. Stop-hook cousin: claims vs **session** evidence | Quote extraction, citation graph, audit log. When the answer *is* a span you already hold, **point** at line ids and copy verbatim — the model never writes the excerpt (**choxos/jev-reviewer**, ≠ egma-ai; human tick never overwritten). Keyword retrieve is not semantic (clear-head) | **Empirical recipe** (citation_check cookbook; choxos sample study 18-q **4.6 s / $0.0101** *theirs*, `notes.md` §48, §74; clear-head Stop, `notes.md` §51). Atlas: paraphrase_support and reversed_meaning_high_overlap correctly judged when both texts are in `state` (`notes.md` §49) | +| Extractive selection + offline re-threshold | Keep/drop over sentences, ids, character offsets, observed DOM controls, or stdout chunks code already holds | Per-item Noul/Choice + one broadcast; join in order; `redecide` on the log with no new calls. Compaction: retention Choice + span locate. Computer-use: score among a11y/DOM candidates. Stdout prune: Noul per chunk after hard envelope. Harness extract: pick elements, copy text. **Evidence-synthesis two-pass:** relative Choice then absolute Noul | Numbering, header skip, thresholds, publish permission; mutation envelope; copy exact bytes; click in code; never generate selectors; archive dropped stdout; schema/gate else LLM; *Not found* is an answer; human check is the product | **Empirical recipe** (testimonial-miner 8-request fixture; **choxos/jev-reviewer** **12★** / two-pass / Cochrane-PRISMA, `notes.md` §48, §74; gliner25-compaction char-offset copies, `notes.md` §50; gliner2-ultrafast observe→score→act, `notes.md` §52; jev-pruner, `notes.md` §53). **Empirical as PR body** Stagehand #2955 pick-and-copy, `notes.md` §57. Cousin of applied-mappings §2 | +| Combinatorial grid / program synthesis | Consistent whole-object from many cells | **Rejected as extractive.** Cell-wise Choice does not assemble ARC grids (4/400 Direct Jev) | Search, a program, a simulator | **Empirical as a negative** (`notes.md` §49) | +| Spec vs artifact conformance (model checking *mindset*) | Property holds / violated / unverifiable for a named requirement | One Noul/Score per requirement, batched; violated → named rule back into context (pi-warden / Abide shape). This is **not** TLC/Apalache/GNATprove | Requirement enumeration, enforcement, logging; the **linter** if the rule is lintable; the real checker if you have one | **Empirical recipe** (pi-warden: 6→0 rule breaks, 150 paired runs; jev-pref: YOU define the rule; Abide: productized compile/calibrate/tune/replay, `notes.md` §47; if-ai: plain-English PR check, fail-closed on error, `notes.md` §51). Ownership split: `formal-methods.md` | +| AST ∩ semantic lint | Semantic remainder after a parser already extracted units | Typed questions on Tree-sitter targets; do not execute scanned code. Skills→oxlint: generic AST facts + prechecks, then remainder Noul; guidance whole-file in state | Parser, selection, fail-on; `tenbin` owns the lint *skill* | **Empirical as a shape** (jevscan 0.2.0rc4; not a calibration claim; `notes.md` §48). **Empirical as Phoenix experiment** (jev-oxlint; answer-key agree; not a hard gate, `notes.md` §58) | | Alloy finder vs Apalache / TLC | Which bound, which counterexample, is the property tautological? | Triage instances/CEs; never "this spec looks right" | Analyzer / SMT / explicit-state engine | **Hypothesis** as product; **Contract** as ownership (`formal-methods.md` §2) | -| Type-checking analog | Does this planned call match the schema/operation/target? | Decomposed nouls over {request, schema, trace}; never trust a Jev pass as authorization | Real validation of operation+target in code | **Empirical recipe** (validation.md self-monitoring) | +| Type-checking analog | Does this planned call match the schema/operation/target? | Decomposed nouls over {request, schema, trace}; never trust a Jev pass as authorization | Real validation of operation+target in code | **Empirical recipe** (validation.md self-monitoring). Product: toolgate allow/block/review — Jev is not authorization; timeout stops (`notes.md` §55). Capability kernel: interlock — secrets never in agent; Jev SENSOR; policy.py decides; type-safe ≠ correct (`notes.md` §59). Human-confirmed kill: port-cleanup (`notes.md` §59). Actiongate / turnstile: evidence ≠ authority (`notes.md` §64, §66). SEAL: no seal, no advance (`notes.md` §67) | +| OOD calibration / sign by type | Honesty of p on an unseen rule | ECE with noise floor; refit T; per-type breakdown | Org-policy labels in code, not in the prompt | **Empirical as 900-ticket table** (jev-ood-calibration; AUC ≠ ECE; `notes.md` §66) | +| Sureness over a probability vector | Act / escalate from the *shape* of p, not one p | max_prob / margin / entropy / gini / perplexity in code | Band thresholds; never treat max_prob as the strictest | **Empirical as 60-q library** (how-sure-is-jev; Choice confidence = max_prob; `notes.md` §67) | +| TLA+ / model-check around an oracle | Protocol invariant under crashes / unstable votes | Decision model as noisy `consult()`; escalate on no quorum | Spec, quorum bound, stability floor | **Empirical as TLC + 1,080 golden** (jev-labs; never confidently wrong; `notes.md` §67) | +| Advance / coverage ledger | May the world change? | Strike is optional; Seal is the gate; coverage.path visible | Effects locked while escalations open | **Empirical as README** (seal; mint ≠ product brain; `notes.md` §67) | +| Pre-review typed CI gate | auto-approve / human-review / block | Four cheap questions before expensive review | Operator thresholds; never paste secrets in comments | **Empirical as own-repo latencies** (ci-gatekeeper-bot-jev; `notes.md` §67) | +| Host-adapter capability rank | Which supplied tool/skill next | Nouls over a caller-built catalog; none is first-class | Host decides; lexical fallback | **Empirical as README** (jev-in-codex; ranking unbenchmarked; `notes.md` §67) | +| Receipts-not-leaderboard map | Where the class holds vs breaks | API receipts; not a ranking | Retrieve first; pin version | **Empirical as axis** (atlas §49) **and eval-integrity cluster** (atlas 10★ + frontier-100 + ood; `notes.md` §66) | +| Engine truth ∩ coaching judgment | Severity / error class / interrupt on engine facts | Typed questions; never re-evaluate the position | Engine eval/lines/swing; templates; capped writing model | **Empirical as PRD** (game-coach Wave 0; Stockfish owns truth; `notes.md` §59). Anti-soundness-theater with egma attention≠correctness | +| Native-probability calibration | Honesty of noul/choice/score vs analytic truth | Brier / ECE / reliability / risk-coverage; fan-out batches | World definitions; oracle stub; teeth miscalibrated stub | **Empirical as their live card** (jev-arena Brier 0.0059 / ECE 0.0620 on 145 noul; overconfident in low bins; `notes.md` §59) | +| Ranking family on soft scores | Pairwise inversion / Score ordinality / ties of a sort key | Gate SQL on a passing results.json; secondary key on two-decimal ties | The engine's LIMIT/tie break; request-shape measurement | **Empirical as independent measurement** (jev-orderby-bench; Score 0.143 weak link; 53-way 0.99 tie; recodelabs batch-40 fails ranking; `notes.md` §60) | +| Constrained-AR PCD vs decision model | O(1) schema-valid decode vs calibrated Noul | Same labels, report Brier/ECE/acc/latency | Schema enforcement; n small is a caveat | **Empirical as their n=50 table** (system-one-benchmark; Jev 84.0% / Brier 0.1096 vs PCD 52% / 0.3884; `notes.md` §61) | +| Operator-owned approval criterion | Whether to suppress a human prompt on a gated call | Parallel verdict + severity + in-scope; presets measured for prompts removed *and* unsafe auto-approvals | Host deny rules; plugin never self-tunes the bar | **Empirical as measured traffic + corpus** (omp-greenlight; 1,013/10; default 40.9% / 0 of 94; not a sandbox; `notes.md` §62) | +| Pre-agent skill grant | Which authorised skills to inject this turn | Jev relevance/confidence over a code-built candidate set | Catalog, policy, limits, grants; foundation-only on failure | **Hypothesis / outline** (skill-broker; Jev never grants access; not a production recipe; `notes.md` §62) | +| Eval-instrument audit | Data/scorer/run/claim soundness, not the headline score | `dinostomp jev` if-statement hygiene (accuracy, p(yes) cut, ECE, blank lean, rewording) | Reproduction commands; self-findings ledger | **Empirical as FINDINGS.md + demo card** (dinostomp; 189 entries / 99 against itself; ECE 0.062 *theirs* on 24; beside jevals, not a Harbor taskset; `notes.md` §62) | +| Constrained min-cost s.t. SLO/quality | Soft semantic features for a backend picker | Jev task/exactness/evidence; fail-open local features | Health/context/tools/quality/SLO floors; exactness never overrides capability | **Empirical as live analysis *shape*** (slo-router; p95 77.93→490.38 same routes; eight-row demo not a benchmark; `notes.md` §63). **Hunch:** S1 on the feature side, never the sole hot-path gate | +| Effect-based command gate | Blast radius / reversibility, not privilege tokens | Choice allow/deny + independent risk Nouls | Fast-allow/deny; operator-owned minConfidence/riskThreshold; fail-closed | **Empirical as certification** (construct-auto-classifier; Jev 0/975 dangerous; chat leaked 16–104; `notes.md` §63). **Hunch:** contracts on effects, not tokens | +| Human-review attention filter | Do I need to look / which files / debris | Per-file Nouls; never green unless sure | Never blocks the agent; never edits; shadow mode | **Empirical as README architecture** (jev-lens + jev-lens.nvim; `notes.md` §63). **Hunch:** VOI for human review, not a permission gate | +| Evidence-gated question pack | Is this criteria set measured on a pinned version | Record accuracy/ECE/cost/latency; `unknown` on every Choice/Score | Pin model; no numbers → not `verified` | **Empirical as registry + runner** (jev-packs nine packs *theirs*; **jevassert LANDED** record/replay CI; 2,990-case matrix: Jev/Sonnet 5 accuracy tie, Jev better calibrated 7/9, ~250× cheaper; sms-spam this-pass 0.953/ECE 0.040; `notes.md` §64, §70). **Hunch:** measurement owns endorsement | +| Record/replay CI | Does a recording still pass accuracy/ECE/Brier/cost/latency gates | Offline from recordings; exit 0/1/2; McNemar compare | Pack SPEC; `unknown` mandatory; cost priced at check | **Empirical as README + Action @v0** (jevassert Apache-2.0; PyPI; backend-neutral; `notes.md` §70) | +| Failure-finding arena | Which questions does the model get wrong | Pairwise JudgeBench-shaped protocol; BYOK | Not a leaderboard; harness ≠ findings | **Empirical as runnable harness** (chenmingtang830/jevarena ≠ meetr1912/jev-arena; `notes.md` §70) | +| BBQ stereotype/uncertainty | Ambiguous vs informative stereotype alignment | Accuracy + BBQ bias + cost/latency on full 58,492 | Dataset CC BY 4.0; not a general bias cert | **Empirical as full-set card** (jev-bbq-experiment 97.28%/0.04/0.34/$0.3429 *theirs*; `notes.md` §70) | +| Sentence-as-rule lint | Does this AST subject contradict a declared sentence | ast-grep `rule:` × Jev `ask:`; Score or Noul | Matcher silent-fail; Jev loud; fail-open no-verdict | **Empirical as 13/15 corpus** (mizchi/jevlint 1.00/1.00 *theirs*; ≠ huntedman/JevLint; `notes.md` §70) | +| Decider ≠ executor loop | Which tool next / progress / risk / done | Jev owns those; LLM fills args only | Risk≥0.5 pause; stuck ladder; tool execution | **Empirical as README** (jeffrey; pick ≠ fill; `notes.md` §70) | +| Decider ≠ executor across timescales | Which flight action vs which strategy | S1 Choice every tick; S2 one-use advice | S1 never stalls; physics owns collisions; S2 never grants | **Empirical as README delta** (khordoo/jev-reflex-autonomy-lab; Local ≠ localjev; 20% still soft; `notes.md` §80) | +| Persist constraints across compaction | Does this call violate a user constraint that survived the window | Jev KEEP/LIFT/… on receipts; never writes policy | Structured ledger; replay without Jev; fail-open | **Empirical as 79-session bench** (pi-heed 98.5%/0 false block *theirs*; `notes.md` §70) | +| Harbor SGR-judge contract | Does evaluation need generation | Frozen Jev vs schema-guided LLM judges; invalid = FN; cost/latency | Incomplete cohort ≠ headline; canaries ≠ quality | **Empirical as contract, not a score** (jev-judge-bench; 21 offline tests; **no quality headline yet**; ≠ jevarena/jevbench; `notes.md` §71) | +| Hand no-text steps | Did it work / which next / safe | Batched Choice/Noul/Score; writing stays generated | Fail-open gate; reconstruct confidence if the gateway drops it | **Empirical as 95-call card** (jev-use 220 ms p50 / 12/12 / Vercel 0.4; ≠ jev-ultrafast; `notes.md` §71) | +| Pi System-One control plane | Route / gate / retry / sieve / review / click | Named sensors; fail polarity per act | Compaction never writes session; GUI never force-click | **Empirical as README** (pi-jev-control; license null; no live quality numbers; `notes.md` §71) | +| Generation as a tree of Choices | Next word without free generation | One typed question per choice over a closed lexicon | Embedding tree in code; rank texts last | **Empirical as README demo** (jev-gpt ~400 calls / 75 s / 2¢ *theirs*; license null; `notes.md` §71) | +| Recipe atlas | Which narrow questions fit this job | Code prepares; policy in code; samples not benches | Numbers/dates exact; review band; pick don't extract | **Empirical as 15 recipes** (jev-cookbook 425/$0.015; 16–36 not benches; `notes.md` §71) | +| Personal-history ranking | What to show next from *your* trail | One request per batch of ten; distribution *is* ranking | History local; no social graph | **Empirical as README** (jevfeed; 17 tests; `notes.md` §71) | +| Competing NAR claim-audit | Is this open head "better calibrated"? | Like-for-like ECE channels; n and CI; throughput ≠ latency | Dual-channel is a design fork, not a 15× badge | **Hypothesis until independent run** (openJev-verdict-2.0 + PR #1; ≠ IamBusy/OpenJev; `notes.md` §71) | +| Evidence ≠ authority | May a proposed tool call execute | Six narrow semantic questions; never one "is this safe?" | RBAC/schema/limits; positive p cannot override a hard fail | **Empirical as README slogan** (actiongate-jev; 500-case is label-baseline not accuracy; `notes.md` §64) | +| Ranking ≠ calibration | Do stated probabilities match frequencies | AUC vs ECE/Brier vs human rates; Platt/isotonic on held-out | Domain labels; do not cite ECE alone (gameable) | **Empirical as 8,000-judgment audit** (does-jev-confidence; stated ~75% vs human ~10%; ~96% ECE removed; jevcal ~100 rows; `notes.md` §64) | +| Hot-click observe→score→act | Next operation + target from an indexed viewport | One request, speculative target heads; compatible-only options | Observe/execute/stale-ref/`--until` in code; generator only for type | **Empirical as n=3 medians** (ego-jev; ~2× vs per-step LLM; high variance; not a bench; `notes.md` §65) | +| OCR+AX desktop observe→score→act | Exclusive kind / item / site among numbered OCR+AX items | Hosted TypeSafe Choice; writer only for free text; post-type Noul | Crop+tile OCR, dates.py, AX bonus never sole; never generate selectors; 0.4 / 0.5 still soft; `done` ≠ success | **Empirical as README** (typesafe-computer-use MIT **427★**; $0.0002 vs Opus $0.032 155× *theirs* one screenshot; **≠** jev-ultrafast **≠** cua-s1 **≠** camoufox; `notes.md` §81) | +| ASR voice-browser observe→score→act | Intent / target / complete / is_command / destructive among snapshot ids + regex spans | One 9–11-question request; pointer-not-generator | Playwright; debounce; overlay numbers exact; spoken confirm ≠ auth; 0.5 / 0.55 / 0.6 still soft | **Empirical as README** (jev-voice-browser MIT **103★**; 27/27 fixtures / ~300 ms / ~$0.0002 *theirs*; **≠** jev-voice-control **≠** nikolas-j **≠** OCR §81; `notes.md` §82) | +| Wrap-as-execution ALLOW/ASK/DENY | Serves-intent / safe / consequential among a proposed tool call | Rules first; Jev remainder; ASK throws | failMode closed; judge swappable; hosted tools out of reach; AUTO_APPROVE is a hatch | **Empirical as README** (AgentGhost MIT **2★**; **≠** actiongate **≠** toolgate **≠** jev-use **≠** jwen5419807; `notes.md` §83) | +| JP application genre atlas | Which hole is this app? | Named genres, not a rank | Stars research-time; not verified evals; likes ephemeral | **Empirical as tweet, not a score** (@studio_yebisu; **≠** @airesearch12 class census **≠** v1.2 board; `notes.md` §84) | +| External pedagogy / how-to-apply | Where does typed judgment beat an LLM hammer? | Code owns branches; parallel questions; thresholds in code | schema-safe ≠ correct; 200×/400× TypeSafe ceiling; shadow first | **Empirical as article, not a bench** (@akshay_pachaar “Jev Clearly Explained”; **≠** official docs **≠** Flavio **≠** AgentGhost; `notes.md` §85) | +| Verbatim compact + same-pass safety | Should this message stay; is the pending act destructive | Keep/drop Choice + Foreman Nouls; regex floor independent of Jev | Pins, tool-pair integrity, never rewrite; dual fail polarity | **Empirical as product-arm table** (jev-compactor later 73%/350ms/4 of 4 vs shipped summarizers; `notes.md` §68) | +| Pre-send view selection | Smallest view that still serves the next step | Code-built outline/focus/testlog/…; Jev picks; recall restores | Compress before first send; code full unless confident | **Empirical as 500-trajectory** (dizk/jev-lens 79% fewer tokens; `notes.md` §68) | +| Observational keep/kind | Keep this candidate? What kind? | Verbatim ledger; kind-keyed files; model-free compact | Failed Jev does not drain buffer | **Empirical as README** (pi-observational-memory-jev; `notes.md` §68) | +| Attention redirect (not block) | Eight risk axes on a turn diff | Assist = one reinspect; fail-open | Uncalibrated 0.85; not a merge blocker | **Empirical as owner-run smoke** (jev-preflight; `notes.md` §68) | +| Local-rules ∩ remainder + anti-self-train | Collapse junk after a free prove pass | `rules.js` then batched Nouls; confirm-queue before examples | Collapse not delete; human confirm; never train on own hides | **Empirical as README + small e2e** (x-reply-filter; 0.90/0.93 vs 0.08/0.10; `notes.md` §65) | ## Economics & game theory | Method | Judgment-shaped component | Jev substitution | Stays in code | Status | |---|---|---|---|---| -| Multi-criteria decision analysis | Attribute scores per alternative | Composite scores with weights in code (re-weight without re-inference) | Weight policy, Pareto views | **Empirical recipe** (mapping 1) | +| Multi-criteria decision analysis | Attribute scores per alternative | Composite scores with weights in code (re-weight without re-inference). Intent columns: a heading scores each row | Weight policy, Pareto views, exact fields | **Empirical recipe** (mapping 1). Showcase *shape*: jevable intent columns; snack MCDA clocks are **claims** (`notes.md` §56) | | Mechanism/game response (adversarial state) | Opponent-intent / bluff / risk read on a state | Choice over reads + risk Score; policy in code; 2.5Hz-style advisory rate | Strategy solvers, exploitative math | **Empirical recipe** (jev-trader, game agents) | | Auction/market event classification | Is this signal material? Direction? | Choice over event classes + urgency nouls; execution in code | Order execution, risk limits | **Empirical recipe** (jev-trader 81ms/block) | -| Assignment / scheduling (OR) | Soft affinity per pair | Score/Noul as a cost *feature*; solver owns capacity/legality | ILP/heuristic, fairness | **Hypothesis** (`mappings.md` §15) | +| Threshold-probe value CDF (Vickrey) | P(V > X) noul fan-out; code bids | Native probabilities; monotonize CDF; Jev never bids | Second-price / first-price payment rule | **Empirical as their live/offline cards** (jev-vickrey; oracle regret 0; live Brier 0.1391; overconfident stub loses money; `notes.md` §59) | +| Assignment / scheduling (OR) | Soft affinity per pair | Score/Noul as a cost *feature*; solver owns capacity/legality | ILP/heuristic, fairness | **Hypothesis** (`mappings.md` §15). **Empirical as *shape*** (slo-router constrained min-cost s.t. quality+SLO; Jev features not the sole gate; `notes.md` §63) | | Situated software (Shirky) | Local meaning inside a named group | Dense full-traffic judgment only inside a named community boundary | Public metrics, law, FM where wrongness is intolerable | **Hypothesis** (`mappings.md` §16) | ## Rejected (standing, so they aren't rediscovered) diff --git a/.agents/skills/augustus/references/mixed-architecture.md b/.agents/skills/augustus/references/mixed-architecture.md index b45bd76..c415e93 100644 --- a/.agents/skills/augustus/references/mixed-architecture.md +++ b/.agents/skills/augustus/references/mixed-architecture.md @@ -31,7 +31,10 @@ If the request is "how do I call Jev?", stop and load `typesafe-ai`. If it is "should this step be a judgment-class model, an LLM, a regex, or a trained classifier — and which family?", stay here. Family table: `judgment-class.md`. Cross-domain frames: `mental-models.md`. Proof vs -judgment: `formal-methods.md`. +judgment: `formal-methods.md`. [wellposed](https://github.com/suraj-phanindra/wellposed) +is an Empirical *recipe* of tenbin's lint hole (missing `other` → +confidence 1.00 wrong); it is not a second Augustus skill +(`notes.md` §46). ## Default architecture @@ -75,6 +78,42 @@ from the page; Jev picks among visible elements; code enforces prices/dates; answers are *located*, never composed. The day the task needs a paragraph, an LLM re-enters — that is mixed architecture, not a different religion. +**Component node, not internals-as-FSM.** Place the judgment model as a +node in *your* state machine; do not describe its internals as one. +Atlas teaching (`notes.md` §49): paraphrase_support and +reversed_meaning_high_overlap are distributed LM understanding, not +keyword rules. Code still owns transitions (`mappings.md` §3). + +**Browser-use strength is DOM-as-text + speculative fan-out, not +vision.** Translate a screenshot task into a structured DOM snapshot +as `state`; fan out over candidate elements in one call. Text-only +models then sit in their strong zone (extractable from fed state). +Same placement as lizard-agent / solari-reflex; not blackwood +screenshot-in. **Backend-agnostic this hour:** +[gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) +is that hole with local GLiNER2 instead of Jev — observe → score +among a11y/DOM candidates → code acts (`notes.md` §52). Hybrid: +local decide; remote fill only for TYPE. `DONE` is not verified +success. + +**Harness productization this hour (draft Watch):** +[Stagehand #2951–#2955](https://github.com/browserbase/stagehand/pull/2955) +puts the same DOM-as-text loop inside Browserbase Stagehand: Jev +picks; code copies or acts; LLM fallback. Extract pick-and-copy is +a fast path, not a replacement (`notes.md` §57). Do not copy the +opt-in flag. + +**Dual-process cascade (Kahneman productized; routing accuracy +unmeasured).** +[`taro1985/dual-process-ai`](https://github.com/taro1985/dual-process-ai) +(MIT): S1 typed decision + confidence; S2 generates only when +`confidence < τ`. Routing **fails open**; safety **fails closed**. +Degraded keyword mode without a key is **not** equivalent S1. Crossover +for business/life, not only SWE: cheap classify/route/gate on S1; +write/reason on S2. Same split as jev-reflex-autonomy-lab (S1 keeps +control) and jev-hermes (route ≠ memory). Do not copy hooks. Tune τ +on *your* escalation log. + ## It's just classification Canonical answer: `references/faq.md`. Short form: classification is not @@ -151,8 +190,150 @@ not a global virtue: | Action | Typical failure policy | Why | |---|---|---| | Drop a RAG chunk or log line | **Fail open** (keep on error) | A false drop loses evidence; a false keep costs tokens | +| Compact / drop a completed tool result | **Fail closed** to keep-full (`gliner25-compaction`) | Compaction is a destructive edit of memory. Uncertain *looks* like keep-on-error from the evidence side; name the *reduction* as the act. Contrast Abide / jevgate fail-open | +| Omit a session-memory fact from the brief | **Fail open** (dump the whole ledger) (`carryforward`) | Scoring failure must not hide a rule. Constraints/corrections always return; Jev never votes on them | +| Open a file / URL into agent context | **Fail open** (closer look on error / truncation / uncertain) (`jev-sift`) | False drop loses evidence. Errors and truncation are not irrelevance. Transport failure ≠ "irrelevant" | +| Withhold a draft because the judge is silent | **Fail open** (heartbeat / proceed or escalate; do not hold forever) | Missing verdict is not a block and not a pass. Contrast Abide `<0.5` silence (the *edit proceeds*) | +| Prune Bash stdout before the LLM | **Fail closed** to original (`jev-pruner`) | Dropping the log is irreversible. ≤10k / JSON-diff-whole-doc prove pass-through; archive/Jev/incomplete-score failure keeps the result. Harbor plugin-eval cannot reach Jev → cannot prune | +| Skip waking a sleeping agent | **Fail open** (wake on error / unsure / no key) (`wakegate`) | Skip is the irreversible act. User-message, skip-limit, nothing-to-judge, and p in 0.2–0.5 all wake. Contrast pi-jev-approver fail-closed without a key | +| Merge a red CI run | **Fail closed** on `--gate` (`latch`); reporter stays fail-open | False PASS merges a real bug. Missing key never fails Playwright; the gate is a separate step. Judge never says ignore alone | +| Plain-English PR check | **Fail closed** on error / empty / low confidence (`if-ai`) | A skipped or timed-out check is not a pass. Threshold is policy, not measured correctness | | Route to a tool / start a side effect | **Fail closed** (don't call) | A wrong tool is an action | -| Rerank a retrieved list | Fail open: keep retrieval order (`WiktorB2004/llama-index-jev`, **Empirical recipe** on BEIR nfcorpus: MiniLM 0.340 nDCG@5 → MiniLM+Jev 0.396; rerank fails open, *select* fails closed). Listwise/cross-encoder scores belong here, not on the row above. | Ranking errors are quality; selection errors are control-flow | +| Execute a proposed tool call | **Fail closed** on block / timeout / guard error (`toolgate`) | Execution is the irreversible act. `review` needs authenticated human approval, not self-approval. Jev is not authorization. Distinct from ndolinschi allow/ask_human/deny *vocab*. Distinct from **capability kernel** (`interlock`): secrets never in the agent; closed action space; Jev is SENSOR; `policy.py` BLOCK/ASK/ALLOW | +| Wrap tool execution as ALLOW/ASK/DENY (`AgentGhost`) | **Fail closed** (ASK/DENY throw; judge error → DENY) | The wrap *is* the execution function; the model cannot skip it. Rules first; `allow` skips the judge. HITL cannot be silently skipped. **≠** actiongate (policy owns authority) **≠** jev-use (fail-open) **≠** advisory sidecar. rh-guard owns the gate cousin | +| Treat `AGENTGHOST_AUTO_APPROVE=1` as a grant | **Fail closed** (demo hatch) | Unattended examples skip `y/N`. Not authorization. Provider-hosted tools stay out of reach | +| Collapse reddpy/AgentGhost into jwen5419807 / vventirozos / actiongate / toolgate | **Fail closed** (qualify the wrap) | Intent-aware Jev wrap, MIT **2★**. **≠** 95★ 2025 namesake **≠** FastAPI swarm **≠** evidence-not-authority slogan | +| Treat a JP genre atlas as a bake-off / paste its ★ | **Fail closed** (list + research-time stars) (`studio_yebisu`) | Tweet *theirs*: not full eval. Stars drift (typesafe-computer-use 203→427; jev-voice-browser 40→103). **≠** @airesearch12 class census **≠** v1.2 board | +| Quote Akshay’s 200× / 400× as a measured win | **Fail closed** (TypeSafe ceiling) (`akshay_pachaar`) | Article *theirs*: treat as a ceiling, not a promise. 70–500 ms / $0.042/MTok are vendor envelope. Not Harbor | +| Treat topical cosine as “customer is asking” | **Fail closed** (proposition ≠ embedding) (`jev-semgrep`) | Contrast-set: all six about refund; only asking pass. Angry agent vs angry customer cosine ~1 | +| Multiply parallel meaning Nouls / negative-query tricks | **Fail closed** (boolean after threshold) (`jev-semgrep`) | AND/OR/NOT in code on bits. Do not multiply p. ≠ jev-combinators metaphor | +| Treat jev-semgrep as Semgrep.dev or as a gate | **Fail closed** (name collision; not a gate) | Rename if both installed. Ranking fail-open. rh-guard skip | +| Treat “cannot hallucinate” without the schema caveat | **Fail closed** (schema-safe ≠ correct) | Cannot invent out of schema; can pick the wrong valid option. Safer *theirs*: cannot break the declared schema, but can still be wrong | +| Treat the explainer as a TypeSafe how-to / copy its Python | **Fail closed** (independent pedagogy) | Placement theses, not an SDK tutorial. **≠** official docs **≠** Flavio Copes **≠** LangChain harness **≠** AgentGhost wrap | +| Kill a live process / port | **Fail closed** to human confirm (`port-cleanup`) | Jev recommends; human is the only trigger. Identity re-check before signal; shields override; mapped explanations, not raw model prose. Kill recs need conf ≥ 0.8 | +| Actuate an observed browser control | **Fail closed** (code validates the node) | Freshness / visibility / disabled / occlusion in code; model never emits selectors (`gliner2-ultrafast`, jev-ultrafast, solari-reflex). `DONE` does not authorize "success". Cua-S1: dry-run default; `execute`/`submit` opt-in; fail-closed unknown checkbox; fill execution fails closed without token `set_value`. Stagehand: same; LLM fallback when Jev abstains. **typesafe-computer-use:** AX press when the item came from AX; mouse fallback; off-screen refusal is a no-op | +| Ship a screenshot to frontier for the *decision* | **Fail closed** (OCR+AX text-state) (`typesafe-computer-use`) | Decision never sends pixels. The one-shot **answer** reader may receive the capture — that is a writer packet, not the Choice. Not omni. Skip Archer | +| Offer overlapping CU action options | **Fail closed** (exclusive set) | Confidence measures concentration; overlap reads as doubt (*theirs*). wellposed cousin. Split `kind`/`item`/`site`/`offscreen` | +| Treat $0.0002 / 155× as a Harbor score | **Fail closed** (one screenshot) | Same screenshot, one decision each *theirs*. Dates.py caveat. Re-measure on *your* taskset | +| Treat `--min-confidence` 0.4 or post-type Noul 0.5 as Harbor τ | **Fail closed** (still soft) | Product copy. schema-safe ≠ correct. Mouse-slam / Accessibility deny are the exact stops | +| Collapse typesafe-computer-use into jev-ultrafast / cua-s1 / camoufox | **Fail closed** (qualify the host) | macOS OCR+AX + hosted Jev. **≠** browser DOM **≠** Cua-S1 **≠** OmniParser loop **≠** Camoufox clone | +| Treat AX as the sole source / mix off-screen into visible items | **Fail closed** (bonus source; separate question) | Spotify 0 *theirs*. A mouse click would land on the wrong pixel | +| Treat `done` as verified success | **Fail closed** (loop termination) | Writer answer is a reader packet. Dry-run prints no answer | +| Ship a waveform / screenshot to Jev because the UI is voice | **Fail closed** (transcript text-state) (`jev-voice-browser`) | ASR is the producer. Jev sees the schema, not audio. Compose with OCR §81. Skip Archer | +| Treat spoken "confirm" as authorization | **Fail closed** (convenience, not a guarantee) | README *theirs*. Anyone who can reach the control port drives the browser. Judgment ≠ permission | +| Truncate free-text on a closed-set wait | **Fail closed** (wait policy is VOI) | Closed-set may act on a partial; search/type wait for final or 600 ms silence *theirs* | +| Call a second model to pick among numbered overlays | **Fail closed** (UI number is exact) | Spoken digit copies an id code already holds | +| Quote 27/27 or $0.0002/call as a Harbor score | **Fail closed** (fixtures) | Integration on captured pages *theirs*. Re-measure. 0.5 / 0.55 / 0.6 still soft | +| Collapse moritzkremb/jev-voice-browser into jev-voice-control / nikolas-j / typesafe-computer-use | **Fail closed** (qualify the host) | Headed Chromium + Web Speech + hosted Jev. **≠** macOS stub **≠** 0★ namesakes **≠** OCR desktop | +| Replay a cached browser action | **Fail open** on the freshness check (`stagehand` cacheCheck) | Errors/timeouts never block replay; a stale verdict re-infers. Opt-in: the check costs a snapshot + a request | +| Skip the LLM on extract / act | **Fail open** to the generator (Stagehand pick/judge) | Schema/gate/screenshot envelope in code; pick is a fast path, not a replacement. Invalid extract → LLM | +| Publish a public wall ask without a key | **Fail open** (allowlist; UI says Jev offline) (`ask-jev-ai`) | Missing judge is not a block and not a silent pass. Safety p≥0.6 still blocks when the judge is on | +| Assign PR attention P2 | **Fail closed** incomplete → `uncertainPriority` never P2 (`egma-ai/jev-reviewer`) | Attention ≠ correctness. Deterministic `alwaysReviewPaths` P0. Do not treat a Noul as a proof the PR is good | +| Pick a session model | **Fail closed** to a declared standard (`jev-adaptive-thinking`) | Timeout / no session / missing first-round text locks `gpt-5.6-sol`. Contrast jev-gateway fail-open passthrough if Jev is down | +| Drop a meaning-search hit | **Fail open** as ranking (`jevgrep`) | False drop loses the file. Keyword tools still win exact strings | +| Rerank a retrieved list | Fail open: keep retrieval order (`WiktorB2004/llama-index-jev`, **Empirical recipe** on BEIR nfcorpus: MiniLM 0.340 nDCG@5 → MiniLM+Jev 0.396; rerank fails open, *select* fails closed). Listwise/cross-encoder scores belong here, not on the row above. One-run cousin: Jev-RAG vs Spark *rerank* (full-context Spark still faster) | Ranking errors are quality; selection errors are control-flow | +| Auto-act an email / ticket | **Fail closed** to review when Noul ≈ 0.5, Score conf = 0.0, or a hard flag fires (`jav-email-cascade`) | Noul 0.5 is cannot-tell, never rounded. Injection always review. LLM leftover is optional | +| `ORDER BY prob LIMIT k` | **Fail open** as ranking; ties need a secondary key (`jev-orderby-bench`) | Two-decimal quantization; 53-way 0.99 tie is engine-dependent. Calibration ≠ sortable | +| Run an irreversible browser/OS act in closed-vote CU | **Fail closed** unless a separate risk vote is low (`JevOnly` `risk ≥ 0.50` never default; waymode host confirmation / p ≥ 0.7) | Code builds options; Jev only picks. `completed` ≠ verified success. Host handlers/permissions still decide | +| OMP/pi acceptance gate or subagent route | **Fail open** if Jev missing / timeout / malformed (`omp-jev-extensions`; `confidence: 0`) | A flaky decision service must not trap the agent. Contrast pi-jev-approver fail-closed without a key | +| Suppress an OMP tool-approval prompt | **Fail closed** to prompt the human unless the operator-owned bar says allow (`omp-greenlight`) | Not a sandbox. Host `bash.patterns: deny` stays the floor and fires first. Plugin never self-tunes the bar. 0/94 is the labelled corpus, not live traffic | +| Grant a specialised skill pack | **Fail closed** to foundation-only; never broaden access (`skill-broker` outline) | Jev scores relevance; code owns grants. Candidates ≠ grants. **Hypothesis / outline — not a production recipe** | +| Treat a Jev score as eval truth | **Fail closed** until the instrument is audited (`dinostomp`) | Check data/scorer/runs/claims, not just the number. `dinostomp jev` tests a question like an if-statement | +| Pick an LLM backend with live Jev on the hot path | **Fail open** to local deterministic features (`slo-router`) | Jev is a feature, not the sole gate. Same routes/accuracy on their fixture; p95 **77.93 → 490.38 ms**. Exactness raises the quality floor; never overrides capability. Eight-row demo is not a benchmark | +| Allow a proposed shell command | **Fail closed** on missing / low-conf / high-risk / failed call (`construct-auto-classifier`) | Privilege ≠ verdict. Landed-script trust is a merge-gate receipt, not a name. Headless escalation is deny-and-report, not auto-approve. 0 dangerous / 975 *theirs* | +| Tell a human the agent work is green / skip | **Fail closed** to "look" unless sure ([rashedInt32/jev-lens](https://github.com/rashedInt32/jev-lens)); **never block** the agent | Attention filter / VOI, not a permission gate. Never edits files. Distinct from dizk/jev-lens (pre-send views) and from jev-gates (stops writes) | +| Endorse a question pack | **Fail closed** until recorded evidence (`jev-packs` + landed `jevassert`) | No numbers, no `verified`. Pin model version. Abstention/`unknown` mandatory. `check` runs **offline from recordings** (exit 0/1/2) | +| Treat a public arena as a leaderboard | **Fail closed** until measured findings (`chenmingtang830/jevarena`) | Failure-finding, not crowning winners. Distinct from `meetr1912/jev-arena`. Harness ≠ findings | +| Certify a model "unbiased" from one BBQ run | **Fail closed** (do not) (`jev-bbq-experiment`) | 97.28% / bias 0.04/0.34 *theirs* is one frozen English/U.S. QA template. Not hiring/lending/healthcare | +| Let the LLM plan *and* fill in jeffrey | **Fail closed** to the split | Jev owns next-tool/progress/risk/done; LLM only fills args. Risk ≥ 0.5 pauses mutating tools | +| Fail a build on a missing jevlint verdict | **Fail open** (no verdict ≠ clean) (`mizchi/jevlint`) | Failed request never reads as a clean repo. Distinct from huntedman/JevLint | +| Skip generative review of a safety-escarpment hunk | **Fail closed** (always keep) (`prune-review`) | Concurrency/auth/a11y/startup stay in the packet regardless of Jev. Cost 1.18% with 305% outlier *theirs* | +| Treat empty intent-search as VERIFIED | **Fail closed** to UNKNOWN (`jev-intent-review`) | Empty search ≠ proof. VERIFIED is only as complete as the search | +| Treat GLiNER2 spec JSON numbers as measurements | **Fail closed** (spec-only) (`Jev_from_GLiNER2`) | Design for implementation; no service, no training. Interface ≠ replica. Distinct from jeff | +| Treat grande/laya-jolt/JEV-CPU/local-jev softmax as a Noul | **Fail closed** until calibrated on *your* labels | Packed-vs-separate / byte-parity / CPU logits / ONNX NLI are substrates. local-jev `confidence` omitted; done 30% vs Jev *theirs* | +| Forget user constraints after compaction | **Fail closed** to persisted structured state (`pi-heed`) | Jev never writes policy. Fail-open if Jev is down. Shadow default | +| Treat capability canaries as a quality headline | **Fail closed** until stages finish (`jev-judge-bench`) | SLA-150 frozen; invalid = FN; 21 offline tests; canaries 5/5–10/10 *theirs* are availability. **No quality result shipped.** Distinct from jevarena / jevbench | +| Treat cookbook sample scores as benches | **Fail closed** (do not) (`jev-cookbook`) | 16–36 handmade items; authors say not benchmarks. 425 calls / $0.015 *theirs* is cost/latency, not F1 | +| Click a GUI target below threshold | **Fail closed** to `unknown` (`pi-jev-control`) | Never force-click. Compaction never modifies the on-disk session | +| Treat a missing Vercel `confidence` as vendor Noul | **Fail open** to margin + a lower bar (`jev-use`) | First loop 17/20 escalate then 0/20 at 0.4 *theirs*. Distinct from jev-ultrafast | +| Sell jev-gpt as a product writer | **Fail closed** (architecture demo) | ~400 calls / 75 s / 2¢ *theirs*. The model never free-generates. Distinct from jeffrey pick≠fill | +| Endorse competing NAR from README badges | **Fail closed** until like-for-like + receipts (`openJev-verdict-2.0`) | Open PR #1: throughput≠latency; Laya gap inside CI (parity); correctness-head ECE ≠ distribution ECE. ≠ IamBusy/OpenJev | +| Treat an empty compaction-proxy slogan as a product | **Fail closed** (empty repo) (`IPECTER/jev-context-pruner` **and** `IPECTER/jev-runway`) | context-pruner 409 empty; jev-runway LICENSE-only (created≈pushed 1s). Sibling of fast-jev-compaction / jev-compactor / dizk/jev-lens — no files | +| Threshold chakuho softmax as a Noul | **Fail closed** until labelled calibration (`chakuho`) | Coverage is format-mass, not correctness. 8B stays coverage 1.00 while `__none__` collapses. Arithmetic stays in code. 503 is could-not-judge, never "no" | +| Treat jevinf speedup as a quality headline | **Fail closed** (argmax-parity only) (`jevinf`) | 2.57×/2.27× at 100% argmax *theirs*. MPS only. Wire-compat ≠ TypeSafe replica | +| Collapse typesafe-elixir-sdk into dannote/jev | **Fail closed** (different jobs) | HTTP client over Req ≠ OTP peer GenServer. Code still owns the `cond` | +| Round a commitjev middle band to pass/fail | **Fail closed** to `"review"` | ≥0.65 is the only verdict. Nouls decide; Choice is headline. Regex proves literals. Hook fails open on instrument failure | +| Treat hermes-plugin-jev as TypeSafe Jev | **Fail closed** (identity) | Live backend is Agnes 3.0 Flash chat-completions. Distinct from hermes-jev-router | +| Confuse pi-jev-compact with pi-jev-compaction | **Fail closed** (qualify owners) | compact = Pi summarizer replacement (verbatim keep/drop). compaction = fast-jev-compaction cousin | +| Trust an English Laya checkpoint on non-English | **Fail closed** to the multilingual checkpoint / script router (`laya-multilingual`) | Khmer 0.000 acc at 0.952 confidence *theirs*. Mean conf never < 0.885. Gating cannot catch it | +| Threshold schema-scorer peaked p as frequency | **Fail closed** (ranking ≠ calibration) | Grouped softmax trained on one-hot. Hub MIT; GitHub 404 | +| Ship jev-cli as if v1 existed | **Fail closed** (not ready) | README: commands not implemented. Crate: not ready. Distinct from jevql | +| Authorize a proposed tool call | **Fail closed** on a deterministic security failure (`actiongate-jev`) | Jev supplies evidence; code owns authority. Positive score never overrides RBAC/schema/limit. Financial/destructive/credential fail closed if Jev is down | +| Auto-act on a raw decision-model p | **Fail closed** until domain recalibration (`does-jev-confidence` / `jevcal`) | Ranking ≠ calibration. Vendor "calibrated" often means rank-correlation. Stated ~75% vs human ~10% *theirs* | +| Compact a long agent history (middleware) | **Fail open** to uncompacted history if Jev is down (`jev-compactor`); **fail closed** on pending destructive/exfil | Dual polarity in one product. Regex floor always local. Contrast gliner25-compaction fail-closed `keep_full`. Never rewrite kept bytes | +| Click / type from an indexed viewport | **Fail closed** to code-owned `--until` / stuck / no-guess fill (`ego-jev`) | Jev `done` is not business success. Malformed text-model JSON does not guess a value. Stale refs aborted | +| Collapse a social reply | **Fail open** as hide-not-delete (`x-reply-filter`); local rules first | Auto-hides are not training labels until a human confirms. Never self-reinforce on the model's own negatives | +| Rank the next skill from live context | **Fail open** on the Claude hook (`skillranker`); CLI keeps real exit codes | Advisory ranking; none-of-these is first-class. A failed recommendation must not block the agent. Distinct from skill-broker (grants) | +| Authorize a proposed tool after policy permit | **Fail closed** on explicit deny; **Review** if Jev is missing (`turnstile`) | Jev never grants what policy denied. Replay thresholds on saved scores; starting 0.85/0.35 are not calibrated. Observe mode is not enforcement | +| Endorse a capability claim / bake-off slogan | **Fail closed** until receipts (`jev-capability-atlas`); attach thinking budget (`jev-frontier-100`) | Not a leaderboard. Schema-valid ≠ correct. "Weaker than 4B" needs the thinking condition | +| Threshold raw p on an unseen rule | **Fail closed** until type-specific recalibration (`jev-ood-calibration`) | AUC ≠ ECE. Choice/Score overconfident (T~3.3); boolean underconfident (T 0.66) on the same tickets. Unknowable policy labels still get mean p 0.74 | +| Gate a vector on Jev `confidence` alone | **Fail closed** until you inspect entropy/margin (`how-sure-is-jev`) | Choice confidence = max_prob (most generous). 75/25 → 0.5 vs entropy 0.19. Bands are policy | +| Advance a sealed effect | **Fail closed** until Seal + coverage (`seal`) | Jev answers questions; SEAL answers whether the world may change. Open escalations keep Effects locked | +| Return a confident pharmacy/protocol verdict under chaos | **Fail closed** to escalate (`jev-labs`) | Never confidently wrong. 1,080 golden 0 wrong *theirs* is not a proof of zero. Stability ≠ answerability | +| Auto-approve a PR from four typed questions | **Fail closed** to human-review unless operator bar says so (`ci-gatekeeper-bot-jev`) | Conservative default escalated trivial diffs. Distinct from latch (finished red run) | +| Rank Codex capabilities / excerpts | **Fail open** to lexical overlap (`jev-in-codex`) | Ranking unbenchmarked. Rec ≥ 0.5 is a heuristic. Caller supplies the catalog | +| Reinspect a turn's risk from a Stop hook | **Fail open** (Claude finishes anyway) (`jev-preflight`) | Attention redirect, not a merge blocker. assist = one reinspect. 0.85 uncalibrated. Distinct from rashedInt32/jev-lens | +| Shrink a tool result before first send | **Fail closed** to full text for *code* unless confident (`dizk/jev-lens`) | Compress-before-first-send. Post-send prune cost 17% more *theirs*. Distinct from rashedInt32/jev-lens | +| Depend on the agent calling `recall` | **Fail closed** to a SessionStart hook (`carryforward`) | tools≠use. 0/4 *theirs*. MCP sitting there is not enough | +| Compact observational history | **Fail open** (failed Jev does not drain the buffer) (`pi-observational-memory-jev`) | Keep/kind only; verbatim ledger; model-free compact. Do not install beside Alvar `/om` | +| Actuate a lock / heater / smoke alarm from a Noul | **Fail closed** (do not) (`HA-Jev`) | Physical-world S1. Sensors and automations yes; safety actuators no | +| Skip an expensive chat completion via same-intent cache | **Fail open** (call upstream) (`jevcache`) | False HIT serves the wrong answer. 0 FP / recall 0.38 *theirs* n=100. Stream/tools/multimodal bypass | +| Strip the skill roster / recommend none | **Fail open** (Pi keeps listing; no-op without a key) (`pi-jev-skill-suggestion`) | Tool mode is tools≠use: the agent still has to call `skill_suggest`. Auto mode pays a Jev call every prompt | +| Return control to the LLM from a typed baton | **Fail open** (typed escalate, never a blocked agent) (`jev-handoff`) | Gate `allow` never grants. Vercel drops confidence — margin fallback is not calibrated | +| Skip the next Hermes main-model call | **Fail open** (continue to the LLM) (`hermes-jev-router`) | WHETHER/HOW/WHAT. Skip-next needs a core patch. Aggressive defaults. License null | +| Badge a feed item skip / save | **Fail open** (show `?`) ([ThinkyMiner/Winnow](https://github.com/ThinkyMiner/Winnow)) | Worth-your-attention VOI. Distinct from kevinpita/winnow. Templates never prose | +| Fail CI on an adversarial browser finding | **Fail closed** only high conf **and** high severity (`browser-jev`) | Sample from the distribution, not argmax. Code-only checks first. Visual blind. Baseline before trusting | +| Treat `/v1/decide` or SemIf `/v1/systemone` as hosted Jev | **Fail closed** until you name the scorer (OpenJev / semif-serve) | Wire-compat ≠ replica. OpenJev is not TypeSafe. Runoff is a product, not a softmax | +| Collapse conflict and ignorance into one Noul | **Fail closed** to a named Choice escape (`jev-typed-evaluation-collapse`) | Noul 0.50–0.57 vs 0.46–0.48 *theirs*. Binary Choice without escape is lexically biased | +| Treat a security-scan / prepared review / guardrail demo as policy | **Fail closed** (they are sensors) | rh-guard owns reward-hack. Reviews never stop commands. TeoMastro numbers unpublished this pass | +| Drop a meaning-grep line | **Fail open** as ranking (`jev-semgrep`); keyword still wins exact strings | AND/OR/NOT over *thresholded* line Nouls. Japanese noisier near threshold. Not a gate | +| Treat classifier.dev as a chatbot / agent hook | **Fail closed** (it is a classification API) (`classifier-dev`) | Public contract is label + calibrated confidence; batch `{id,text}[]` ~1000. Distinct from ask-jev-ai's six-question wall | +| Escalate every multi-label answer on `tier: smart` | **Fail closed** (do not) | Re-judging made it worse (23 s). Smart re-asks **single-label <0.7** only. 0.7 is *theirs*, not a class constant | +| Quote classifier.dev numbers without `eval/README` | **Fail closed** (do not) | `/benchmark` is tracked `vs-jev.json`, not transcription. n=7 train-on-test; ~0.03 is a coin flip | +| Serve a silent fallback as the advertised model | **Fail closed** (mark `FALLBACK`) | granite-4.0-h-micro F1 **0.546** vs advertised ~**0.800** *theirs* for weeks. rh-guard owns the eval-integrity gate; this is the lived cousin | +| Treat choxos/jev-reviewer as egma-ai PR attention | **Fail closed** (different products) | Systematic-review pointer ≠ code-diff attention. Always write **choxos/jev-reviewer** or “systematic-review Jev Reviewer” | +| Let Jev write the quote / skip *Not found* | **Fail closed** (copy verbatim; *Not found* is an answer) (`choxos/jev-reviewer`) | Pointer-not-generator. Noul ≥ 0.5 *theirs*. Unclear keeps closest lines. No paraphrase invent | +| Auto-accept an unchecked extraction quote | **Fail closed** (human tick is the product) | Checked answers are never overwritten by a reworded question. Jev is SENSOR; the reviewer is the constraint | +| Collapse githubnext/localjev into kunchenguid/local-jev | **Fail closed** (qualify owners) | Bun Chat Completions bridge ≠ ONNX ModernBERT approximation. Always write **githubnext/localjev** | +| Treat githubnext/localjev JSON probs as OpenJev logits | **Fail closed** (wire-compat ≠ logit-equiv) | razorback16 structured-read + logprobs vs prompted JSON → validate/retry → normalize + entropy confidence. SDK drop-in is the *wire* | +| Threshold LocalJev self-reported p as a calibrated Noul | **Fail closed** until labelled calibration on *your* workload (`githubnext/localjev`) | README: evaluate before consequential use. Bake-off: do not treat outputs as calibrated (wrong-BoolQ high conf → large NLL). JSON-valid ≠ picked-right | +| Swap LM Studio in and call it OpenJev parity | **Fail closed** (runner gap) | DiffusionGemma load still open (mlx-engine#336 / bug-tracker#2037 *theirs*, 18 Sep 2026). Chat Completions keeps the prompted-prob path. Structured-read primitives are the path | +| Collapse NandhaKishorM/laya into Hub-only / localjev / TypeSafe drop-in | **Fail closed** (qualify the face) | GitHub/PyPI packaging of Hub Laya; **≠** new species; **≠** githubnext/localjev; **≠** `/v1/systemone` SDK drop-in. Always write **NandhaKishorM/laya** | +| Treat typed-decisions 0.766 as zero-shot | **Fail closed** (fine-tune on that split) | Base ckpts 0.362 / 0.342 vs majority 0.461 *theirs*. Fast base to specialise | +| Hard-act at Laya conf 0.85 | **Fail closed** until Harbor cal on *your* labels | README recipe. Khmer 0.000@0.952 already proves gating cannot catch script OOD. 0.85 is *theirs* | +| Quote the vs-Jev table as independently measured here | **Fail closed** (third-party unpublished-here) | No TypeSafe API access; sample sizes/prompts differ; Banking77 72 vs 77 labels | +| Treat post-T ECE 0.081 as raw ECE | **Fail closed** (name the temperature) | Raw typed-decisions ECE 0.213 vs Jev 0.144. 0.081 is after domain T fit | +| Treat the @airesearch12 census tweet as a scored bake-off | **Fail closed** (it is a list + a promise) | ≠ [jevbench](https://github.com/fstandhartinger/jevbench) v1.1. Watch [jev-models](https://benchmarkheaven.com/jev-models); do not paste live ranks into the census card. **≠** jev-judge-bench / jevarena | +| Collapse GLiNER2 or routers into NAR / TypeSafe clones | **Fail closed** (class-boundary) | GLiNER2 locates/categorizes; Succinct 14M / jev-model-router / Director / Loki route. Same job family ≠ replica. Needle 3 is function-calling (already §67) | +| Treat an incomplete openjev list as our watch being wrong | **Fail closed** (census lag) | Laya, githubnext/localjev, kev, TypeAR, openvons, chakuho, jevinf, grande, laya-jolt, blackwood, classifier-dev missing. Lesson, not a dunk | +| Quote tweet likes/views as quality | **Fail closed** (ephemeral) | SIGNAL ~417/9/3; this pass 564/15/5. Do not copy Stripe | +| Mix v1.1 Main 87.6 with v1.2 Score 75.3 as a drop | **Fail closed** (not comparable) | Different tiers and scoring *theirs*. Cal now ON the composite. Keep v1.1 as historical (`notes.md` §67, §78) | +| Treat 75.3 as a class ceiling without the four axes | **Fail closed** (geo-mean product) | I/C/S/K 25% each. Weak axis dominates. SemIf −0.7. Weighting views reorder ranks | +| Treat Luna I=96.8 as rank #1 | **Fail closed** (Cost 28.2 → rank #7) | Accuracy cannot buy back a weak axis. DeepSeek C=96.7 is rank #11 | +| Ignore the ×2 latency assumption / est. costs | **Fail closed** (Harbor honesty) | Self-host/demo ×2 (+0.15 s) is an **assumption, not a measurement**. Many costs est. from OpenRouter/DeepInfra size-class. Production APIs unadjusted | +| Treat Qwen3.8 27B as Archer | **Fail closed** (official Qwen / Chutes TEE) | Partial, Cost 0 from price, hard 21.4%. Archer still Watch | +| Treat Laya absence as a quality verdict | **Fail closed** (gap, not a named exclusion) | Absent from the scored table **and** from the named exclusion list. Completeness ≠ dunk. GLiNER2 is mapping-excluded *theirs* | +| Collapse instruction models into NAR clones because they share the table | **Fail closed** (class-boundary) | Luna/Gemini/DeepSeek/Qwen3.8 are JSON-schema instruction models. Needle 3 is function-calling (C none → 0). OpenJev on board = razorback16 ≠ IamBusy `/v1/decide` | +| Treat the geometric mean as a natural law | **Fail closed** (weights are a choice) | Limits *theirs*. Balanced no-cal puts SemIf #1; Emphasis Cost puts system-one-open #1 / Jev #5 | +| Re-card localjev / classifier.dev / Laya / choxos / census / v1.2 because they reappear on the hourly | **Fail closed** (already folded) | Apply the five as a recipe (`notes.md` §79). Do not dump the hit list again | +| Hard-gate a Noul as a PR merge / quality seal | **Fail closed** (soundness theater) | Soft Noul attends or escalates; an exact envelope proves the irreversible act. [totally-tim/jev-gate](https://github.com/totally-tim/jev-gate) (0★; MIT; Action/CLI/OpenCode) and [connectedGraph/claude-jev-warden](https://github.com/connectedGraph/claude-jev-warden) (1★; MIT; “Art Director Warden”) are this hour’s skip with that risk. **≠** [vinilana/jev-gateway](https://github.com/vinilana/jev-gateway), [MongLong0214/jev-gate](https://github.com/MongLong0214/jev-gate) (model routing), [SargeDev/jev-gate-student-b](https://huggingface.co/SargeDev/jev-gate-student-b). Cousin: ci-gatekeeper (cheap typed pre-review, operator-owned). Attention filter never blocks ([rashedInt32/jev-lens](https://github.com/rashedInt32/jev-lens)) | +| Stall S1 waiting for S2 / let S2 fly | **Fail closed** (S1 keeps the stick) (`jev-reflex-autonomy-lab`) | Escalate-under-threshold **without blocking**. S2 is one-use advice. Distinct from classifier.dev smart re-ask (that path *does* wait). `notes.md` §80 | +| Treat purple S2 arrival as consumed guidance | **Fail closed** (log consumption) | Purple confidence = that Jev decision used returned S2. Purple S2 bar = arrival. Red = fail. Arrival without a later purple point is unused VOI | +| Collapse Local controller into githubnext/localjev | **Fail closed** (qualify the face) | Built-in **rule-based** reflex, no credentials. **≠** prompted-JSON Bun `/v1/systemone` **≠** ONNX ModernBERT. Always write **Local controller** vs **githubnext/localjev** vs **kunchenguid/local-jev** | +| Treat the 20% starting gate as Harbor τ / a flight interlock | **Fail closed** (still soft) | Selecting Live API sets 20% *theirs* and does **not** start a mission. schema-safe ≠ correct. Calibrate τ on *your* labels | +| Treat the seed as a deterministic async replay | **Fail closed** (geometry only) | Live latency still changes the trajectory. Seed repeats obstacle layout, not timing | +| Send pixels / planner prose into the reflex | **Fail closed** (text-state; no graphical input) | ARCHITECTURE *theirs*: no pixels to either provider; planner narrative excluded from Jev input; code never labels safest. Not omni. Skip Archer | +| Assume Jev confidence = selected probability | **Fail closed** (not assumed) | Application contracts ≠ TypeSafe SDK methods. Physics owns collisions. S2 never grants | Worked placements (2026-09-18 topic:jev hour + prior archive): @@ -162,6 +343,29 @@ Worked placements (2026-09-18 topic:jev hour + prior archive): relevance against the task; last-N lines and error signatures kept in *code* before Jev sees anything; full output recoverable by id. Same shape as winnow/fast-jev-compaction (`agent-self-assessment.md`). + Encoder-backend cousin: + [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) + — GLiNER2.5 retention Choice + exact spans; fail closed to + `keep_full`; `shadowMode` default true (`notes.md` §50). Not Jev. + Stdout-prune cousin, same family, different job: + [jev-pruner](https://github.com/tamaratran/jev-pruner) — Jev Noul on + residual noisy Bash after a hard ≤10k/format envelope; fail-safe + original; archive for recovery (`notes.md` §53). Marketplace id + still `fast-jev-output`. +- **Session ledger before the next task** — + [carryforward](https://github.com/Dharundp6/jev-carryforward): + verbatim facts; Jev scores which are still live; rules never + judged; fail-open dump (`notes.md` §55). +- **Evidence set before the generator** — + [decision-native-rag-skills](https://github.com/emergency-lee/decision-native-rag-skills): + retrieve wide → decide → evidence set → LLM. Embeddings stay + candidate generators. No universal benchmark (`notes.md` §55). +- **Files / URLs before the main agent reads** — + [jev-sift](https://github.com/kbhuw/jev-sift): classify first, + read selectively. Batch path/url/text → Jev; the main LLM opens + survivors. Uncertain/errors/truncation ≠ irrelevant. Topology A + MCP (LLM outer loop). Transport tests ≠ accuracy. No LICENSE this + pass (`notes.md` §56). Do not copy plugin how-to. - **Diff hunks before `git add`** — `ibrahemid/git-jev-stage`: one Choice per hunk (`include` / `exclude` / `mixed`); mixed and low-confidence stay unstaged; lines never split; staging is an exact patch after confirm. @@ -171,7 +375,9 @@ Worked placements (2026-09-18 topic:jev hour + prior archive): runs for root cause. Prefilter of *analyst attention*. The named cut this hour: Noul `env_broken` *as opposed to* the agent's own bug; silent vs disclosed vs recovered vs clean. Pre-mortem of a new - sandbox *before* users meet it (`notes.md` §42). + sandbox *before* users meet it (`notes.md` §42). Merge-gate of the + same split: [latch](https://github.com/CaseReed/latch) clusters in + code, labels with Jev, policy owns PASS/BLOCK (`notes.md` §51). - **Realtime hold-before-publish** — community moderation claims ~200ms (**Hypothesis** as a number; **Empirical** as a family via Near Here / jev-experiments firehose in the archive). Thresholds stay yours. @@ -203,7 +409,8 @@ Live ecosystem (examples of the *shape*, not SDKs to copy): - Official skill_suggestion cookbook; GodsBoy router 94.4% vs 70.8% lexical. - `Dicklesworthstone/skillranker` — next-step skill rank from live session - context; fail-closed; calibration loop. + context; two-pass + none-of-these; Claude hook **fail-open** (quiet + exit 0); CLI keeps real exit codes; local calibration/replay. 52★. - `WiktorB2004/llama-index-jev` selectors — which query engine handles the query; fail closed or a declared default. - Toolrouter (product, X 2026-09-18) / open JevRouter harnesses — request → @@ -211,8 +418,30 @@ Live ecosystem (examples of the *shape*, not SDKs to copy): - `rajdhakad9826/routeKit` — Jev estimates requirements; a deterministic policy picks the model. Jev does not choose the LLM. **Hypothesis** until your catalog (`notes.md` §33). +- [jev-adaptive-thinking](https://github.com/jxu-dev-c/jev-adaptive-thinking) + (~18:46) — session-sticky first-prompt Choice; lock for the + session; fail-closed to `gpt-5.6-sol`. License null. Live testing + left to the deployer (`notes.md` §58). Do not copy dylib/YAML. +- [`trietphan/jev-claw`](https://github.com/trietphan/jev-claw) — same + hole for OpenClaw: Jev classifies task/complexity/risk; `decide()` in + code maps to a route; a sensitive-path regex floors risk. Confidence + is the **min** across heads. 11 offline policy tests; live 10/10 is + the author's 10 samples (`notes.md` §44). +- [`nekowasabi/jev-routing`](https://github.com/nekowasabi/jev-routing) + — host **adapter** (Go binary), not an MCP server, in front of + Claude Code / Codex / Grok Build. Compacts tool results, then one + Choice + done-Noul, then one schema. Adding it via `mcp add` makes + the catalog worse. No key → on-device classifier. Do not copy ports. +- Higgsfield API auto-routing ([tweet](https://x.com/higgsfield_ai/status/2101022473248727177)) + is the same hole on a video/image catalog. **Claim**, no labeled + receipt. - Function-calling cookbook (**Contract**): function *names* and closed-set args as questions; code still validates the call. +- [jev-gateway](https://github.com/vinilana/jev-gateway) (~16:48) — host + adapter: Jev picks the tool (and closed-set args); LLM fills open + args or is skipped (`direct`). Fail-open passthrough if Jev is down. + Harbor on/off measurement: [jev-gateway-bench](https://github.com/vinilana/jev-gateway-bench) + (one-run signal, not a measurement; `notes.md` §51). Do not copy ports. Routing ROI does not transfer across datasets (`validation.md`, calibre). `FirasSX914/Janus` exists to *measure* when Jev vs another model wins on @@ -238,13 +467,68 @@ SREGym-Lite is topology A: Jev ranks next tests/evidence; the agent still runs them and still diagnoses (`notes.md` §33). [`runta-dev/jot`](https://github.com/runta-dev/jot) is topology B with a *closed* tool catalog (the host executes; the calculator does the -math). "First general-purpose System One agent" is a claim. **Does not:** the decision model as the planner — neither inventing tools +math). "First general-purpose System One agent" is a claim. +[`nekowasabi/jev-routing`](https://github.com/nekowasabi/jev-routing) is +a host adapter in front of an existing coding CLI (not topology B, not +MCP): it peels the catalog *before* the generator sees it. +[`kbhuw/jev-sift`](https://github.com/kbhuw/jev-sift) **is** topology A +MCP: the LLM still owns the outer loop; Jev is a tool that classifies +paths/URLs/text before the agent reads them (`notes.md` §56). Do not +merge with jev-routing. **Does not:** the decision model as the planner — neither inventing tools nor picking its own next tool in a loop (standing red flag, above and in `boundary-audit.md`); skipping schemas so the model "just knows"; treating a workflow AST as a proof. The outer loop stays with the LLM or with code. **Hypothesis** as "Jev builds the AST"; **Empirical** as named topologies. Not an MCP how-to. +**S1 keeps control; S2 is one-use advice.** Topology B under latency: +the reflex (typed Choice over legal actions) never hands the stick to +the planner. Optional System 2 is *advice* on low confidence, one-use, +asynchronous — the reflex does not pause +([jev-reflex-autonomy-lab](https://github.com/khordoo/jev-reflex-autonomy-lab); +experimental drone viz, not a flight controller; GitHub license null +this pass). Same Kahneman split as the toolbox row (S2 proposes, S1 +discriminates; never the reverse). `notes.md` §46. **Delta +(`notes.md` §80):** escalate-under-threshold **without stalling**; +telemetry marks when guidance was **consumed** (purple confidence = +that Jev decision used returned S2; purple S2 bar = arrival; red = +fail) — arrival ≠ used. **Local controller** is a built-in +rule-based reflex, **≠** githubnext/localjev **≠** +kunchenguid/local-jev. Live API is hosted `POST /v1/systemone` +`jev-latest`; selecting it sets a 20% starting gate *theirs* and +does **not** start a mission. Seed = geometry, not async replay. +No pixels to either provider. Confidence is not assumed equal to +selected probability. S2 never grants. 20% is still soft. +`agent-self-assessment.md`. **Route ≠ memory** is the same split on a +turn: [jev-hermes](https://github.com/de-niji/jev-hermes) cheap-gates +calendar/mail/status off the memory tour; complex keeps Honcho +(`notes.md` §48). **Advisory sidecar:** +[agent-workflow-typesafe-ai](https://github.com/ngallodev-software/agent-workflow-typesafe-ai) +emits `no_action` receipts and **never** changes host routing. +**Productized Kahneman cascade:** +[dual-process-ai](https://github.com/taro1985/dual-process-ai) — S1 +decides, S2 writes; routing accuracy **not measured**; keyword +fallback is not S1 (`notes.md` §49). +**Harbor-shaped decide→policy leftover (measured compare arms, +2026-09-18 ~20:43):** +[jav-email-cascade](https://github.com/skiingfalcon/jav-email-cascade) +— shared Answer schema; jev / gen-json / gen-logprob; policy +auto/review/llm; Noul 0.5 never rounded. Mock gen-json +flat-confidence is *their mock*. License null this pass +(`notes.md` §60). +**Closed-vote computer-use, no planner LLM (2026-09-18 +~21:39):** +[JevOnly](https://github.com/buluoray/JevOnly) — code builds +options, Jev only picks; fact register + verify/undo; type +without generation. Apache-2.0. Distinct from Stagehand LLM +fallback (`notes.md` §61). +**Host-owned product surface:** +[waymode](https://github.com/mossburgh/waymode) — app retains +handlers/permissions/validation/state; Jev over live typed +actions; `completed` is Jev's reading. 24/26 and 34/36 +*theirs* — bounded evidence, not a self-driving proof +(`notes.md` §61). + **Effect-oriented loop (same author, later post).** Topology B inside an effect system ([tweet](https://x.com/JamesWard/status/2100981305009664299)): the host @@ -278,15 +562,48 @@ ask Jev to own the standard of quality. Good checks: "does this diff introduce new mutable module-level state?", "API impact: none / additive / behavioral / breaking?" +[`coldteadotai/abide`](https://github.com/coldteadotai/abide) is the +fuller **productized** path of that contract (compile / calibrate / +tune / replay / audit; Claude Code / Codex / OpenCode). Soft +AGENTS.md / CLAUDE.md rules → one Score per rule on the **diff**, +never the conversation; hard, linter-checkable rules stay with the +linter — same layering family as jevgate's hard envelope (`mappings.md` +§18). Edit-phase vs turn-phase is an observation-window question +("added more than asked" has no answer after edit 1 of 12; +`question-design.md`). Product bands: ≥0.8 repair, 0.5–0.8 note, <0.5 +silence — **their** operating point, not a universal 0.8 (`mappings.md` +§2). Fail-open: no key / no network → the edit proceeds; hooks exit 0. +`.abide/rubric.json` quotes source lines; false positives are rule +rewrites, not a model swap. Replay of 93 sessions with an independent +reviewer: edit precision ~26%, turn ~73% (`notes.md` §47; +`validation.md`). Text/diff only — not multimodal. Complementary to +[`24601/rh-guard`](https://github.com/24601/rh-guard) (eval-integrity / +reward-hacking on tool use vs project soft rules on diffs): same +hook-host lessons, different judgment class; do not merge products. +Do not copy hooks. + Related placements: -- **AGENTS.md / project prefs as criteria** — jev-pref; pi-warden rule - breaks 6→0 on 150 paired runs (`agent-self-assessment.md`). +- **AGENTS.md / project prefs as criteria** — jev-pref states the + contract; Abide productizes it; pi-warden rule breaks 6→0 on 150 + paired runs (`agent-self-assessment.md`; `notes.md` §47). + [if-ai](https://github.com/Victor-Casado/if-ai): one plain-English + condition + required min-confidence; fail-closed on error + (`notes.md` §51). [jev-marshal](https://github.com/LightningK0ala/jev-marshal) + is Watch / empty. - **Confidence gates + shadow mode** — `AntonioCoppe/jev-harness` (48.9s Claude CLI vs 1.3s Jev on a 24-row filter). Log would-do until evals pass. Selective abstention (`mappings.md` §2): low confidence is `review`, not a guess. Coppe on SREGym regressions: inspect whether - confidence was high on the wrong Choice (`notes.md` §33). + confidence was high on the wrong Choice (`notes.md` §33). This hour + the *practice* is first-class: eval CLI asserts on the **action**, + not on prose; recipes span alerts / RTB / sports-bet / prediction + markets (`notes.md` §44). LLM-as-judge is not the primary System One + score (`faq.md`). Compaction rollout of the same instinct: + [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) + public default `shadowMode: true` (analyze + log; do not replace + history until explicitly enabled) (`notes.md` §50). Do not copy the + client. - **Hybrid countable + judgment rules** — `DanRWilloughby/snifftest`: deterministic tells score 1.00; judgment rules flag only outside the unsure band. Explicit: a reading near 0.5 is *no judgment*, never a pass. @@ -298,6 +615,13 @@ Related placements: rules → file-level Noul ≥ 0.8; write→check→fix; no line-level, no generated names, no auto-fix. Sibling of jev-pref. Independent, not TypeSafe. Pointer: `notes.md` §26. +- **Sentence-as-rule lint (ast-grep × Jev)** — + [`mizchi/jevlint`](https://github.com/mizchi/jevlint): matcher + decides *which* code is looked at (silent miss); a sentence + `ask:` decides whether it is a problem (loud). 13/15 1.00/1.00 + on their corpus; review mode 2 req / $0.00013. Fail-open no + verdict. **Always qualify** vs huntedman/JevLint. Independent of + eslint-plugin-jev (`notes.md` §70). - **Malicious-before-run** — `luantak/is-malicious`. High-stakes gate: fail closed, shadow first, never treat a Jev yes as authorization to execute untrusted code. Code still sandboxes. @@ -318,47 +642,247 @@ decision-design card. Do not clone APIs from READMEs. | Placement | Judgment | Stays in code | Artifact | |---|---|---|---| -| Hold-before-publish moderation | Hazard Nouls + harm Score | Block/review/pass policy | Near Here / firehose family | +| Hold-before-publish moderation | Hazard Nouls + harm Score | Block/review/pass policy | Near Here / firehose family; **x-reply-filter** (local rules first; never auto-train on own hides) | | Tool / engine / skill select | Choice + fits-Noul | Dispatch, auth, reject-all | skillranker, LlamaIndex selectors, Toolrouter | -| Preference lint | Per-rule Noul/Choice on a diff | Rule text, outcome map | jev-pref, JevLint | -| Context / log prune | Per-line or per-block relevance | Always-keep set, recall keys | jevprune, winnow | +| Preference lint | Per-rule Score/Noul on a diff; or ast-grep subject × sentence | Rule text, linter for hard rules, bands + fail-open; matcher silent / Jev loud | jev-pref (contract), Abide (productized), huntedman/JevLint; **mizchi/jevlint** (13/15 1.00/1.00 *theirs*); if-ai (plain-English PR check, fail-closed on error); jev-marshal (Watch / empty repo) | +| Context / log prune | Per-line or per-block relevance; or a retention Choice + spans; or a Noul per stdout chunk; or keep/kind on a verbatim ledger; or pre-send views of a tool result | Always-keep set, recall keys; mutation envelope in code; shadow before replace; size/format envelope then Noul; archive dropped spans; kind-keyed topic files | jevprune, winnow; fast-jev-compaction / pi-jev-compaction / fast-jev-compaction-pi; gliner25-compaction; jev-pruner; **jev-compactor** (73% / 350 ms product-arm *theirs*); **dizk/jev-lens** (pre-send; 79% fewer tokens); **pi-observational-memory-jev** (keep/kind verbatim) | | Exact hunk staging | Per-hunk include/exclude/mixed | `git diff`, atomic apply | git-jev-stage | -| Semantic `WHERE` | Noul/`jev_prob` over a row | SQL, indexes, LIMIT | `kylemclaren/jevql` (CLI rewrites; DB sees ordinary SQL) | -| Home automation read | Choice/Score/Noul as an entity | Automations, device I/O | `AboveColin/HA-Jev` | +| Semantic `WHERE` | Noul/`jev_prob` over a row | SQL, indexes, LIMIT | **jevql** (CLI judges; vanilla Postgres never sees `jev()` — judgment outside the store); sqlite-jev / pg-jev (in-engine) | +| Formula / query embedding | JUDGE as a function | Spreadsheet/SQL engine | judge-sheets, jevql, sqlite-jev | +| Soft ABR / live encoder | Choice over a ladder | Probe × headroom, thermal, battery | bitrate-advisor | +| Voice → typed act | Choice/Noul on a transcript | ASR producer; macOS actions | jev-voice-control (README stub; §44) | +| ASR voice-browser CU (productized) | 9–11 questions on a partial transcript; pointer spans | Playwright; debounce; numbered overlay; spoken confirm still soft | moritzkremb/jev-voice-browser (MIT **103★**; `notes.md` §82). **≠** jev-voice-control **≠** nikolas-j **≠** typesafe-computer-use | +| Wrap-as-execution ALLOW/ASK/DENY | Intent-conditioned verdict before the real tool | Rules first; ASK throws; fail-closed; judge swappable | reddpy/AgentGhost (MIT **2★**; `notes.md` §83). **≠** actiongate **≠** toolgate **≠** jev-use. rh-guard owns the gate cousin | +| Host-adapter routing | Choice next-tool + done-Noul | Shrink `tools[]`; compaction | jev-routing (not MCP; **delta:** Cursor Agent CLI / Devin CLI this hour); **jev-in-codex** (Codex MCP; ranking unbenchmarked; lexical fallback) | +| Multi-model route | Classify axes; policy maps | Escalation `if`, path regex | jev-claw, routeKit | +| Finish-line gate | Noul/Score/Choice on evidence | Deterministic shell checks first | hermes-jev-north-star | +| Home automation read | Choice/Score/Noul as a sensor | Automations, device I/O, daily budget; **not** locks/heaters/smoke | [HA-Jev](https://github.com/AboveColin/HA-Jev) (MIT; **17★**; confidence gating; Jev-gates-LLM examples) | | Browser loop without generation | Action Choice over visible elements | Perception, constraints, click | lizard-agent | +| Screenshot / DOM candidates → Choice | Omni decide over letters code marked | Click/act in code; fail-open to specialist OCR | blackwood-rlcd (CC BY-NC; not Archer) | | Android / macOS computer-use | Choice over prevalidated candidates | UI tree / AX / OmniParser; no generated coordinates | jev-mobile, jev-macos-loop | -| Model router | Requirement Scores; policy in code | Eligibility, cost/quality/latency objective | routeKit | +| S1 reflex + optional S2 advice | Typed action Choice; planner one-use on low p; escalate **without stalling** | Collision, legality, physics; the stick stays with S1; S2 never grants | jev-reflex-autonomy-lab (experimental viz; **7★**; license null; `notes.md` §46 + §80) | +| Mixed-initiative consumption telemetry | Same Choice; mark when advice was *used* | Green = local context; purple confidence = consumed S2; purple S2 bar = arrival; red = fail | jev-reflex-autonomy-lab charts. Arrival ≠ used | +| Local controller vs Live API (reflex A/B) | Same questions; backend is rule-based vs hosted `jev-latest` | Physics/seed fixed; 20% gate still soft; credentials do not auto-switch | Harbor-adjacent of backends, **not** a scored bake-off. **≠** githubnext/localjev | +| Decision-as-business-tool | Named judgment; gate is part of the result | Registry, arithmetic, hard guards | jev-decision-layer (unofficial) | +| NL cases → checked e2e | Jev selects observed controls | Playwright expectations; PASS/FAIL/BLOCKED | jev-e2e (alpha) | +| Extractive quotes / pointer evidence | Per-sentence, per-line-id, or char-offset Noul/Choice | Verbatim join; place; `redecide` / CSV; model never writes the excerpt | testimonial-miner; jev-reviewer; gliner25-compaction | +| Structured observe → decide → act | Score / Choice among numbered a11y/DOM/OCR+AX controls | Guard check; deny-list absence; no screenshots **on the decision**; no generated selectors; TYPE is the only generation; `DONE` ≠ verified success | solari-reflex (Jev); jev-ultrafast (Jev); gliner2-ultrafast (GLiNER2); laya-mind2web (Laya, DOM indices); cua-s1 (option-attention fill/check/click/skip; not TypeSafe Jev; source-only); Stagehand experimental Jev (harness; draft #2951–#2955); **ego-jev** (ego-lite indexed table; operation+target; `--until` in code); **typesafe-computer-use** (macOS OCR+AX; hosted Jev; MIT **427★**; `notes.md` §81) | +| OCR+AX desktop CU (productized) | kind / item / site / offscreen Choices | Crop+tile OCR; AX bonus never sole; dates.py; writer only for free text; post-type Noul still soft | awlevin/typesafe-computer-use. 155× *theirs* one screenshot. **≠** jev-ultrafast **≠** cua-s1 **≠** camoufox | +| ASR voice-browser CU (productized) | intent / target / site / complete / is_command / destructive + span Choices | Web Speech producer; Playwright acts; numbered overlay; spoken confirm ≠ auth | moritzkremb/jev-voice-browser. 27/27 fixtures *theirs*. **≠** jev-voice-control **≠** typesafe-computer-use | +| Harness pick-and-copy extract | Choice among a11y candidates; completion Noul | Schema plan + validation gate in code; screenshot always LLM; LLM fallback; pick ≠ replacement | Stagehand #2955 (`off`/`judge`/`pick`; 37/75 no-LLM ~0.5s vs 4.37s *their* card) | +| Specialist form S1 (plan ≠ execute) | Option-attention among observed elements | Dry-run default; snapshot-bound tokens; reobserve; submit opt-in; fail-closed checkbox/fill | cua-s1 (`cua-s1-form-v0` profile; no weights this pass) | +| Hybrid local decide + remote fill | Local encoder scores observed controls | Code owns actuators; remote OpenAI-compat helper writes field text only | gliner2-ultrafast (GLiNER2 local + Mercury 2.5 default) | +| Dataframe semantic columns | Noul / Choice / Score per row; full `p__` | pandas/Polars, indexes, never silent renormalize | jevpandas; jevframe (PyPI + Polars) | +| Route ≠ memory | Intent Choice before a turn | Config + flat tools on easy routes; memory stays on for hard ones | jev-hermes | +| Advisory sidecar receipts | Typed answers as `no_action` evidence | Host routing / executor / policy unchanged | agent-workflow-typesafe-ai | +| Structure induction over a bag | Pairwise dependency Noul/Choice | DAG / scheduler in code | dag-jev (experiment) | +| Simulated world control vs content | Intent / page-type Choice | Generator writes documents; Zod + deterministic compiler; SQLite world | jev-agentworld-web-simulator | +| AST ∩ semantic lint | Typed questions on Tree-sitter units | Parser, selection, fail-on; does not execute scanned code | jevscan (`tenbin` owns the lint skill); jev-oxlint (skills→oxlint remainder; Phoenix experiment; not a hard gate) | +| Model router | Requirement Scores; policy in code | Eligibility, cost/quality/latency objective | routeKit; jev-adaptive-thinking (session-sticky first-prompt; fail-closed fallback) | | Bulk-judgment coprocessor | Choice/Noul off the frontier context | Counts, policy, fail-open gate | jev-mode | | Closed-catalog System One shell | Choice over host tools | Execute, arithmetic, credentials | jot | | Jump-by-description | Noul relevance on a local shortlist | zoxide index, local paths only | joxide | | Game move | Choice over legal actions | Rules, legality, win check | jev-plays-games | -| Finish-line gate | Noul/Score/Choice on evidence | Deterministic shell checks first | hermes-jev-north-star | | Analyst attention cascade | Step-level silent-failure Nouls | Grouping, LLM autopsy | OpenSmoke | -| Formula / query embedding | JUDGE as a function | Spreadsheet/SQL engine | judge-sheets, jevql | +| CI merge-gate (flaky vs real) | Cause Choice per clustered signature | Cluster + fingerprint + PASS/BLOCK table; reporter never fails the runner | latch (`notes.md` §51) | +| Fail-open wake / resume | p(wake) on waitingFor × event | Sleep duration, skip-limit, user-message always wakes | wakegate | +| Claim/evidence Stop | supports / contradicts / not-addressed per claim | Keyword retrieve session lines; firm-confidence floor never blocks | clear-head | +| Harbor on/off routing | Tool Choice per turn | Hidden verifier; fail-open if Jev down | jev-gateway + jev-gateway-bench (one-run signal) | +| Device-loop Choice | Folder among a closed catalog | Never invent folders; extension-map fallback | jev-downloads-sorter | +| S1 extract + escalate-S2 index | GLiNER spans / relations on the bulk | Graph in code; LLM only if backend loaded and low conf; query does not invent edges | s1-graphify-indexer (10–50× unfilled) | +| S1 specialists + S2 coordinator | Typed {value, probability} | Coordinator / hysteresis in code | reification-labs/foreman (**description-only** Phoenix scaffold; not the super-jev loop) | +| Bounded Pi supervisor | Skills / recovery / review / verify | Shadow default; never generates commands | jevons | +| Judgment as language primitive | `chance` / `pick` / `rate` (Noul / Choice / Score) | English-as-config; stub backend; fail polarity per action (`rescue nil` at save ≠ spam gate) | hunch (Ruby library, not a new language; cousin of probably-lang) | +| Decision-native RAG | Relevance / evidence / freshness / authority Nouls + Score | Evidence-set builder, conflict/temporal logic, provenance; embeddings generate candidates only | decision-native-rag-skills (no bundled harness; Hypothesis as a measured win) | +| Classify-first agent I/O | Relevance / typed questions on path/url/text | Hard envelope (50 / 60k / 2MB / public-IP); main LLM opens survivors; uncertain/errors/truncation ≠ irrelevant | jev-sift (MCP topology A; mocks ≠ accuracy; no LICENSE this pass) | +| Generative UI decide | Intent / layout Choice | Zod + deterministic compiler; model cannot add components | json-render + jev-agentworld-web-simulator | +| Robotics text-state | Choice on geometry-as-text | Code owns kinematics / Hz; two-call split; not pixels | MuJoCo showcase; jev-drone; Doom JSON. Drawing-pixel claim is not Archer | +| Draft-gate heartbeat | Quality Noul / Score | Fail-open / heartbeat on silence; missing verdict ≠ hold forever | jevable draft-gate fail mode; contrast Abide `<0.5` (edit proceeds) | +| Living class-pattern atlas | (not a model) | Cross-link exemplars; do not dump 342 titles | jevable.com (342 is *their* count; JSON-LD first page 36) | +| Verbatim session recall | Noul "still live for this task?" | JSONL ledger; constraints/corrections always-keep; fail-open dump; **SessionStart hook** (tools≠use) | carryforward (9×3 hint; **0/4** recall *theirs*) | +| Pre-exec tool gate | allow / block / review | Permissions, arg validation, transaction limits in code; timeout stops | toolgate (72-case synthetic, not independently annotated; Jev ≠ authorization) | +| Capability kernel | parallel hazard Nouls (sensor) | Closed action space; canaries/placeholders; BLOCK/ASK/ALLOW in `policy.py`; secrets never in the agent | interlock (ring 0 vs LLM ring 3; type-safe ≠ correct; 38-case local-judge set, not a blind paper). Distinct from toolgate | +| Typed control plane around an LM program | classifier over a closed ontology | Ontology validation, security override, confidence, state machine, tool allow-list in code; DSPy drafts AFTER route+action | jev-dspy-control-plane (OpenJEV / DSPy / JSON Schema share ontology; offline heuristic ≠ quality) | +| Engine owns truth / Jev owns judgment | severity / error-class / interrupt Noul+Score+Choice | Engine eval/lines/swing; templates + capped writing model; silence is a feature | game-coach (Wave 0 PRD; Stockfish WASM; GPL-3.0; anti-soundness-theater with egma attention≠correctness) | +| Human-confirmed OS kill | Stop / Keep / Your-decision Choice | Identity re-check; shields override; mapped explanations; TCP only | port-cleanup (conf ≥ 0.8 for kill recs; tiny pidfd-less race) | +| Native-probability calibration arena | noul/choice/score on analytic worlds | Oracle stub; Brier/ECE/reliability/risk-coverage; fan-out batches | jev-arena (live 145 noul Brier 0.0059 / ECE 0.0620 *theirs*; overconfident in low bins). Fan-out suite: jev-sonar (heatmap-as-policy); jev-vickrey (Jev never bids); jev-bracket (Brier vs Elo; live trailed Elo) | +| Meaning-as-spec UI test | pick among observed controls | Playwright acts; lockfile replay; refuse below 0.6 | jevcumber (.feature only; no step glue) | +| Conceptual PR labels | type/area/size Choice | Fixed taxonomy; never invent names; security/breaking never auto-removed | jev-pr-labeler (size = conceptual scope, not line counts; 0.75 abstain) | +| Healthcare S1 + S2 | NEWS2 remainder / med recon / inbox route | Code owns NEWS2, recon, routing; S2 blinded review | explore-typesafe-ai (synthetic FHIR; not clinically validated) | +| Public judgment wall | Six parallel questions (yes/no/depends, safety, mood, topic) | Policy-in-code; allowlist; p≥0.6 block; cost-to-1M from tokens | ask-jev-ai (license null; no-key UI says Jev offline) | +| Meaning-search without embeddings | Packed parallel relevance; two-stage outline→zoom | Keyword still wins exact strings; finds, does not explain | jevgrep (79% top-5 vs BM25 40% / grep 20% on stripped repos) | +| PR attention ≠ correctness | P0/P1/P2 attention | OpenAI writes deltas; alwaysReviewPaths P0; incomplete never P2 | egma-ai/jev-reviewer (**not** choxos pointer-not-generator) | +| Skills → oxlint | Remainder Noul/Choice after AST/precheck | Guidance whole-file in state; survey/calibrate/propose; not a hard gate | jev-oxlint (Phoenix fixtures; experiment; tenbin owns lint skill) | +| Session-sticky model route | First-prompt Choice | Lock for session; fail-closed declared fallback | jev-adaptive-thinking (license null; same family as routeKit) | +| Measured RAG rerank | Relevance vs a generative reranker | Evidence set / retrieval order on error; name the no-RAG arm | Jev-RAG (one-run ≥70%/72% vs Spark rerank; full-context Spark still faster) | +| Decide→policy→LLM leftover | 8 typed questions; shared Answer schema | auto / review / llm in code; Noul 0.5 never rounded; Score conf 0.0 never acted on; injection always review | jav-email-cascade (license null; mock gen-json flat-confidence is *their mock*; ~$0.034/1k *theirs*) | +| Domain specialist LoRA | soft-target Choice on independent gold | Threshold/deferral/EU in code; hosted few-shot when only argmax | Domain-jev-maker (CLINC labels, not Jev teacher-copy; KL 0.168 vs 0.580 banking *theirs*) | +| Wire-compat encoder backend | choice / score / noul on GLiFormer-400M | typesafe-sdk `base_url`; T=3.2; isolate nouls; tokens ≠ Jev billing | jeff (license null; ~$2.6 vs $15.6 L4 HTTP ~6×; A10G direct ~$0.65 ~24×; AG News 75.5% vs 90.5% *theirs*; not a Jev replica) | +| Loopback System One gateway | pass-through of whoever answers | Policy auto / prefer-local / prefer-hosted / local-only / hosted-only; credential from env never config; no weights | sysone (MIT; early; not a model) | +| ORDER BY ranking measurement | pairwise inversion / Score ordinality / ties | Gate SQL on results.json; secondary key on two-decimal ties; measure request shape | jev-orderby-bench (six gates pass; Score 0.143 weak link; 53-way 0.99 tie; recodelabs batch-40 fails ranking) | +| Closed-vote CU (no planner) | Choice among code-built options; done / off-path / next | Fact register; verify/undo; type without generation; irreversible risk never default | JevOnly (Apache-2.0; 11 steps / 43 calls / ~$0.014 / 17 s *theirs*) | +| Host-owned System One product | Choice among live typed actions | Host handlers, permissions, validation, state; prove writes from server state; p ≥ 0.7 default | waymode (MIT; 24/26 + 34/36 *theirs*; not on npm; not a self-driving proof) | +| OMP/pi acceptance + route | Choice `{accepted, rejected}`; topology/tier Choice | Fail-open missing Jev (`confidence: 0`); out-of-set answers → default | omp-jev-extensions (MIT; contrast pi-jev-approver fail-closed) | +| OMP prompt suppression | Parallel verdict + severity + in-scope | Operator owns presets; plugin never self-tunes; host deny fires first; agent prose withheld | omp-greenlight (MIT; 1,013/10; default 40.9% / 0 of 94 *theirs*; not a sandbox) | +| Pre-agent skill intervention | Relevance/confidence over authorised candidates | Code owns catalog/policy/grants; Jev never grants access; foundation-only on failure | skill-broker (**Hypothesis / outline**; not a production recipe) | +| Eval-instrument audit | Accuracy / ECE / blank lean / rewording of a Jev question | Check data/scorer/claims; 99 of 189 findings against itself | dinostomp (README Apache-2.0 / GitHub NOASSERTION; `dinostomp jev`; ECE 0.062 *theirs* on 24 examples) | +| Constrained optimizer + S1 features | task / exactness / external-evidence | Controller owns SLO/quality floors; fail-open local features; exactness never overrides capability | slo-router (license null; p95 **77.93 → 490.38 ms** same routes *theirs*; 3/8 label disagreements did not change routes; eight-row demo is not a benchmark) | +| Effect-based shell gate | Choice allow/deny + nine independent risk Nouls | Fast-allow/deny <1 ms; landed-script trust; headless ≠ auto-approve; fail-closed | construct-auto-classifier (Apache-2.0; Jev **0** dangerous / 975; $0.047/1k; privilege ≠ verdict) | +| Attention filter / human-review VOI | Per-file need-a-look / kind / debris Nouls | Never blocks the agent; never edits; never green unless sure | [rashedInt32/jev-lens](https://github.com/rashedInt32/jev-lens) + nvim. Distinct from [dizk/jev-lens](https://github.com/dizk/jev-lens) (pre-send views) and from jev-gates | +| Evidence-packet explorer | Rank BM25 shortlist; packet source_of_truth / tests / callers | Index once; agent still reads cited files; read-only | [jevex](https://github.com/jimmyhealer/jevex) (rename of jev-semantic-explorer; n=16 160s→69s / $8.74→$3.13 / 16/16 *theirs*; keep n=8 1/8→6/8; HitFile diagnostic) | +| Meaning-grep | Per-line Noul; AND/OR/NOT on *thresholded* bits | Thresholds / `--level`; cross-lingual; Semgrep.dev collision; not a gate | jev-semgrep (MIT LICENSE; GitHub NOASSERTION; 0.94/0.98 *theirs*; **51★** ephemeral). `notes.md` §86 | +| Active-learning triage | Confidence routes accept / teacher / human | Soft-label full distributions; real outcomes stay training targets; do **not** distill Jev as teacher | jev-triage (MIT; ~68% ceiling anti-pattern) | +| Evaluation-model-first SDK | predicate / classifier / rubric as data | check / evaluate / filter / partition / rank; cancellable; never auto-retry | sysone-help/sysone (MIT TS; first adapter Jev via Vercel AI Gateway). **Not** hraness/sysone (loopback gateway) | +| Evidence-gated question pack | accuracy / ECE / cost / latency on a pinned version | Pack is `provisional` until evidence.md; `unknown` mandatory; `jevassert check` offline from recordings | jev-packs (CC0; nine verified *theirs*) + **jevassert landed** (Apache-2.0; Action `@v0`). 2,990-case matrix: Jev/Sonnet 5 accuracy tie, Jev better calibrated 7/9, ~250× cheaper *theirs* | +| Runtime authorize (evidence ≠ authority) | Six narrow semantic Nouls | RBAC/schema/limits in code; positive p never overrides a hard fail | actiongate-jev (Apache-2.0; slogan: Jev supplies evidence, code owns authority; 500-case is label-baseline, not accuracy) | +| Ranking ≠ calibration | AUC vs ECE/Brier vs human rates | Recalibrate on labelled domain data; do not threshold raw p | does-jev-confidence (8,000 judgments; stated ~75% vs human ~10%; ~96% ECE removed) + jevcal (~100 rows) | +| Hot-click CU (indexed viewport) | operation + per-op target in one request | Code owns observe/execute/stale-ref/loop/`--until`; text model only for type; never guess fill | ego-jev (MIT; HN 4.9s vs 9.7s / wiki 5.4s vs 10.1s *theirs* n=3; not a bench). Cousin of jev-ultrafast | +| Framework-agnostic compact + gate | keep/drop per message + Foreman Nouls | Pins/dedup/regex floor in code; never rewrite; compaction fail-open if Jev down; safety fail-closed | jev-compactor (MIT; later product-arm **73%** / 350 ms / 4 of 4 *theirs*; 30–250× cheaper than shipped summarizers). Claude Code: fast-jev-compaction; OpenCode: fast-jev-opencode | +| Local-rules-then-remainder feed | four remainder Nouls after `rules.js` | Collapse not delete; auto-hides need human confirm before they become examples | x-reply-filter (MIT; 3-sample e2e 0.90/0.93 vs 0.08/0.10). Cousin of bohutang/sift | +| Control-plane combinators | Then / Gate / Vote / Cascade / Weighted + Router / Loop / Retry / Fallback / Memory | AND/OR aggregation stays in code (do not multiply); trace is the oscilloscope; Fallback is the fail-closed node | [jev-combinators](https://github.com/voidning/jev-combinators) (renamed from decision-combinators; TS; package MIT / GitHub SPDX null; npm 0.1.0; no measurements). Digital-design metaphor ≠ literal AND/OR. Not a new Jev API | +| Skill-library VOI | two-pass Choice then fit Nouls; none-of-these | Hook fail-open; Quill prefilter >254; advisory; local replay | skillranker (Rust; 52★; MIT + rider / GitHub NOASSERTION). Correction vs §7: hook is not fail-closed | +| Receipts-not-leaderboard map | hold vs break with API receipts | Schema-valid ≠ correct; retrieve first | jev-capability-atlas (10★ this pass; axis already §49; this hour is the eval-integrity cluster) | +| Jev vs thinking-budget small models | 100 four-choice tasks × 3; freeze questions | Attach the thinking condition; exploratory not preregistered | jev-frontier-100 (MIT; Jev 77.0% vs Qwen3.5 4B/2048 96.7%; 4B off 56.0%). Not a ceiling | +| OOD calibration / AUC ≠ ECE | ECE with noise floor; refit T by type | Calibrate per question; do not threshold `confidence` | jev-ood-calibration (MIT; 900 tickets; ECE 0.107 = 4.4× floor; priority 44.7% / mean p 0.74 / T 3.40) | +| Replayable evidence≠authority gate | policy first; Jev remainder; allow/review/deny | Missing Jev → Review; hard denials stay hard on replay | turnstile (Apache-2.0; experimental alpha; no npm). Actiongate-class clone | +| MLX one-pass schema→JSON | per-field probs in one forward pass | Softmax ≠ Noul; no local leaderboard yet | jevmlx (MIT; 28★; Apple Silicon replica economics). Distinct from system-one-benchmark Harbor table | +| TLA+ consensus around a noisy oracle | five paraphrased votes; stability gate; quorum 3/5 | Escalate when unstable; TLA+ owns the protocol; never confidently wrong | jev-labs (MIT; 1,080 golden 0 wrong *theirs*; TLC 1,049,750 states / 0 errors; synthetic, not clinical) | +| Advance / coverage ledger | Strike fills Candidates; Seal advances | coverage.path auto\|code\|human\|escalate visible; Effects locked while escalations open; mint ≠ product brain | seal (MIT; BEYOND-JEV.md; zero runtime deps) | +| Sureness over a probability vector | max_prob / margin / entropy / gini / perplexity | Bands are policy; Choice confidence = max_prob (generous) | how-sure-is-jev (MIT; zero-dep; 60-q reverse-engineer) | +| Scored Jev-class bake-off | Capability / Speed / Cost → Main Score | Calibration reported, **not scored**; native vs verbalized; partial runs not ranked | jevbench v1.1 (MIT; unofficial; Jev 1.13.0 Main 87.6 *theirs*) | +| Pre-review typed PR gate | should_review / risk / route / touches_secrets | Operator-owned thresholds; secondary LLM only on human-review + elevated risk | ci-gatekeeper-bot-jev (`package.json` MIT / GitHub SPDX null; 504–629 ms *theirs*) | +| Codex MCP host adapter | select_capability / search / triage Nouls | Caller supplies catalog; lexical fallback; scores advisory; ranking unbenchmarked | jev-in-codex (MIT; experimental MVP). Distinct from jev-routing (not MCP) | +| Stop-hook attention redirect | eight risk Nouls on a redacted turn diff | Fail-open; assist = one reinspect; 0.85 uncalibrated; not a merge blocker | jev-preflight (Go MIT; 1★; Claude Code 2.1.267 owner-run) | +| Pre-send tool-result views | smallest view that still serves the next step | Code builds outline/focus/testlog/…; recall restores lines; code full unless confident | dizk/jev-lens (MIT; 79% fewer tokens / 500 SWE-rebench *theirs*). Distinct from rashedInt32/jev-lens | +| Observational memory | keep / kind Choice | Verbatim ledger; kind-keyed files; model-free compact; failed Jev does not drain buffer | pi-observational-memory-jev (MIT). Sibling of fast-jev-compaction / jev-compactor | +| Independent open-Jev class | LM / vision / voice finite-choice+prob; JevPick menu decode | NOTA; execute/confirm/reject; `/v1/systemone` wire-compat ≠ replica | openvons (Apache-2.0 LICENSE / GitHub SPDX NOASSERTION; 7★; JevPick 3.2–4.8× *theirs*) | +| Same-intent VOI cache | `same_intent` Noul + pick candidate; exact SHA-256 first | Policy bypass stream/tools/multimodal; **fail-open** to upstream | jevcache (MIT; n=100 *theirs*: 0 FP / precision 1 / recall 0.38 / fpr 0 vs Jaccard@0.35 fpr 0.48). Not in v0: streaming HITs | +| Worth-your-attention VOI | read / skim / save / skip from typed answers | Templates never prose; feed batches ≤12; 7-day cache | [ThinkyMiner/Winnow](https://github.com/ThinkyMiner/Winnow) (MIT; 80%/90% *theirs*). **Distinct from** [kevinpita/winnow](https://github.com/kevinpita/winnow) (context sieve) | +| Jev WHETHER / Python HOW / LLM WHAT | compact original chunks; suppress duplicate observational tools; skip next main-model | Fail-open; skip-next needs a Hermes core patch; aggressive defaults | hermes-jev-router (Python; license null; community plugin, not vendor). Cousin dizk/jev-lens + jev-compactor | +| Typed escalate/continue/abort baton | LLM writes; Jev judges; control returns typed | Gate `allow` never grants; fail-open; inverted loop wakes LLM only on escalate | jev-handoff (MIT; alpha v0.1; three backends; Vercel drops confidence). Same author as wakegate | +| Adversarial browser explore | six oracle Nouls + severity Score + next-action Choice | Playwright executes; sample from the distribution not argmax; fail only high conf **and** high severity | browser-jev (TS; license null). Visual blind. CI exit 1 on non-baselined findings | +| Local OpenJev `/v1/decide` | Qwen3-0.6B + LoRA + scalar head | **Not** a TypeSafe drop-in; not RLCD; distinct from hraness/sysone OpenJev runners | IamBusy/OpenJev (Apache-2.0; 45/60 vs v0.2 39/60 *theirs*; Hub OpenJev-Branch-v0.3) | +| SemIf `/v1/systemone` runoff | choice/score/noul from logits; no option ceiling | Wire-compat ≠ replica; confidence inferred; runoff is a product not a softmax | semif-serve (pyproject MIT / GitHub SPDX null; RTX 3080 Ti 1164 ms vs hosted 178 ms *theirs*) | +| Conflict ≠ ignorance | Noul collapses; named Choice escape separates | Binary Choice without escape is lexically biased | jev-typed-evaluation-collapse (license null; NCML field note v0.3 *theirs*) | +| Decision-as-memory flywheel | log every typed decision as a memory event | Append-only; 6 rows is a schema not a corpus | DGUI_HYPERMEM-JEV (HF MIT; sibling INSTRUCT_JEV) | +| Toolbelt sensors (not policy) | local-rules-then-Nouls; prepared reviews; guardrail+intent demo | Not a cert; reviews never stop commands; bench summary 404 this pass; **rh-guard owns reward-hack** | jev-security-scan (MIT) / jev-decisions (MIT; 1★) / TeoMastro (license null) | +| Record/replay pack CI | accuracy / ECE / Brier / cost / latency vs golden | Record outside CI; `check` offline; exit 0/1/2; McNemar compare | jevassert (Apache-2.0; SPEC v0 with jev-packs; adapters typesafe/openai/anthropic) | +| Failure-finding judgment arena | pairwise chosen/rejected; native vs verbalized p | Reviewed failure atlas, not a leaderboard; mock runs never enter a ranking | chenmingtang830/jevarena (Apache-2.0; JevJudge-Bench harness **not** findings). **≠** meetr1912/jev-arena | +| BBQ stereotype/uncertainty/cost | 3-way Choice on BBQ passages | Frozen instruction; unknown is the abstention option; not a bias cert | jev-bbq-experiment (R; license null; 58,492 / 97.28% / $0.3429 *theirs*) | +| Decider ≠ executor agent | next_action / progress / risk / stuck / done | LLM fills args only; risk≥0.5 pause; stuck ladder 2 Jev / 0 steps | jeffrey (MIT). Distinct from jev-handoff baton | +| Sentence-as-rule lint | ast-grep `rule:` × sentence `ask:` | Matcher silent-fail (over-match); Jev loud; fail-open no-verdict | mizchi/jevlint (MIT; 13/15 1.00/1.00 *theirs*). **≠** huntedman/JevLint | +| VOI hunk prune before generative review | per-hunk actionable-finding + required-context | Safety escarpment always keeps; cost not quality | prune-review (README Apache-2.0 / GitHub NOASSERTION; 22-run 1.18% / 305% outlier *theirs*) | +| Whole-repo intent vs the diff | VERIFIED / VIOLATION / UNKNOWN / NOT_APPLICABLE | Empty search ≠ proof; CLI works, Action not written | jev-intent-review (MIT/Apache-2.0; under construction) | +| Persist constraints across compaction | KEEP/LIFT/NARROW/EXCEPTION/REPLACE/UNKNOWN | Jev never writes policy; resources from user words; fail-open | pi-heed (MIT; 3★; v0.8.0+Jev 98.5% / 0 false block *theirs*) | +| Open replica substrates | `/v1/systemone` on Rust/WebGPU, Clojure/Jolt, CPU SemIf, ONNX NLI, GLiNER2 spec, prompted-JSON Bun bridge | Softmax / generated JSON ≠ Noul; isolation/byte-parity/agreement are the tests; spec ≠ product; wire-compat ≠ logit-equiv | grande (license null; JGLUE *theirs*); laya-jolt (Apache-2.0 byte parity); leesk212/JEV-CPU (Meanblock 404); kunchenguid/local-jev (done 30%/shape 57%); githubnext/localjev (**261★**; prompted JSON ≠ razorback16 logits); Eran-BA spec ≠ jeff | +| Harbor SGR-judge contract | Jev vs schema-guided LLM judges; invalid = FN | Frozen SLA-150; human labels; cost/latency first-class; incomplete cohort ≠ headline | jev-judge-bench (README MIT / GitHub SPDX NOASSERTION; 21 offline tests; canaries ≠ quality; **no quality headline yet**). **≠** jevarena / jevbench | +| Hand no-text steps to Jev | did-it-work / which-next / severity / safe | Writing stays with the LLM; PreToolUse deny/ask **fail-open**; Vercel reconstructs confidence as margin | jev-use (MIT v0.4.1; p50 220 ms; 186 vs 2,672 ms; gate 12/12; first loop 17/20 then 0/20 at 0.4 *theirs*). Same author as jev-handoff. **≠** jev-ultrafast | +| Pi System-One control plane | task/model/tool/failure/retry/context/skill/memory/review/GUI | Compaction never writes session; GUI < threshold → unknown; unresolved failures KEEP | pi-jev-control (TS; license null; v0.3.0 private; no live quality numbers). Distinct from omp-jev-extensions / jevons / pi-heed / pi-om | +| Generation as a tree of Choices | one typed question per word, then rank texts | Model never free-generates; embedding tree in code | jev-gpt (Python; license null; ~400 calls / 75 s / 2¢ *theirs*). Architecture demo. Distinct from jeffrey | +| OpenRouter recipe atlas | one call, many narrow questions; policy in code | Samples 16–36, not benches; pick don't extract; review band around every cut | jev-cookbook (JS MIT; 1★; 425 calls / $0.015; browser 5/6 *theirs*). Code prepares, Jev answers | +| Personal-history feed (no social graph) | Choice distribution over candidates = ranking | History local; one request per batch of ten; dwell = nearest-to-middle | jevfeed (JS MIT; 17 tests no network). Distinct from ThinkyMiner/Winnow and kevinpita/winnow | +| Competing NAR claim-audit | Choice/Score/Noul NAR; dual-channel ECE | Like-for-like channels; throughput ≠ latency; n=2000 CI before "SOTA" | openJev-verdict-2.0 (README Apache-2.0 / GitHub SPDX NOASSERTION; 77.10%/0.0636/0.0144 *theirs* unverified; **open PR #1**). **≠** IamBusy/OpenJev `/v1/decide` | +| Empty compaction-proxy skip | slogan only | No files, no fail polarity | IPECTER/jev-context-pruner (409 empty) **and** IPECTER/jev-runway (LICENSE-only). Sibling contrast only | +| 1-token logprob local endpoint | label-mass over caller-enumerated options | Arithmetic/numeric rules in code; 503 ≠ "no"; coverage ≠ correctness | chakuho (MIT; GUI 336 *theirs* 27B 95%/92% vs Jev 89%/82%; `__none__` 97% vs 8B 10%). Cousin jevify / TypeAR / pcdServer / jevmlx. Softmax ≠ Noul | +| Open replica inference engine | segmented forwards + Jev wire | Prefix reuse; families nanojev/decider-2b/laya; MPS only | jevinf (MIT; 2.57×/2.27× 100% argmax *theirs*). Wire-compat ≠ replica | +| Unofficial Elixir HTTP client | typed Noul/Choice/Score over Req | Policy in `cond`; confidence ≠ P(correct) | typesafe-elixir-sdk (MIT; 1★). **≠** dannote/jev OTP peer | +| Files-to-read VOI (rename + n=16) | BM25 shortlist then Jev packet | Index once; agent still Reads; read-only | jimmyhealer/jevex (was jev-semantic-explorer). n=16 160s→69s / $8.74→$3.13 / 16/16 *theirs*; n=8 finish 1/8→6/8 stays | +| Commit pre-review attention≠verdict | seven Nouls + headline Choice; six regex | Middle band = review; Nouls decide; hook fail-open on instrument failure | commitjev (MIT; 0 false on 5 clean *theirs*; small control; same owner as jev-orderby-bench) | +| Hermes plugin branded as Jev | Choice/Noul/Score *shape* | Not TypeSafe; not a Noul | hermes-plugin-jev (README MIT / GitHub SPDX null). Agnes 3.0 Flash chat-completions. **≠** hermes-jev-router | +| Pi verbatim summarizer replacement | one Noul per paired tool call | Keep-windows/pins in code; fail-open to LLM summary if <25% saved | pi-jev-compact (MIT). **≠** vava-nessa/pi-jev-compaction. Pair pi-jev-control / pi-heed | +| Decision-native inbox | nine typed signals | 100-point policy + SLA/tier in code; humans own ambiguity | mailordinal (MIT). Cousin jav-email-cascade. Not affiliated with TypeSafe | +| Unofficial Jev CLI (not ready) | planned exit-status semantic `if` | No release; do not copy MCP add | jev-cli (Apache-2.0 OR MIT; 0.0.0; 17 issues). **≠** jevql | +| Multilingual Laya class expansion | Choice/Noul/Score, mmBERT-base | Route by script before the forward pass; refit T | laya-multilingual (Apache-2.0; 322M; MASSIVE 0.366/0.387 vs English 0.227/0.733 *theirs*; ships uncalibrated) | +| Schema-conditioned encoder scorer | scalar logit per (state, candidate); code softmaxes | Peaked p = ranking | mobarmg/jev-schema-scorer-deberta-v3-large (Hub MIT; GitHub 404; v2 Choice 0.841 *theirs*) | +| Host-adapter surface delta | Choice next-tool + done-Noul | Same binary; more hosts | jev-routing now lists Cursor Agent CLI / Devin CLI (still not MCP; already §44) | +| Productized System One HTTP | caller labels → label + calibrated p; batch `{id,text}[]` | LLM chains are fallback only; policy stays in code | classifier-dev (MIT; **185★**; https://classifier.dev). 400 headlines **650 ms** *theirs*; packing 100 = one-at-a-time. Distinct from ask-jev-ai wall | +| Escalate-under-threshold (smart tier) | re-ask single-label p<0.7; mark `escalated` | Multi-label **ignores** tier (re-judge worse, 23 s) | classifier-dev. Emotion ≥0.9 → 82% / <0.5 → 29% *theirs*. gemini-3.8-flash 87.5→90.0 / 61.8→63.7; other flashes no better. Cousin jev-use | +| Measurement-first public bench | vs_jev / single / escalate / multi-label | Site table = tracked JSON; read eval/README first | classifier-dev. Multi-label F1 **0.887** / **230 ms** vs cascade **0.799** / 1.5 s *theirs* (eval 232 ms). n=7 train-on-test; ~0.03 coin flip. Not a Harbor taskset | +| Silent-fallback honesty | digest names the model that answered | `FALLBACK` marker; alerts on quiet chain | granite-4.0-h-micro F1 **0.546** vs advertised ~**0.800** *theirs*. rh-guard owns the gate; dinostomp owns instrument-not-score | +| Evidence-synthesis pointer (choxos) | Jev picks line ids; code copies verbatim | *Not found* / *Unclear* first-class; human tick never overwritten | choxos/jev-reviewer (MIT; **12★**; https://jevreviewer.xera.ac). **≠** egma-ai. 18-q template **4.6 s / $0.0101** *theirs* (spot check, not a validation study) | +| Two-pass Choice + Noul | relative “which line?” then absolute “does this line itself answer?” | Multi-row tables (Mean SD vs Median IQR) need both | choxos/jev-reviewer. Quotes = Noul ≥ 0.5 *theirs*. Cousin Stagehand extract / jev-sift | +| Institutional local `/v1/systemone` (GitHub Next) | TypeSafe SDK drop-in on DiffusionGemma via Chat Completions | Wire-compat ≠ logit-equiv; entropy-conf is generated | githubnext/localjev (MIT; **261★**). **≠** kunchenguid/local-jev. **≠** razorback16/openjev structured-read. **≠** IamBusy/OpenJev `/v1/decide`. Do not copy bun / `.env` | +| Prompted-JSON bake-off (Harbor-shaped) | AG News / BoolQ / SST-5; 5 models × 120 × 2 lengths = 1,200 | Prompted pipeline, **not** logits; no definitive winner; not calibrated | githubnext/localjev eval *theirs* M5 Max: Qwen3.6 short macro **76.7%**; Gemma 4 26B-A4B **75.0%** (SST-5 MAE **0.533**); DiffusionGemma **74.2%**. Qwen vs Gemma 26B = 2/120. Long-input both **69.2%**. Serving default unchanged | +| Runner gap vs structured-read | Chat Completions host ≠ OpenJev parity | Seeded canvas + read-only denoise + selected-token logits | LM Studio cannot load DiffusionGemma (18 Sep 2026 *theirs*). Cousin djev-spark already §36 | +| Laya packaging (not a new species) | Choice/Score/Noul NAR + Router over three Hub ckpts | Script-before-p; auto_task_detection off | NandhaKishorM/laya (Apache-2.0; **710★**; PyPI). Weights: convaiinnovations/{laya, laya-multilingual, laya-typed-decisions}. **≠** TypeSafe `/v1/systemone`. Do not copy pip | +| Where Jev still leads | High-cardinality Choice; soft-acc; raw ECE | Token budget `head_max_len`; 255 options | Banking77 Jev **0.870** (72) vs Laya **0.425** (77, ~3–4 tok/label); soft-acc 0.580 vs 0.471; raw ECE 0.144 vs 0.213 *theirs* (third-party Jev rows unpublished-here) | +| Where Laya leads on *their* T4 card | Latency; post-T ECE; multilingual router | Route by script before p | 1q **32.8 ms** vs Jev p50 236–276 ms (~7.8×); post-T ECE **0.081** vs 0.246; Khmer 0.000@0.952 is why Router exists | +| External openjev census (tweet, not scores) | Named list of ~18; first leaderboard promised "today" | Class-boundary + completeness watch; likes ephemeral | [@airesearch12](https://x.com/airesearch12/status/2101259522933186879) (Florian S / Benchmark Heaven). **≠** jevbench v1.1. Watch [jev-models](https://benchmarkheaven.com/jev-models); scored card is sibling. Do not copy Stripe | +| JP application genre atlas (tweet, not scores) | High-star apps by genre; SAM 3.1 + OpenRouter Jev noted | Stars research-time; not verified evals; likes ephemeral | [@studio_yebisu](https://x.com/studio_yebisu/status/2101065176069886152). **≠** @airesearch12 class census **≠** v1.2 board. Do not dump the 30 repos | +| External pedagogy / how-to-apply (article, not a product) | LLM hammer; code owns branches; parallel questions; schema-safe ≠ correct; shadow + questions-as-code | 200×/400× TypeSafe ceiling; text-only; do not rebuild the agent first | [@akshay_pachaar](https://x.com/akshay_pachaar/status/2101037514945597645) “Jev Clearly Explained”. **≠** official docs **≠** Flavio Copes **≠** LangChain harness **≠** AgentGhost. `notes.md` §85 | +| Meaning-grep dedicated (proposition ≠ embedding) | Line+question Noul; boolean AND/OR/NOT after threshold | Contrast-set refund; calibrated ~0.5; no index; EN safer near threshold; **≠** semgrep.dev; not a gate | [uehaj/jev-semgrep](https://github.com/uehaj/jev-semgrep) (MIT LICENSE / GitHub NOASSERTION; **51★** ephemeral; HEAD `21120e9`; README SHA `923e6a5`). **≠** jevgrep **≠** jev-combinators. `notes.md` §86 | +| Class-boundary on a public list | GLiNER2 + routers counted as openjevs | Locate/categorize ≠ Noul; route ≠ replica ECE | GLiNER2 (Fastino); Succinct Router 14M; jev-model-router, Director, Loki. Qualify open-jev Dasein vs JoshuaSP; OpenJev razorback16 vs IamBusy | +| Incomplete census vs watch | Absence ≠ out of class | Completeness is a board watch item | Laya / localjev / kev / TypeAR / openvons / chakuho / jevinf / grande / laya-jolt / blackwood / classifier-dev | +| Harbor honesty watch (pre-score) | What the board must disclose | Calibration on/off rank; cost/latency assumptions; silent fallback; partial runs | Kinship with §67 v1.1 (cal off Main Score) and classifier-dev FALLBACK. Soft-score-as-hard-rank is a *design*. **Promoted:** answers are now Empirical as §78 | +| JevBench v1.2 geometric-mean product | Intelligence × Calibration × Speed × Cost, 25% each | Weak axis cannot be bought back; other views reorder ranks | Live [jev-models](https://benchmarkheaven.com/jev-models) scored 19 Sept 2026. Jev **75.3** / SemIf **74.6** (−0.7) *theirs*. Cal **ON** rank (delta from §67). **≠** tweet census **≠** v1.1 87.6. `notes.md` §78 | +| Weight sensitivity (same axes, not the Score) | Balanced no-cal / Emphasis Accuracy / Speed / Cost | SemIf #1 without cal; system-one-open #1 on cost; Jev #5 on cost | *Theirs*. Limits: "The weights are a choice." Do not treat geo-mean as physics | +| Option-order fragility | yes/no answer-judging 72% → 21% when A/B reversed | Small models are very sensitive to option order *theirs* | open-alternative-jev ranked on author's `A. yes, B. no`. Cousin of paraphrase-brittleness. Both runs in `results/v1.2/runs/open-alternative-jev/` | +| Instruction models in the class table | Typed decision task, not architecture purity | Luna/Gemini/DeepSeek/Qwen3.8 JSON-schema; Needle 3 tool-calling | Luna I **96.8** rank **#7**. Needle 3 C none → 0. OpenJev = razorback16 DiffusionGemma ≠ IamBusy | +| Harbor honesty (×2 / est.) | Name assumptions; ranks are configuration-specific | Self-host latency ×2 (+0.15 s) is an assumption; many costs est. | Production APIs unadjusted. Partial not ranked. Kinship classifier-dev FALLBACK | +| Laya / GLiNER2 / apps gaps | Absence ≠ quality; mapping ≠ scored | Laya absent (not named-excluded); GLiNER2 needs normalization; apps out | Completeness vs watch. Qwen3.8 27B Chutes TEE **≠** Archer | +| Apply-the-five (hourly 0842, already folded) | Wire≠logit · product+FALLBACK · packaging honesty · pointer-not-generator · leaderboard VOI | Do not re-card §73–§78; skip thin noise | `notes.md` §79. Compose, don’t dump | +| Hard-gate Noul as PR/quality (skip) | Soft sensor used as a merge seal | Soundness theater unless an exact envelope already proved the act | totally-tim/jev-gate (0★) / claude-jev-warden (1★). **≠** jev-gateway / MongLong0214/jev-gate / jev-gate-student-b. Do not copy action.yml | +| S1 keeps flying / S2 one-use (delta) | Typed flight Choice; async planner on low p | Physics/collisions; no stall; consume-mark; Local ≠ localjev | khordoo/jev-reflex-autonomy-lab. Seed = geometry. 20% still soft. No pixels. `notes.md` §80 | On-device / Home Assistant / mobile are newly-feasible via the economics inversion, not proven ports of every app. Named placements this hour (`notes.md` §33): `Friedjof/jev-mobile` (USB Android, Mobile MCP task delegation, Jev sees only prevalidated candidates); `jcpsimmons/jev-macos-loop` (local OmniParser/OCR/AX; text-only Jev; pixels stay on the Mac; Finder -demo independently verified). HA-Jev unchanged. Do not copy env, MCP +demo independently verified). HA-Jev is now a real card +(`notes.md` §68; **17★**; not for locks/heaters). Do not copy env, MCP URLs, or install steps. Reproduce/open heads (`rongxinzy/LightJev`, openjev family, [`convaiinnovations/laya`](https://huggingface.co/convaiinnovations/laya), encoder [`open-jev-deberta-v3-large`](https://huggingface.co/com-kotobalabs/open-jev-deberta-v3-large), LoRA [`jev-gate-student-b`](https://huggingface.co/SargeDev/jev-gate-student-b), -companion packaging [`laya-typed-decisions`](https://huggingface.co/convaiinnovations/laya-typed-decisions)) +companion packaging [`laya-typed-decisions`](https://huggingface.co/convaiinnovations/laya-typed-decisions), +GitHub/PyPI face [`NandhaKishorM/laya`](https://github.com/NandhaKishorM/laya) (**710★**; Router; not a new species; `notes.md` §76), +[`jaredpalmer/kev`](https://github.com/jaredpalmer/kev), +[`BlackwoodAI/blackwood-rlcd`](https://huggingface.co/BlackwoodAI/blackwood-rlcd)) are evidence that the *interface* (Choice/Score/Noul, or yes/no logits as P(relevant)) is the transferable part — not a request to implement a backbone or a second API skill. Laya: self-hostable, text-only, 512 -tokens/question; vendor benches vs Jev are **claims**. Encoder open-jev: -public gold, OOD drop. LoRA student: teacher-copy. Hume's 27B +tokens/question; vendor benches vs Jev are **claims** (this hour the +author published the vs-Jev table *and* named it third-party / +unpublished-here — `notes.md` §76). Encoder open-jev: +public gold, OOD drop. LoRA student: teacher-copy. **kev**: public gold, +pointer readout, System One API drop-in; ID ECE only; not a teacher-copy +(`notes.md` §45). Hub fetch: +[`jaredpalmer/kev-0.5b`](https://huggingface.co/jaredpalmer/kev-0.5b). +**blackwood-rlcd**: open multimodal RLCD, Jev-compatible +shim, CC BY-NC; Jev still leads general text; not Archer Watch +(`notes.md` §46). Hume's 27B decision-model drop is **Watch**. Closed calibrated API vs open weights -is a self-eval tradeoff (`research/notes.md` §18, §33). When-to-use +is a self-eval tradeoff (`research/notes.md` §18, §33, §45). When-to-use axes: `judgment-class.md`. TypeSafe remains the documented *exemplar*, -not the class monopoly. GLiNER (locate) / GLiClass (categorize) / -GLiNER2.5 (local multi-head), listwise, and vision families: +not the class monopoly. **Local CUDA/PyTorch replica this hour:** +[`Mintzs/jevify`](https://github.com/Mintzs/jevify) — Qwen2.5-1.5B +Choice/Score/Noul *shape*; **uncalibrated likelihoods ≠ Noul**; no +LICENSE this pass (`notes.md` §55). **Local contract drop-in this hour:** +[`us/jev-local`](https://github.com/us/jev-local) speaks `/v1/systemone`; +**default scorer is a deterministic stub** until `JEVLOCAL_SCORER=hf` +(`notes.md` §48). **ONNX replica path:** +[`Mattepiu/laya-onnx`](https://huggingface.co/Mattepiu/laya-onnx) — do +not copy the inherited vs-Jev table. **Complete browser +int8 cousin (distinct):** +[`gqgs/laya-onnx`](https://github.com/gqgs/laya-onnx) +(496.8 MiB; conversion smoke, not accuracy; `notes.md` +§64). **Local ModernBERT approximation, not +equivalence:** +[`kunchenguid/local-jev`](https://github.com/kunchenguid/local-jev) +— ONNX ModernBERT; measured done 30% / shape 57% vs Jev *theirs*; +`confidence` omitted; **not** equivalence (`notes.md` §70). Distinct +from jev-local stub and jeff. **This hour's substrates (not Archer):** +[`bokuweb/grande`](https://github.com/bokuweb/grande) Rust/WebGPU +kev-shaped branches; [`jlt-commons/laya-jolt`](https://github.com/jlt-commons/laya-jolt) +Clojure byte-parity Laya; [`leesk212/JEV-CPU`](https://github.com/leesk212/JEV-CPU) +SemIf on CPU (Meanblock 404); [`Eran-BA/Jev_from_GLiNER2`](https://github.com/Eran-BA/Jev_from_GLiNER2) +spec-only GLiNER2 decide adapter. GLiNER (locate) / GLiClass (categorize) / +GLiNER2.5 (local multi-head; extractive compaction is a named *job* on +that family, `notes.md` §50; computer-use selection is a *different* +named job on GLiNER2 `gliner2-multi-v1`, `notes.md` §52), listwise, and vision families: `judgment-class.md`. ## Design-card extras for mixed systems diff --git a/.agents/skills/augustus/references/optimizer-integration.md b/.agents/skills/augustus/references/optimizer-integration.md index ebf3b7f..e54f4cd 100644 --- a/.agents/skills/augustus/references/optimizer-integration.md +++ b/.agents/skills/augustus/references/optimizer-integration.md @@ -13,7 +13,9 @@ signatures, and `typesafe-ai` plus the live docs own Jev's request body. Do not write either from this page. DSPy and Ax tune the LM-program slice only. They are never the primary System One calibration score; that seat is a jevals-shaped labeled suite, and a product loop is a -Harbor taskset (`validation.md`, Eval & hill-climb). +Harbor taskset (`validation.md`, Eval & hill-climb). Shared bake-off +exemplar this hour: [`pngwn/open-jev-laya-bench`](https://huggingface.co/datasets/pngwn/open-jev-laya-bench) +scores ECE/NLL/Brier — not an LLM-as-judge paragraph (`notes.md` §46). ## Judgment: what these optimizers may climb (Hypothesis) @@ -37,7 +39,9 @@ prompt loop: schema, criteria, and policy thresholds, scored on labeled eval (jevals) — not a search over a decoder. Open recipes such as Nimble: climb data curation and LoRA, measured on holdout ECE and agreement. Nimble's published holdout is agreement on synthetic -labels, not a measured ECE (`judgment-class.md`). +labels, not a measured ECE (`judgment-class.md`). kev: climb LoRA / +public-gold labels; the published ID ECE is a receipt, not your +workflow (`notes.md` §45). No call shape in this paragraph. The adapter notes below stay names of seats, not a request you copy. @@ -118,6 +122,46 @@ hashes. PoC benchmarks on 3 cases without caching are a lead, not a result. - Nothing yet covers optimizing *against* Jev as the metric model end-to-end; if you build it, measure judge variance first (see above). +## Typed control plane around DSPy (not more knobs) + +Ax and DSPy remain LM-program climbers. A **typed control plane** +is deterministic code *around* that program: +classifier → ontology validation → security override → +confidence thresholds → state-machine transition → tool +allow-list. The LM may draft wording **after** route and action +are fixed; it cannot add a route, change the action, or invoke +an unapproved tool. +[jev-dspy-control-plane](https://github.com/manikanda-kumar/jev-dspy-control-plane) +is the Harbor-shaped bake-off of that split (OpenJEV / DSPy / +JSON Schema share ontology, dataset, metrics). Offline smoke +uses a heuristic + labelled contract stubs — a high score is +plumbing, not quality. Metrics named: intent/sub-intent +accuracy, invalid-output/policy-violation, abstention/coverage/ +selective accuracy, Brier/ECE, consistency, p50/p95, per-category +stress. Accuracy alone is not enough; a negative result is +valuable. Do not copy venv / `.env`. `notes.md` §59; +`validation.md`. + +## Specialist as metric vs few-shot as classifier + +A typed judge used as an optimizer metric **reads the +probability** (Brier/ECE, risk-coverage, expected cost). That +is the Domain-jev-maker placement: train a specialist when the +downstream consumer is the distribution; few-shot hosted is +enough when the program only takes argmax. Do not substitute a +verbalized `"confidence"` (jav-email-cascade gen-json mock) for +a native Noul. `notes.md` §60. + +**Do not distill Jev as teacher of record.** +[jev-triage](https://github.com/ThyFriendlyFox/jev-triage) +logs full distributions as a *bootstrap* for a local student; +**real outcome labels** stay the training targets. Author +~68% ceiling compounds errors. Soft labels are features, not +the gold. Distinct from Domain-jev-maker (independent gold) +and from PAW coupling (1) below — if you use Jev as a labeling +teacher, measure the student against real outcomes and cut +the cord when it wins. `notes.md` §61. + ## ProgramAsWeights: materializing a Jev judgment locally (Hypothesis) PAW (programasweights, pre-dates Jev — Python SDK 0.4.6, Mar 2026 repo, MIT) @@ -136,11 +180,13 @@ a local artifact for high-volume/offline/zero-latency paths. No measured Jev+ PAW integration exists in the wild (checked the full 187-repo archive), so this is Hypothesis-grade. Two candidate couplings: -1. **Jev as the labeling teacher.** Use Jev fan-outs to score a labeled set - (its calibration is the reason to trust the labels), pass `examples=[…]` - into the PAW compile/finetune compiler (`paw-ft-bs48`), then serve locally. - Jev = oracle, PAW = distilled student. This is ordinary distillation with - an unusually cheap teacher. +1. **Jev as the labeling teacher — with teeth.** Use Jev fan-outs to + score a labeled set, then train a local student on the *full + distributions*. **Do not** treat those labels as teacher-of-record + gold (jev-triage ~68% ceiling). Real outcome labels remain the + target; cut the cord when the student wins on held-out outcomes. + This is ordinary distillation with an unusually cheap *filter*, + not a substitute for independent gold (Domain-jev-maker). 2. **Jev as the calibration gate on PAW.** Shadow both on live traffic; a calibrated Jev judgment arbitrates disagreements and the disagreement rate is the drift signal for when to recompile the PAW program. Threshold diff --git a/.agents/skills/augustus/references/question-design.md b/.agents/skills/augustus/references/question-design.md index 54cf681..1e65ad1 100644 --- a/.agents/skills/augustus/references/question-design.md +++ b/.agents/skills/augustus/references/question-design.md @@ -17,7 +17,7 @@ request, and treat a stale pin as a prior, never a setting. 2. One question per judgment. Split any question that weighs two properties. 3. Pick the primitive whose answer code acts on directly. 4. Build the smallest state that answers every question; compute in code whatever code can compute. -5. Put every question sharing the state into **one request** (speculative fan-out — parallel questions cost little latency; code ignores unneeded answers). Second requests only when later data depends on an earlier answer. +5. Put every question sharing the state into **one request** (speculative fan-out — parallel questions cost little latency; code ignores unneeded answers). Second requests only when later data depends on an earlier answer. Extractive / pointer: number the candidates in **code**; ask per-id Noul/Choice; copy verbatim. "Not found" is an option. The model never writes the quote. **Evidence-synthesis scale ([choxos/jev-reviewer](https://github.com/choxos/jev-reviewer), ≠ egma-ai):** fan-out every question over shared chunks, then a **second** absolute Noul ("does this line itself answer?") for multi-row tables; human tick never overwritten (`applied-mappings.md` §2; `notes.md` §48, §74). 6. Combine in code: branches, weights, confidence gates. 7. Test on labeled examples; read `probabilities` on the misses; revise one or two questions at a time. @@ -28,7 +28,7 @@ request, and treat a stale pin as a prior, never a setting. - `confidence` measures how **peaked** the distribution is — a property of the model's answer, not a correctness guarantee. - `score` is the probability-weighted mean of level numbers; 1.0 can mean certainty at level 1 or a split. Levels are weakly calibrated as numbers: threshold, rank, or round — **do not interpolate quantities** from a Score. - Jev does not count or do arithmetic or date math. Ask per-item Nouls in one request and sum in code; extract date parts with Choices (with a "not stated" option) and compare in code. -- Every answer stays inside the supplied options — code never parses prose. +- Every answer stays inside the supplied options — code never parses prose. **Conflict ≠ ignorance:** a Noul has nowhere to put "both and neither." Name those states as Choice options or the model will collapse them (typed-evaluation-collapse; `notes.md` §69). Same family as missing `other` → confident wrong. - Questions on one request never see each other's answers. - Text in the state can steer the answer; Jev does not treat state as hostile. State in criteria what counts; test injected and self-describing content before deployment. @@ -40,11 +40,12 @@ request, and treat a stale pin as a prior, never a setting. - Keep numerals-for-levels out of instructions ("Rate from 0 to 2" gives nothing to match); write the full question in `instructions` (the question ID never reaches the model); keep decision policy out of questions (policy lives in code). - Instructions accept prose or a structured form (the live docs own which keys that form accepts — do not write them from this page). The design rule is what transfers: name the sub-parts of the judgment explicitly instead of packing them into one sentence, and pass schemas/taxonomies as JSON, not as serialized strings. - When you catch yourself explaining what you meant after a wrong answer — that explanation is the missing half of the instruction. Add it. +- **Sentence as rule:** a natural-language sentence can *be* the criterion when a matcher already extracted the subject ([jevlint](https://github.com/mizchi/jevlint) ast-grep `rule:` × Jev `ask:`). Do not pack two properties into the sentence. Mechanical defects stay with the compiler; contradiction of a declared contract is the System One hole. Not SWE-only: any artifact that names itself (policy, checklist, form, recipe) can be a subject × a sentence. Qualify vs huntedman/JevLint. `notes.md` §70. ## Criteria shape - Criteria are an extension of the instruction and must ask the same thing, in the same direction (a Noul whose `true` side describes "no" performs worse). -- Choice options: contrastive `what` / `not_for` / short concrete `examples` (instances, not descriptions of instances). Add an `other` / `none-of-the-above` when the list may not cover inputs. +- Choice options: contrastive `what` / `not_for` / short concrete `examples` (instances, not descriptions of instances). Add an `other` / `none-of-the-above` when the list may not cover inputs. Skipping that hatch is not a style nit: the model will pick a listed option at confidence 1.00, and no downstream gate will see a problem (`notes.md` §46). Request-shape lint (wellposed / `tenbin`) puts the hatch on the offered set; **training must confront it as a wrong alternative too**, with varied wording, or the model learns "this wording ⇒ pick it" ([kev](https://github.com/jaredpalmer/kev) first-run shortcut; dedicated `none_of_the_above` eval; `notes.md` §45 delta). Corpus-scale cousin (maker claim, not re-run): SEO internal-link audit **584** placed / **139** refused because nothing honestly fit — Choice-with-`other` at catalog scale (`notes.md` §45–§46, §56). Computer-use cousin: Stagehand pick asks `best` (no none) **and** `strict` (with none; vetoes above 0.9); ambiguity **stops rather than guesses** (`notes.md` §57). - Score levels (2–10): describe **situations**, one dimension each, each standing alone (Jev sees neither the level's number nor its neighbors — "worse than previous" means nothing). No numerals. Levels may be objects `{"summary", "signals"}`. Give a rare extreme its own level when code treats it differently. - Composite scoring: one Score per dimension, normalize by `len(criteria)-1`, weight and combine in code. Change policy by changing weights — never by rewriting questions. - Taxonomy walk: one Choice per tree level, walk in code; each option's value is its subtree (direct children + sample leaves); follow several branches when probabilities are close. @@ -54,6 +55,11 @@ request, and treat a stale pin as a prior, never a setting. | Symptom | Likely cause | Fix | | --- | --- | --- | | Wrong answers, high confidence | Instruction read literally | State exact condition; put boundary cases in criteria | +| Wrong answers, high confidence, **nothing in state that could answer** | Bare recall / missing evidence | **Retrieve first**; put the passage in `state`. Atlas history: wrong@0.90 without context → right@0.97 with passage. Confidence gating on recall is not enough (Case A was 0.90 *and wrong*). `notes.md` §49 | +| Wrong answers, **dangerous-high** ECE on overlapping labels | Population calibration failed (blurred categories) | Do not threshold. DAIR Emotion: 48% acc / mean conf 0.819 / 16% p(correct)=0. Plot reliability on *your* labels | +| Wrong answers, **confidence ~1.00**, no `other` | Forced pick: the offered set does not cover the input; the model *must* choose | Add `other` / none-of-the-above. **Confidence gating cannot catch this** ([wellposed](https://github.com/suraj-phanindra/wellposed) live probe: unsubscribe email → `"support issue"` at 1.00 without `other`, `"other"` at 0.93 with it). Overlapping options collapse confidence (loud). `notes.md` §46. `tenbin` owns the lint skill | +| Residual `"other"` always picked (or never) | Training saw none-of-the-above only as the true label — a wording shortcut | Confront the hatch as a *wrong* alternative too; vary wording; eval present-vs-removed ([kev](https://github.com/jaredpalmer/kev) `none_of_the_above`; `notes.md` §45 delta). wellposed still owns request-shape lint | +| Question names a state path that does not exist | Dead reference; the API still answers | Lint the request (walk JSON). Structural, not semantic. wellposed recipe; do not copy the CLI | | Low-confidence Choice | Options overlap / none fits | `what`/`not_for`/`examples`; add `other` | | Low-confidence Score | Overlapping levels, two dimensions, thin state | Distinct-situation levels; split question; add state field | | Scores cluster mid-scale | Levels are degrees/numbers | One concrete situation per level; remove numerals | @@ -65,6 +71,62 @@ request, and treat a stale pin as a prior, never a setting. | Answer follows state text | Content steers the model | Tighten criteria; adversarial tests; confidence-gate the action | | Rewording trades one error for another | One question, several properties | Split into atomic questions | | Synonymous wording swings p / the act | Stimulus includes question text; no invariance promised | Paraphrase-pair eval; abstain or raise t; rewrite (`mappings.md` §17) | +| Question has no answer yet (edit 1 of 12) | Observation window is wrong: a turn-level property asked at edit time | Name when the evidence exists. Edit-phase vs turn-phase is a question-design cut, not a hook detail ([Abide](https://github.com/coldteadotai/abide): "added more than asked" is a turn rule). `notes.md` §47 | +| Review is green on the diff; the rest of the repo violates the stated intent | Observation window is the *diff*, not the places the intent applies | Search the whole repo after the change; one small question per place; **UNKNOWN** is cheaper than a false VERIFIED. Empty search ≠ proof ([jev-intent-review](https://github.com/yottayoshida/jev-intent-review)). `notes.md` §70 | +| Naming/comment "rule" as a paragraph the linter cannot prove | The sentence is the criterion; AST/ast-grep already extracted the subject | Put the sentence in `ask:`; matcher silent-fail vs Jev loud; fail-open if no verdict ([mizchi/jevlint](https://github.com/mizchi/jevlint); ≠ huntedman/JevLint). `notes.md` §70 | +| Catalog tagged "because it mentioned AI" | Criteria omitted what *doesn't* count | Add one exclusion sentence; 36/100 → 6/100 *theirs* ([jev-cookbook](https://github.com/nexibeo/jev-cookbook) TemplatesGrokBot). `notes.md` §71 | +| One severity Score bunches in the middle | "How bad" hides several yes/no properties | Split into concrete Nouls (cookbook log triage 4/7 → 7/7 *theirs*). `notes.md` §71 | +| Gateway returns no `confidence` | The statistic is missing, not "uncalibrated" | Reconstruct margin; lower the bar on *that* backend; tune on your traffic ([jev-use](https://github.com/shitianfang/jev-use) 17/20 → 0/20). `notes.md` §71 | +| README badge ECE vs a different channel | Like-for-like channels; n and CI | Dual-channel ECE is a design fork; do not put correctness-head 1.44% beside distribution 21.40% ([openJev-verdict-2.0](https://github.com/Heman10x-NGU/openJev-verdict-2.0) PR #1). `notes.md` §71 | +| Coverage 1.00, wrong `__none__` | Format-mass ≠ correctness; small-n theater | Measure `__none__` gold and shuffle/mix; do not threshold coverage ([chakuho](https://github.com/taku-me/chakuho) 8B 3/30 vs 27B 29/30). `notes.md` §72 | +| English checkpoint on Khmer/Hebrew | Confident-wrong OOD; p never drops | Route by **script before** the forward pass; do not wait for gating ([laya-multilingual](https://huggingface.co/convaiinnovations/laya-multilingual)). `notes.md` §72 | +| Commit "0.4, so pass" | Middle band is not a verdict | Report `"review"`; Nouls decide, Choice headlines ([commitjev](https://github.com/yodablocks/commitjev)). `notes.md` §72 | +| Plugin named Jev, key is Agnes | Branding ≠ backend | Read the client ([hermes-plugin-jev](https://github.com/Mrmimee/hermes-plugin-jev) is chat-completions). `notes.md` §72 | +| Quoted F1 without eval/README | `/benchmark` is tracked JSON; n=7 train-on-test | Read eval/README first; ~0.03 is a coin flip ([classifier-dev](https://github.com/mrmps/classifier-dev)). `notes.md` §73 | +| Docs say 0.800, serving 0.546 | Silent fallback is a lie about the instrument | Mark `FALLBACK`; rh-guard owns the gate ([classifier-dev](https://github.com/mrmps/classifier-dev) granite *theirs*). `notes.md` §73 | +| Escalate every multi-label on smart | Re-judge made it worse (23 s) | Smart is single-label <0.7 only; 0.7 is *theirs*. `notes.md` §73 | +| Paraphrased "quote" from a paper | Generator invented the excerpt | Point at line ids; copy verbatim; *Not found* is an answer ([choxos/jev-reviewer](https://github.com/choxos/jev-reviewer), ≠ egma-ai). `notes.md` §74 | +| One pass on a table with two Age rows | Relative Choice is not an absolute check | Two-pass: which-line Choice, then "does this line itself answer?" Noul. `notes.md` §74 | +| Unchecked extraction entered the review | Human tick skipped as chrome | Checked answers never overwritten; tick is the product. `notes.md` §74 | +| LocalJev JSON p used as a Noul | Self-reported vector ≠ logit read | Calibrate on *your* labels; wire-compat ≠ logit-equiv ([githubnext/localjev](https://github.com/githubnext/localjev), ≠ kunchenguid/local-jev). `notes.md` §75 | +| "localjev" without the owner | Namesake collision | Always **githubnext/localjev** (Bun Chat Completions) vs **kunchenguid/local-jev** (ONNX ModernBERT). `notes.md` §75 | +| 77-option Choice at default Laya head budget | ~3–4 tokens/label; labels collide | Hierarchical Choice, or a head that owns 255 options (Jev). Quote the token-budget fact; do not copy `head_max_len` ([NandhaKishorM/laya](https://github.com/NandhaKishorM/laya)). `notes.md` §76 | +| Auto-act because Laya conf ≥ 0.85 | Recipe ≠ Harbor cal; gating misses script OOD | Route by script first; fit T; pick τ on *your* labels. 0.85 is *theirs*. `notes.md` §76 | +| Treat 0.766 / 0.081 as zero-shot / raw ECE | Fine-tune on that split; post-T | Base ckpts below majority. Name the temperature. Jev rows unpublished-here. `notes.md` §76 | +| Rank openjevs from the census tweet | A list is not a bake-off | Use the scored sibling §78; still ≠ v1.1. Watch [jev-models](https://benchmarkheaven.com/jev-models). `notes.md` §77, §78 | +| Collapse GLiNER2 / routers into NAR clones because they are on the list | Class-boundary | Locate/categorize and route are placements, not replicas. Needle 3 already not Jev-class. `notes.md` §77 | +| Treat missing Laya/localjev/kev as out of class | Census lag | Incomplete ≠ our watch wrong. Completeness is a board watch item. `notes.md` §77 | +| Quote 15 likes as quality | Engagement is ephemeral | SIGNAL ~417/9; this pass 564/15. Do not copy Stripe. `notes.md` §77 | +| Mix v1.1 87.6 with v1.2 75.3 | Different tiers and scoring | Cal now ON the composite. Hard 220 new. `notes.md` §67, §78 | +| Treat Luna I=97 as rank #1 | Weak Cost axis (28.2) | Geo-mean product; rank #7 *theirs*. `notes.md` §78 | +| Ignore ×2 latency / est. costs | Assumption, not measurement | Harbor honesty; ranks are configuration-specific. `notes.md` §78 | +| Treat Qwen3.8 27B as Archer | Official Qwen / Chutes TEE | Partial; Cost 0 from price. Archer still Watch. `notes.md` §78 | +| Read Laya absence as quality | Gap, not a named exclusion | Absent from table **and** exclusion list. `notes.md` §78 | +| Reverse A/B on a small yes/no rebuild and quote one number | Option-order 72%→21% | Rank with author's order; keep both runs. Cousin of paraphrase brittleness. `notes.md` §78 | +| Re-card localjev / classifier.dev / Laya / choxos / census because they reappear on the hourly | Already folded | Apply the five as a recipe; skip thin noise. `notes.md` §79 | +| Fail CI / stamp quality from a Noul | Soft sensor as a hard seal | Attend or escalate; exact envelope proves the irreversible act. Qualify [totally-tim/jev-gate](https://github.com/totally-tim/jev-gate) ≠ jev-gateway / MongLong0214/jev-gate. `notes.md` §79 | +| Stall the reflex waiting for S2 / let S2 fly | Planner as executor | S1 keeps the stick; S2 is one-use advice. [khordoo/jev-reflex-autonomy-lab](https://github.com/khordoo/jev-reflex-autonomy-lab). `notes.md` §80 | +| Treat S2 arrival as consumed guidance | Telemetry conflates bar with decision | Purple confidence = used; purple S2 bar = arrived; red = fail. `notes.md` §80 | +| Call the lab's Local controller "localjev" | Namesake collision | Rule-based built-in **≠** githubnext/localjev **≠** kunchenguid/local-jev. `notes.md` §80 | +| Hard-act at the 20% starting gate / treat seed as replay | Soft slider as interlock; geometry as DST | 20% *theirs* still soft; schema-safe ≠ correct. Seed repeats layout, not timing. `notes.md` §80 | +| Send pixels or planner prose into the reflex | Omni / stale bearings | No graphical input; code never labels safest; physics owns collisions. Skip Archer. `notes.md` §80 | +| Assume confidence = selected probability | SDK field smuggled as the app contract | Application contracts ≠ TypeSafe methods. `notes.md` §80 | +| Overlapping CU actions / one 255-way soup | Confidence collapse; noise in the kind | Exclusive set; split kind/item/site ([typesafe-computer-use](https://github.com/awlevin/typesafe-computer-use)). `notes.md` §81 | +| Ship pixels to Jev for the click | Omni CU | OCR+AX text-state; answer-reader capture ≠ the Choice. Skip Archer. `notes.md` §81 | +| Quote 155× as a Harbor score / 0.4 as τ | One screenshot; product copy | Re-measure. schema-safe ≠ correct. `notes.md` §81 | +| Ship audio to Jev / treat 27/27 as Harbor | Omni voice; fixtures as a board | Transcript text-state; integration on captured pages *theirs*. [jev-voice-browser](https://github.com/moritzkremb/jev-voice-browser). `notes.md` §82 | +| Truncate free-text on a partial / spoken confirm as auth | Wait-policy collapse; soft Noul as interlock | Closed-set may fire; search/type wait. Confirm is convenience. `notes.md` §82 | +| Call a second model for "two" / collapse into jev-voice-control | Extra generation; namesake | Numbered overlay is exact. **≠** chris-wozniczek **≠** nikolas-j **≠** OCR §81. `notes.md` §82 | +| Let the model skip the wrap / silent ASK | Advisory sidecar; HITL skipped | Wrap *is* execution; ASK throws. [AgentGhost](https://github.com/reddpy/AgentGhost). `notes.md` §83 | +| Treat AUTO_APPROVE as auth / wrap hosted tools | Demo hatch; out-of-reach actuators | Provider tools stay unwrapped. rh-guard owns the gate. `notes.md` §83 | +| Collapse AgentGhost into actiongate / toolgate / jev-use / namesakes | Slogan mix; fail polarity | Wrap ≠ evidence-only; fail-closed ≠ jev-use fail-open. **≠** jwen5419807 **≠** vventirozos. `notes.md` §83 | +| Paste JP atlas ★ as a bake-off | Research-time stars as scores | Genre list, not verified evals. [@studio_yebisu](https://x.com/studio_yebisu/status/2101065176069886152). **≠** §77 **≠** §78. `notes.md` §84 | +| Quote 200× / 400× as Harbor / “cannot hallucinate” | Marketing multiples; schema as correctness | TypeSafe ceiling *theirs*. schema-safe ≠ correct. [@akshay_pachaar](https://x.com/akshay_pachaar/status/2101037514945597645). `notes.md` §85 | +| Copy the explainer’s Python / collapse into a wrap how-to | Recipe dump; product mix | Independent pedagogy. **≠** official docs **≠** Flavio **≠** AgentGhost §83. `notes.md` §85 | +| Treat topical cosine as “customer is asking” | Embedding as proposition | Contrast-set: all six about refund; only asking pass. [jev-semgrep](https://github.com/uehaj/jev-semgrep). `notes.md` §86 | +| Multiply parallel meaning Nouls / negative-query tricks | Independence; set-diff theater | Threshold each Noul, boolean-compose bits in code. ≠ jev-combinators metaphor. `notes.md` §86 | +| Call jev-semgrep Semgrep.dev / a merge gate | Namesake; soundness theater | **≠** [semgrep.dev](https://semgrep.dev). Ranking fail-open; not a gate. `notes.md` §86 | +| Paste 0.94/0.98 or ★42/51 as Harbor | LLM-as-judge / ephemeral stars | 10 cases × 51-line corpus *theirs*. Stars research-time. `notes.md` §86 | | Each answer right, decision wrong | Policy wrong | Change weights/thresholds in code, leave questions alone | ## Revision discipline diff --git a/.agents/skills/augustus/references/toolbox-mapping.md b/.agents/skills/augustus/references/toolbox-mapping.md index f998655..c64ab14 100644 --- a/.agents/skills/augustus/references/toolbox-mapping.md +++ b/.agents/skills/augustus/references/toolbox-mapping.md @@ -69,15 +69,60 @@ component; keep the rest of the method in code. | Probabilistic method: priors | Choice distribution as P(s,a) policy prior (MCTS/PUCT); Choice confidences as calibrated gating | **Empirical recipe** (jev-mcts, calibrated vs exact truth) | | Search: value function | Score rubric as leaf value V(s) — only where a simulator validates outcomes; speculative depth hard-capped at 2 | **Empirical recipe** (jev-mcts fidelity split) | | Measurement theory: probe vs estimate | Only post-execution probes concede milestones; model estimates never do — "estimation wearing a measurement costume" is the rejection template | **Empirical recipe** (jev-mcts, pi-warden done-check) | -| Experimental design: perturbation | Behavioral tests as the stats layer: candidate removal, option-order shuffle, distractor injection, boundary cases | **Contract-level** (validation.md) | +| Experimental design: perturbation | Behavioral tests as the stats layer: candidate removal, option-order shuffle, letter-shuffle on screenshot Choice, distractor injection, boundary cases | **Contract-level** (validation.md); letter-shuffle receipt: blackwood-rlcd 0.133 vs Jev 1.13 0.587 on 300 web steps (`notes.md` §46) | +| Experimental design: frozen protocol bake-off | Decision-model vs constrained LLMs vs deterministic baselines; accuracy + ECE + latency + cost + honesty; raw logs; recompute | **Empirical as Harbor/jevals practice** (DMB v2; jevals-data CC-BY-4.0 boards + JSONL; `notes.md` §49). Do not merge Banking77 across protocols | +| Experimental design: already-folded hourly | Named HIGHs already on the branch → extract how-to-apply; do not re-card | **Empirical as watch accounting** (hourly 0842: wire≠logit, FALLBACK, packaging honesty, pointer-not-generator, census≠score; skip thin noise; `notes.md` §79). Hard-gate Noul as PR/quality = soundness theater | +| Experimental design: Local vs Live reflex A/B | Same physics/seed; toggle only the decision backend | **Empirical as README architecture** (khordoo/jev-reflex-autonomy-lab; rule-based vs hosted `jev-latest`; 20% still soft; **≠** githubnext/localjev; not a scored bake-off; `notes.md` §80) | +| Experimental design: pre-registered AMBIGUOUS eval | Kill/go printed; cascade margin sensitivity; AUROC ≠ ECE; serving-path ≠ model-speed; same-day errata | **Empirical as Harbor/jevals practice (honest negative)** (jev-baselines-eval; cascade sign-flip; confidence=1.0 theater; encoder-with-labels; `notes.md` §55) | +| Experimental design: extractable-from-state axis | Same question with vs without a supporting passage; citation paraphrase vs reversed-meaning | **Empirical as a boundary map** (jev-capability-atlas history suite N=3; `notes.md` §49). Qualitative, not a knowledge-breadth estimate | +| Experimental design: combinatorial negative | Cell-wise Choice assembly of a grid vs extractive keep/drop | **Empirical as a negative** (ARC-AGI Direct Jev 4/400; `notes.md` §49) | +| Experimental design: collab arms | `llm_autonomous` vs `scripted_plus_jev` vs `llm_plus_jev`; Wilson + McNemar; the decision model is **not** a peer arm | **Empirical as a harness shape** (jev-testbench; bake into jevals/Harbor, `notes.md` §48) | +| Experimental design: Harbor on/off routing | Same coding-agent task with routing on vs off; hidden verifier; cheaper unsolved is not a saving | **Empirical as a *shape* and one-run signal** (jev-gateway-bench; `notes.md` §51). Pair CI merge-gate with Harbor + rh-guard | | Discrete math: width vs depth | Fan out in width (parallel ≈ free), pay depth linearly; two-stage only when next options depend on an earlier answer | **Empirical recipe** (fan-out: 12.2× cheaper, 10× faster) | -| Psychology: Kahneman | System 2 generates/proposes (LLM), System 1 discriminates (Jev); never the reverse | **Empirical recipe** (mcts-agent role split; 2026-09-18 mixed-architecture discourse) | -| IR / cascades | Cheap relevance / irrelevance before an expensive ranker or generator | **Empirical recipe** (RAG cookbook; jevprune; git-jev-stage; LlamaIndex Jev rerank) | -| Spec / lint | Project-defined semantic rules as predicates over a diff | **Empirical recipe** (jev-pref contract; JevLint file-level Noul; pi-warden; snifftest unsure-band) | -| Formal methods / DST / safety | Judgment triages counterexamples, failing seeds, and named-rule conformance; proof/MC/DST stay with their tools. Alloy finder ≠ Apalache BMC ≠ Quint run. DST trio: Antithesis hypervisor / Resonate HQ Lean+oracle+SDK (durable async) / PufferLib env+seed. Noul is a sensor, not a discharged PO. Semi-formal diagrams are vocabularies, not enforcers | **Hypothesis as product**, **Contract** as ownership (matching `mappings.md` §8 and `methods-catalog.md`; worked shape pi-warden — `formal-methods.md`, `formal-semi-formal.md`) | -| Decision analysis: VOI | Gather as an enumerated act; pay iff expected decision-loss drop > cost | **Hypothesis** as calculator (`mappings.md` §6; `mental-models.md`) | -| Signal detection | Noul as evidence variable; criterion from costs and base rate; ROC/PR on your labels | **Hypothesis** for non-SWE plots (`mappings.md` §7) | -| Safety engineering: STPA | Sensor ≠ constraint; table of unsafe control actions if the sensor lies | **Contract** as ownership (`mappings.md` §8; Leveson) | +| Psychology: Kahneman | System 2 generates/proposes (LLM), System 1 discriminates (Jev); never the reverse. S1 keeps control; optional S2 is one-use advice and does not fly. **Delta:** escalate-under-threshold **without stalling**; telemetry marks **consumed** advice (purple confidence ≠ purple S2 bar); Local controller (rule-based) vs Live API as Harbor-adjacent A/B of reflex backends (**≠** githubnext/localjev). Productized cascade: `conf ≥ τ` → S1 decides else S2 writes; routing fails open / safety fails closed; **routing accuracy unmeasured**; keyword fallback ≠ S1. **Harbor-shaped cousin:** decide→policy→LLM leftover with three compare arms on labelled emails (jav-email-cascade; Noul 0.5 never rounded; mock gen-json flat-confidence is *their mock*). S1 specialists + S2 coordinator is the same split (reification-labs/foreman is description-only Phoenix scaffold this pass — do not invent an Elixir API). Indexer: S1 GLiNER extract on the bulk, escalate LLM on the tail (10–50× unfilled). Healthcare: S1 remainder after NEWS2/code, S2 blinded review (explore-typesafe-ai; synthetic; not clinically validated) | **Empirical recipe** (mcts-agent role split; 2026-09-18 mixed-architecture discourse; jev-reflex-autonomy-lab as a control-loop shape, `notes.md` §46, **delta §80**; [dual-process-ai](https://github.com/taro1985/dual-process-ai) as a business/life cascade, `notes.md` §49; s1-graphify-indexer, `notes.md` §51; explore-typesafe-ai, `notes.md` §55; jav-email-cascade, `notes.md` §60). **Route ≠ memory:** a cheap intent gate skips memory/tool *tours* on easy routes; memory still writes; complex still searches (jev-hermes, `notes.md` §48) | +| IR / cascades | Cheap relevance / irrelevance before an expensive ranker or generator | **Empirical recipe** (RAG cookbook; jevprune; git-jev-stage; LlamaIndex Jev rerank; jev-pruner stdout after size/format envelope, `notes.md` §53). **Empirical as architecture** (decision-native-rag-skills retrieve-wide→decide→evidence-set; Hypothesis as a measured win, `notes.md` §55). **Empirical as README** (jev-sift classify-first MCP; mocks ≠ accuracy; errors/truncation ≠ irrelevant, `notes.md` §56). **Empirical as stripped-repo card** (jevgrep 79% top-5 vs BM25 40% / grep 20%; keyword still wins exact strings, `notes.md` §58). **Empirical as one-run** (Jev-RAG ≥70% cost / 72% latency vs Spark *rerank*; full-context Spark still faster, `notes.md` §58). **Empirical as line meaning-grep** (jev-semgrep AND/OR/NOT after threshold; proposition ≠ embedding; contrast-set; Semgrep.dev; not a gate; 0.94/0.98 *theirs*, `notes.md` §61, §86). **Empirical as evidence packets** (jevex 1/8→6/8 n=8; packet HitFile diagnostic, `notes.md` §61; **n=16** 160s→69s, `notes.md` §72) | +| IR / extractive keep-drop | Number candidates in code; model selects; copy verbatim; `redecide` thresholds on the log with no new calls. Compaction: same pointer job on tool results (Jev Noul/Score *or* GLiNER encoder). Computer-use: same pointer job on observed a11y/DOM controls (Jev *or* GLiNER2 *or* Cua-S1 option-attention *or* Stagehand harness pick *or* closed-vote JevOnly / host-owned waymode *or* ego-jev indexed viewport *or* typesafe-computer-use OCR+AX desktop hosted Jev *or* jev-voice-browser ASR+Playwright). Stdout: same pointer job on Bash chunks after a hard envelope (Jev Noul). Showcase: score-among-observed (ads, on-screen posts). **Evidence-synthesis two-pass:** relative Choice then absolute Noul; *Not found* is an answer; human tick never overwritten (**choxos/jev-reviewer**, ≠ egma-ai) | **Empirical recipe** (testimonial-miner; jev-reviewer pointer-not-generator, `notes.md` §48; gliner25-compaction char-offset + fail-closed keep_full, `notes.md` §50; gliner2-ultrafast observe→score→act, `notes.md` §52; jev-pruner, `notes.md` §53; cua-s1 specialist form, source-only, `notes.md` §54). **Empirical as PR body** Stagehand #2955 pick-and-copy (37/75 no-LLM ~0.5s vs 4.37s *theirs*; pick ≠ replacement, `notes.md` §57). **Empirical as README architecture** JevOnly no planner / waymode host-owned (`notes.md` §61). **Empirical as n=3 medians** ego-jev hot-click (~2× vs per-step LLM; not a bench; `notes.md` §65). **Empirical as one-session bench** jev-compactor 64.5%/366ms (`notes.md` §65); later product-arm **73%**/350ms/4 of 4 (`notes.md` §68). **Empirical as 500-trajectory bench** dizk/jev-lens 79% fewer tokens (`notes.md` §68). **Empirical as README** pi-om keep/kind verbatim (`notes.md` §68). Atlas class pattern: Your Signal / Near Here (`notes.md` §56). **Empirical as README delta** choxos/jev-reviewer **12★** two-pass / 18-q **4.6 s / $0.0101** *theirs*; human check is the product (`notes.md` §74). **Empirical as README** typesafe-computer-use OCR+AX macOS (**427★**; $0.0002 / 155× *theirs* one screenshot; exclusive actions; split kind/item/site; `notes.md` §81). **Empirical as README** jev-voice-browser ASR+Playwright (**103★**; 27/27 fixtures *theirs*; `notes.md` §82) | +| Spec / lint | Project-defined semantic rules as predicates over a diff; linter owns hard rules. AST remainder: Tree-sitter units, then typed questions; do not execute scanned code. Plain-English PR check: one condition + min-confidence; fail-closed on error. Skills→oxlint: AST/precheck prove, guidance whole-file in state, remainder judged — not a hard gate. **Sentence-as-rule:** ast-grep subjects × Jev `ask:` (mizchi/jevlint; 13/15 1.00/1.00 *theirs*; fail-open no-verdict; ≠ huntedman/JevLint) | **Empirical recipe** (jev-pref contract; Abide productized path — replay 93 sessions, edit precision ~26% / turn ~73% before tune, `notes.md` §47; JevLint file-level Noul; pi-warden; snifftest unsure-band; jevscan AST∩semantic, `tenbin` owns the lint skill, `notes.md` §48; if-ai, `notes.md` §51). **Empirical as Phoenix experiment** (jev-oxlint fixtures agree with the human answer key; routing sharp; coarse hint not, `notes.md` §58). **Empirical as 13/15 corpus** (mizchi/jevlint, `notes.md` §70). jev-marshal is Watch / empty this pass | +| Formal methods / DST / safety | Judgment triages counterexamples, failing seeds, and named-rule conformance; proof/MC/DST stay with their tools. Alloy finder ≠ Apalache BMC ≠ Quint run. DST trio: Antithesis hypervisor / Resonate HQ Lean+oracle+SDK (durable async) / PufferLib env+seed. Noul is a sensor, not a discharged PO. Semi-formal diagrams are vocabularies, not enforcers. Eval integrity: check the instrument, not just the score. Effect contracts, not surface tokens. **TLA+ compose:** protocol in TLC, oracle is Jev, escalate is the valve (never confidently wrong). **Coverage ledger:** exception queue visible; mint ≠ product brain | **Hypothesis as product**, **Contract** as ownership (matching `mappings.md` §8 and `methods-catalog.md`; worked shape pi-warden — `formal-methods.md`, `formal-semi-formal.md`). **Empirical as FINDINGS ledger** (dinostomp; 99 of 189 against itself; `dinostomp jev` if-statement hygiene; `notes.md` §62). **Empirical as certification** (construct-auto-classifier; privilege ≠ verdict; landed-script / headless≠auto-approve; `notes.md` §63, §68). **Empirical as TLC + chaos table** (jev-labs; 1,080 golden 0 wrong *theirs*; `notes.md` §67). **Empirical as README** (seal; `notes.md` §67) | +| Decision analysis: VOI | Gather as an enumerated act; pay iff expected decision-loss drop > cost. Wake/resume: skip the LLM turn only if the judge answers and p is low (Horvitz); user-message / skip-limit already answer without a model. Selective memory: score a verbatim ledger; dump on failure; never judge the rules. Classify-first: pay for a full agent open iff relevance might change the act. **Specialist vs few-shot:** pay for a local head iff downstream *reads* p (Domain-jev-maker). **Training-data VOI:** pay for teacher/human labels iff confidence says they change the outcome (jev-triage); do not distill Jev as teacher. **Decision-model latency cost:** pay for sync Jev on a router iff quality gains beat hundreds of ms tail (slo-router negative). **Human-review VOI:** pay for a look iff the filter is unsure (jev-lens; never green unless sure). **Skill-library VOI:** pay to load a skill iff it changes the next step (skillranker; abstention first-class). **Same-intent cache:** pay for the LLM iff intent is new (jevcache 0 FP/100). **Human-feed VOI:** pay for a click iff worth-your-attention (ThinkyMiner/Winnow; ≠ kevinpita/winnow). **Second-call VOI:** pay for the narrating main-model turn iff generation is still required (hermes-jev-router). **Hunk-review VOI:** pay for generative review of a hunk iff Jev says it is worth looking at (prune-review; safety escarpment always keeps concurrency/auth/a11y/startup; ~20% cost target). **Intent-search VOI:** pay to judge a place the diff did not touch (jev-intent-review; UNKNOWN cheaper than false VERIFIED). **No-text-step VOI:** pay the LLM only when writing is the job (jev-use 186 vs 2,672 ms). **Batch ranking VOI:** one request per ten (jevfeed). **Empty compact-proxy skip:** IPECTER slogan only. **Split-question CU VOI:** kind/item/site/offscreen in one request; exclusive actions; perception rebuild is the expensive gather (typesafe-computer-use). **Partial-speech VOI:** complete Noul; closed-set may fire; free-text waits (jev-voice-browser) | **Hypothesis** as calculator (`mappings.md` §6; `mental-models.md`). wakegate 21/21 is smoke (`notes.md` §51). carryforward 9×3 is a hint (`notes.md` §55). jev-sift mocks ≠ accuracy (`notes.md` §56). Domain-jev-maker is Empirical as their RESULTS.md (`notes.md` §60). jev-triage is Empirical as README architecture (`notes.md` §61). **Empirical as a *negative*** (slo-router p95 77.93→490.38 same routes; `notes.md` §63). **Empirical as README** (jev-lens never blocks; `notes.md` §63). **Empirical as README** (skillranker 52★; hook fail-open; `notes.md` §66). **Empirical as eval finding** (carryforward 0/4; tools≠use; `notes.md` §68). **Empirical as 500-trajectory** (dizk/jev-lens compress-before-first-send; `notes.md` §68). **Empirical as n=100** (jevcache 0 FP; `notes.md` §69). **Empirical as unreviewed goldens** (ThinkyMiner/Winnow 80%/90%; `notes.md` §69). **Empirical as 22-run cost** (prune-review 1.18% with 305% outlier *theirs*; `notes.md` §70). **Empirical as CLI** (jev-intent-review; `notes.md` §70). **Empirical as 95-call card** (jev-use; `notes.md` §71). **Empirical as README** (jevfeed; `notes.md` §71). **Empirical as n=16** (jevex; `notes.md` §72). **Empirical as 13 labelled** (commitjev; `notes.md` §72). **Empirical as latency table** (pi-jev-compact; `notes.md` §72). **Empty skip** (IPECTER runway; `notes.md` §72). **Empirical as README** (classifier-dev escalate-under-threshold; `notes.md` §73). **Empirical as README** (typesafe-computer-use 155× *theirs* one screenshot, not a taskset; `notes.md` §81). **Empirical as README** (jev-voice-browser 27/27 fixtures *theirs*; `notes.md` §82) | +| Signal detection | Noul as evidence variable; criterion from costs and base rate; ROC/PR on your labels. Operator owns the criterion; a plugin must not self-tune the safety bar. Exactness raises a quality floor — it must not override capability. Privilege ≠ verdict. Ranking ≠ calibration: never hard-threshold raw p as a frequency. **Sureness:** inspect the vector (entropy/margin), not only max_prob. **Conflict ≠ ignorance:** name those states as Choice options. **BBQ:** stereotype/uncertainty/cost as one card; not a general bias cert. **Dual-channel ECE:** correctness-head ≠ distribution ECE; like-for-like before a 15× badge. **Escalate-under-threshold:** smart re-asks single-label <0.7; multi-label ignores (classifier-dev; 0.7 is *theirs*) | **Hypothesis** for non-SWE plots (`mappings.md` §7). **Empirical as measured OMP suppression** (omp-greenlight default 40.9% / 0 of 94 on labelled corpus; live traffic unlabelled; `notes.md` §62). **Empirical as live analysis** (slo-router exactness floor; 3/8 label disagreements did not change routes; `notes.md` §63). **Empirical as certification** (construct-auto-classifier; `sudo status` can be safe; `notes.md` §63). **Empirical as 8,000-judgment audit** (does-jev-confidence; AUC ~0.91; stated ~75% vs human ~10%; `notes.md` §64). **Empirical as 60-q library** (how-sure-is-jev; Choice confidence = max_prob; `notes.md` §67). **Empirical as HA measurements** (HA-Jev; not for locks; `notes.md` §68). **Empirical as owner-run smoke** (jev-preflight fail-open attention; `notes.md` §68). **Empirical as NCML field note** (typed-evaluation-collapse; `notes.md` §69). **Empirical as full BBQ** (jev-bbq-experiment 97.28%/0.04/0.34/$0.3429 *theirs*; `notes.md` §70). **Hypothesis until independent run** (openJev-verdict-2.0 + PR #1; `notes.md` §71). **Empirical as 13 labelled** (commitjev middle band; `notes.md` §72). **Empirical as MASSIVE** (laya-multilingual confident-wrong OOD; `notes.md` §72). **Empirical as README Router** (NandhaKishorM/laya; Khmer 0.000@0.952; 0.85 still soft; post-T ≠ raw ECE; `notes.md` §76). **Empirical as GUI 336** (chakuho coverage ≠ correctness; `notes.md` §72). **Empirical as Hub eval** (schema-scorer peaked ranking; `notes.md` §72) | +| Safety engineering: STPA | Sensor ≠ constraint; table of unsafe control actions if the sensor lies. Host deny stays above Jev prompt-suppression. Judgment ≠ permission. Contracts on effects, not tokens. Attention filter ≠ permission gate. **Jev supplies evidence, code owns authority**. **Wrap-as-execution:** the wrap *is* the actuator path (AgentGhost; ASK throws; fail-closed; rh-guard owns the gate cousin). **Turnstile:** policy first, Jev remainder, replay. **SEAL:** coverage ledger; no silent advance. **Never confidently wrong:** escalate instead of hard-gate. **Typed baton:** escalate/continue/abort; gate never grants. **Persist constraints:** conversational policy survives compaction; Jev never writes it (pi-heed). **Pi control plane:** named sensors; GUI never force-click (pi-jev-control). **jev-use gate:** fail-open deny/ask. **Silent FALLBACK:** digest names the model that answered (classifier-dev granite F1 0.546 vs advertised ~0.800 *theirs*; rh-guard owns the gate) | **Contract** as ownership (`mappings.md` §8; Leveson). **Empirical as architecture** (interlock: Jev SENSOR, policy.py constraint, secrets never in agent; type-safe ≠ correct; `notes.md` §59). **Empirical as measured suppression** (omp-greenlight: not a sandbox; operator owns bar; `notes.md` §62). **Hypothesis / outline** (skill-broker: Jev never grants access; sibling to turnstile/skillranker; `notes.md` §62, §67). **Empirical as certification** (construct-auto-classifier: effect contracts; fail-closed; `notes.md` §63). **Empirical as README** (jev-lens: never blocks the agent; `notes.md` §63). **Empirical as slogan** (actiongate-jev: positive p never overrides a deterministic failure; `notes.md` §64). **Empirical as README** (AgentGhost wrap-as-execution; ASK throws; `notes.md` §83). **Empirical as README** (turnstile: evidence ≠ authority + replay; `notes.md` §66). **Empirical as TLC + chaos** (jev-labs; `notes.md` §67). **Empirical as README** (seal; `notes.md` §67). **Empirical as README** (jev-handoff: gate never grants; fail-open; `notes.md` §69). **Empirical as 79-session bench** (pi-heed 98.5%/0 false block *theirs*; `notes.md` §70). **Empirical as README** (pi-jev-control / jev-use; `notes.md` §71). **Empirical as README** (mailordinal inbox policy; `notes.md` §72). **Identity lock** (hermes-plugin-jev is Agnes not TypeSafe; `notes.md` §72) | +| Experimental design: native-probability arena | Analytic worlds; Brier/ECE/reliability; fan-out as measurement economics; teeth stubs | **Empirical as their live card** (jev-arena 145 noul Brier 0.0059 / ECE 0.0620; sonar/vickrey/bracket suite; `notes.md` §59) | +| Experimental design: evidence-gated question packs | Pack.yaml + golden cases + evidence.md; pin version; no numbers → not verified. Runner: record/replay CI offline (accuracy/ECE/Brier/cost/latency; exit 0/1/2; McNemar) | **Empirical as registry + runner** (jev-packs nine packs *theirs*; **jevassert LANDED**; 2,990-case matrix Jev/Sonnet 5 accuracy tie, Jev better calibrated 7/9, ~250× cheaper; sms-spam 0.953/0.040; `notes.md` §64, §70). **Hunch:** measurement owns endorsement | +| Experimental design: failure-finding arena | BYOK pairwise harness; find wrong questions, do not crown a winner | **Empirical as runnable harness** (chenmingtang830/jevarena ≠ meetr1912/jev-arena; not measured findings; `notes.md` §70) | +| Experimental design: BBQ stereotype/uncertainty/cost | Full 58,492; amb vs inf; bias + $ + latency; not a general bias cert | **Empirical as full-set card** (jev-bbq-experiment 97.28%/0.04/0.34/$0.3429 *theirs*; `notes.md` §70) | +| Experimental design: Harbor SGR-judge contract | Frozen Jev vs schema-guided LLM judges; invalid = FN; cost/latency; incomplete ≠ headline | **Empirical as contract, not a score** (jev-judge-bench; 21 offline tests; canaries ≠ quality; **no quality headline yet**; ≠ jevarena/jevbench; `notes.md` §71) | +| Experimental design: recipe samples vs benches | Handmade 16–36; authors declare not-a-bench; cost/latency still reportable | **Empirical as recipe atlas** (jev-cookbook 425/$0.015; browser 5/6 *theirs*; `notes.md` §71) | +| Experimental design: competing NAR claim-audit | Like-for-like ECE; n/CI; throughput ≠ latency; vendor-baseline rows named | **Hypothesis until independent run** (openJev-verdict-2.0 77.10%/0.0636/0.0144 *theirs* + PR #1; ≠ IamBusy/OpenJev; `notes.md` §71) | +| Experimental design: 1-token logprob vs hosted Jev | Coverage ≠ correctness; `__none__` gold; shuffle/mix; numeric-rule probe | **Empirical as 336-case GUI + mario** (chakuho; `notes.md` §72) | +| Experimental design: replica engine argmax-parity | Speedup at 100% argmax; family adapters; MPS-only honesty | **Empirical as README** (jevinf 2.57×/2.27×; `notes.md` §72) | +| Experimental design: files-to-read n=16 | Same cheap agent both arms; 90s-cap finish; keep older n=8 | **Empirical as author-run** (jevex; `notes.md` §72) | +| Experimental design: commit middle-band | Three-way criterion; regex first; small clean control named | **Empirical as 13 labelled** (commitjev; `notes.md` §72) | +| Experimental design: language OOD / confident-wrong | Script-route before p; MASSIVE + XNLI; T refit | **Empirical as Hub card** (laya-multilingual; `notes.md` §72) | +| Experimental design: schema-conditioned encoder | Grouped softmax; peaked p = ranking; Hub-only if GitHub 404 | **Empirical as Hub eval** (schema-scorer v2 0.841; `notes.md` §72) | +| Experimental design: public vs_jev + named caveats | Tracked JSON on the site; read eval/README; train-on-test n=7 honesty | **Empirical as eval/README** (classifier-dev F1 0.887 / 230–232 ms; AG News 87.7%; granite 0.546 vs advertised 0.800 *theirs*; `notes.md` §73) | +| Experimental design: prompted-JSON local bake-off | Same engine, five backbones; AG News/BoolQ/SST-5; two input lengths; caveats first | **Empirical as evaluation-results-2026-09-18** (githubnext/localjev 1,200 req; Qwen3.6 76.7% / Gemma 26B 75.0% / DiffusionGemma 74.2% short *theirs*; not logits; not calibrated; `notes.md` §75) | +| Experimental design: vs-Jev honesty + post-T vs raw ECE | Name third-party rows; 72 vs 77 labels; T fit; fine-tune ≠ zero-shot | **Empirical as README** (NandhaKishorM/laya; Jev unpublished-here; Banking77 0.425 vs 0.870; post-T 0.081 vs raw 0.213; base < majority; `notes.md` §76) | +| Experimental design: external openjev census ≠ scored bake-off | Named list + promised board; likes ephemeral; do not paste live ranks into the census card | **Empirical as tweet** (@airesearch12 status/2101259522933186879; watch [jev-models](https://benchmarkheaven.com/jev-models); **≠** jevbench v1.1; scored sibling §78; `notes.md` §77) | +| Experimental design: JP application genre atlas ≠ bake-off | Apps by hole; stars research-time; not verified evals; likes ephemeral | **Empirical as tweet** (@studio_yebisu status/2101065176069886152; typesafe-computer-use 203→427; jev-voice-browser 40→103; **≠** class census §77 **≠** v1.2; `notes.md` §84) | +| Experimental design: explainer multiples as ceiling | Quote 200×/400× as TypeSafe ceiling; schema-safe ≠ correct; not a Harbor score | **Empirical as article** (@akshay_pachaar status/2101037514945597645; 70–500 ms / $0.042/MTok *theirs*; **≠** official docs; `notes.md` §85) | +| Experimental design: Harbor honesty on a public class table | Calibration on/off the rank; cost/latency assumptions; silent fallback; partial runs; class column for GLiNER2/routers/Needle 3 | **Empirical as v1.2 board** (cal ON rank; self-host ×2 assumption; many costs est.; kinship §67 v1.1 cal off Main Score + classifier-dev FALLBACK; `notes.md` §78) | +| Experimental design: geometric-mean scored bake-off | I/C/S/K 25% each; weight views; option-order probe; instruction models in the same table | **Empirical as RESULTS-v1.2** (Jev 75.3 / SemIf 74.6 *theirs*; Luna I=96.8 #7; 72%→21%; Laya absent gap; Qwen3.8 27B ≠ Archer; **≠** v1.1 87.6; `notes.md` §78) | +| Experimental design: ranking vs calibration on human labels | AUC vs ECE/Brier vs annotator rates; Platt/isotonic held-out; wording as a factor | **Empirical as 8,000-judgment audit** (does-jev-confidence; stated ~75% vs human ~10%; ~96% ECE removed; jevcal ~100 rows; `notes.md` §64) | +| Experimental design: OOD calibration / sign by type | Unknowable policy label; ECE/floor; refit T per Choice/Score/Noul | **Empirical as 900-ticket table** (jev-ood-calibration; priority 44.7%/mean p 0.74/T 3.40; boolean T 0.66; `notes.md` §66) | +| Experimental design: thinking-budget bake-off | Frozen 100-task set; attach reasoning budget; intervals | **Empirical as exploratory release** (jev-frontier-100; Jev 77.0% vs 4B/2048 96.7%; not preregistered; `notes.md` §66) | +| Experimental design: scored Jev-class bake-off | Capability/Speed/Cost composite; calibration off the rank; native vs verbalized | **Empirical as v1.1 artifact** (jevbench; Main Score 0.6/0.2/0.2; Jev 1.13.0 87.6 *theirs*; unofficial; superseded for the live board by v1.2; `notes.md` §67) | +| Experimental design: pointer compact vs LLM summarize | Saved tokens vs invented paths vs early-fact recall vs latency/cost | **Empirical as product-arm table** (jev-compactor later **73%**/350ms/4 of 4 vs shipped summarizers; earlier vs-Sonnet 64.5%/366ms; `notes.md` §65, §68) | +| Experimental design: pre-send views vs post-send prune | Tokens before first send; cache cost; harm on later edits | **Empirical as 500-trajectory** (dizk/jev-lens 79% fewer; post-send +17% cost; `notes.md` §68) | +| Product decision: independent open-Jev class | Finite-choice+prob without TypeSafe; wire-compat ≠ replica. **Replica substrates** (design, not vendor-lock): Rust/WebGPU (grande), Clojure/Jolt Laya (laya-jolt byte parity), CPU SemIf (leesk212/JEV-CPU; Meanblock 404), ONNX ModernBERT (local-jev measured not equivalent), GLiNER2 spec (Eran-BA; spec ≠ jeff). **Prompted-JSON wire** (githubnext/localjev; **≠** kunchenguid/local-jev; wire-compat ≠ logit-equiv vs razorback16/openjev). **Competing NAR claims** are an audit object (openJev-verdict-2.0; ≠ IamBusy/OpenJev) | **Empirical as their docs** (openvons LM 0.916 vs 27B 0.875; JevPick 3.2–4.8×; `notes.md` §68). **Empirical as RESULTS.md** (IamBusy/OpenJev 45/60 `/v1/decide` ≠ drop-in; `notes.md` §69). **Empirical as latency table** (semif-serve 1164 vs 178 ms; runoff ≠ softmax; `notes.md` §69). **Empirical as JGLUE + isolation** (grande JNLI 0.614 ECE 0.088 T=2.81 / JCQA 0.853 / 270M 0.710/0.710 *theirs*; `notes.md` §70). **Empirical as byte parity** (laya-jolt; `notes.md` §70). **Empirical as 136-ckpt** (local-jev done 30%/shape 57%; `notes.md` §70). **Empirical as README + 1,200-req eval** (githubnext/localjev **261★**; Qwen3.6 76.7% / Gemma 26B 75.0% / DiffusionGemma 74.2% short *theirs*; not calibrated; `notes.md` §75). **Empirical as README packaging** (NandhaKishorM/laya **710★**; Router over Hub Laya; not a new species; `notes.md` §76). **Empirical as tweet census** (@airesearch12 ~18 named; GLiNER2+routers class-boundary; incomplete vs Laya/localjev/kev; **≠** jevbench v1.1; watch jev-models; scored sibling §78; `notes.md` §77). **Empirical as v1.2 board** (geo-mean I/C/S/K; Jev 75.3 / SemIf 74.6 *theirs*; instruction models in the table; OpenJev = razorback16 ≠ IamBusy; Laya absent gap; `notes.md` §78). **Hypothesis until independent run** (openJev-verdict-2.0 + PR #1; `notes.md` §71) | +| Experimental design: same-intent cache vs cosine | False-positive rate of admitting a cached completion | **Empirical as n=100** (jevcache Jev 0 FP vs Jaccard@0.35 fpr 0.48; `notes.md` §69) | +| Experimental design: zeroshot vs BERT-family | Acc/AUC vs clean `-c`; contamination DiD; label-equivalence | **Empirical as 7-set table** (jev-zeroshot-vs-bert +0.05–+0.13; ~230 / >2048 labels; `notes.md` §69) | +| Experimental design: conflict vs ignorance schema | Noul collapse vs named Choice escape vs binary lexical bias | **Empirical as NCML field note v0.3** (typed-evaluation-collapse; `notes.md` §69) | +| Experimental design: BM25 vs Jev skill routing | Roster size 50–500; hit / miss / false_load | **Empirical as harness, not a score** (pi-jev-skill-bench 43 gold; no live numbers this pass; `notes.md` §69) | +| Experimental design: ranking family on soft scores | Pairwise inversion / Score ordinality / two-decimal ties; request-shape as a factor | **Empirical as independent measurement** (jev-orderby-bench six gates; Score 0.143 weak link; recodelabs batch-40 fails ranking; calibration ≠ sortable; `notes.md` §60) | +| Product decision: wire-compat backend | Self-host the System One *wire* on an encoder when GPU economics beat hosted and the accuracy gap is acceptable | **Empirical as their RESULTS.md** (jeff GLiFormer ~6× L4 HTTP / ~24× A10G direct; AG News 75.5% vs 90.5%; CPU more expensive; not a Jev replica; `notes.md` §60) | +| Constrained-AR PCD vs calibrated decide | O(1) schema-valid speed is not a Noul | **Empirical as their n=50 table** (system-one-benchmark; Jev 84.0% / Brier 0.1096 vs local MLX PCD 52% / 0.3884; `notes.md` §61) | +| Optimizer vs control plane | Ax/DSPy climb LM knobs; typed deterministic plane around the program; DSPy drafts AFTER route+action | **Empirical as architecture** (jev-dspy-control-plane; offline stubs ≠ quality; `notes.md` §59) | | Bandits / RL | Value from observed rewards only — Jev provides none; rejected without an environment. PufferLib Ocean is a trainer contract, not a baseline | **Rejected** (standing boundary) | Invalid-but-tempting (record these so they don't get rediscovered): treating diff --git a/.agents/skills/augustus/references/validation.md b/.agents/skills/augustus/references/validation.md index 165d13f..d7efd99 100644 --- a/.agents/skills/augustus/references/validation.md +++ b/.agents/skills/augustus/references/validation.md @@ -4,7 +4,12 @@ 1. Exact computation, semantic judgment, or both? Exact parts stay in code. 2. Can the right answer be represented? (candidate present? level exists? - `other` option where coverage is open?) + `other` option where coverage is open?) Missing `other` on an open + coverage set forces a wrong Choice at confidence 1.00 — **confidence + gating cannot catch it**. Lint the *request* (broken state paths, + bundled judgments) before you trust the answer + ([wellposed](https://github.com/suraj-phanindra/wellposed) recipe; + `tenbin` owns the skill; `question-design.md`; `notes.md` §46). 3. Missing / contradictory / malicious / stale evidence — what happens? 4. Which constraints must code enforce regardless of model output? 5. What does each number mean — and which reading would be invalid? @@ -15,7 +20,9 @@ ## Behavioral tests (measure; Jev promises no invariances) Candidate removal (drop the winner — does probability spread sensibly?); -option-order shuffle; **irrelevant-option / IIA** (append an option that +option-order shuffle; **letter-shuffle on screenshot Choice** +(blackwood-rlcd card: flip **0.133** vs Jev 1.13 text-only **0.587** on +300 web steps — vendor receipt, not re-run; `notes.md` §46); **irrelevant-option / IIA** (append an option that should not move odds among the rest — Hume's reconstruction, `research/notes.md` §31, not a new invariance the API promises); public cousin for the read-the-letter graph: @@ -33,6 +40,15 @@ rejected two correct mates (`notes.md` §42). 149-row cousin: [`typesafe-jev-tools`](https://github.com/wotai-dev/typesafe-jev-tools) — Jev confidence monotonic vs Haiku invert in 0.80–0.95; do not copy the hook. +Open reconstruction cousin: [`jaredpalmer/kev`](https://github.com/jaredpalmer/kev) +— isolation packed vs separate max Δ 3.7e-6; secret-in-sibling p=0.03 +vs in-state 0.99; permute argmax flips 7.4%; IIA log-odds shift mean +0.13; boundary forgery held. Those tests mirror Archer probes; they do +not prove kev = Jev (`notes.md` §45). Hub fetch path this pass: +[`jaredpalmer/kev-0.5b`](https://huggingface.co/jaredpalmer/kev-0.5b) +(`--run` accepts Hub ids). Dedicated `none_of_the_above` eval (true +option present vs removed) is the training-side cousin of wellposed's +request hatch; **no published rates this pass** (`notes.md` §45 delta). ## Offline eval: selective binary decisions @@ -52,9 +68,12 @@ the offline Brier / reliability / cost evaluator. Pointer only the Harbor substrate, and the one composition table are **Eval & hill-climb** below — do not restate them here. A bake-off candidate on that same labeled-case surface, beside Laya, openjev-lm, -and TypeAR, is [Bespoke Nimble](https://github.com/bespokelabsai/nimble) +TypeAR, and [kev](https://github.com/jaredpalmer/kev), is +[Bespoke Nimble](https://github.com/bespokelabsai/nimble) — an open LoRA recipe, not a Jev distill; their 324-example holdout is -a named receipt, not a ranking (`research/notes.md` §35). Same +a named receipt, not a ranking (`research/notes.md` §35). kev is the +runnable Archer-reconstruction candidate on the same surface (public +gold, measured ID ECE, not a teacher-copy; `notes.md` §45). Same acceptance-test *surface*, different UI: [jeiel85/jevscope](https://github.com/jeiel85/jevscope) (local-first visual debugger + JSONL regression; policy buckets are JevScope-derived, @@ -105,7 +124,11 @@ Adoption pattern (AntonioCoppe/jev-harness): run the Jev judgment in parallel wi the live system and only **log what you would have done** (policy + gate applied) until behavioral evals over replayed fixtures pass; then flip to enforcement. Assert on the *action* (block/warn/pass), not on free text. This is the safe path for any -confidence gate added to an existing pipeline. +confidence gate added to an existing pipeline. Harbor/jevals-adjacent +practice, not a second eval product: LLM-as-judge is not the primary +System One score (`faq.md`). Recipes in that repo (alerts, RTB, sports-bet, +prediction-markets) are existence proofs of the same substrate across +business and life, not SWE-only (`notes.md` §44). Do not copy the client. ## Frontmatter (by agents and by Jev rankers) @@ -281,9 +304,15 @@ Rules: - **Room / omni products** (same ask): structural gates first; video-as-judge last. Same sandwich as allowlist-then-remainder (`mappings.md` §18). Perception-then-judgment is composition; - information dies at the interface, and a Noul is not over raw pixels - (`notes.md` §39). LLM-as-judge is not the primary score for a - calibrated System One. + information dies at the act. A shared multimodal decide head still + judges marked candidates, not an open click (`notes.md` §39, §46). + LLM-as-judge is not the primary score for a + calibrated System One. Shared bake-off exemplar: + [open-jev-laya-bench](https://huggingface.co/datasets/pngwn/open-jev-laya-bench) + (ECE/NLL/Brier). Harbor-style class bake-off: + [DMB](https://github.com/nibzard/decision-model-benchmark) (v2). + Public log feedstock: + [jevals-data](https://github.com/Jevals/jevals-data) (CC-BY-4.0). ### Composition @@ -294,21 +323,507 @@ Rules: | End-to-end product / agent loop | Harbor taskset | behavioral assertions, cost/perf bounds | | LM-program knobs only | DSPy/Ax (narrow) | never primary System One calibration score | | Reward-hack / eval gaming | [rh-guard](https://github.com/24601/rh-guard) | structural deny + System One sidecar | - -rh-guard is a reward-hack hook, a different surface from jevgate. One -row is enough. ECE above is wanted, not a Nimble result. +| Project soft-rule lint | [Abide](https://github.com/coldteadotai/abide) | Score per rule on the diff; bands; fail-open; replay + independent review | +| Collab / computer-use product loop | [jev-testbench](https://github.com/ufx7/jev-testbench); [solari-reflex](https://github.com/hitakshiA/solari-reflex); [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast); [cua-s1](https://github.com/trycua/cua/tree/main/libs/cua-s1); [Stagehand #2955](https://github.com/browserbase/stagehand/pull/2955) | Wilson/McNemar arms; independently checked task time; `DONE` ≠ success; Cua-S1 source-only (metric names, no checkpoint scores); Stagehand 37/75 no-LLM ~0.5s vs 4.37s *their* card; pick ≠ replacement; draft | +| Agent routing on vs off | [jev-gateway-bench](https://github.com/vinilana/jev-gateway-bench) | Hidden perft; cost/quality; one-run signal this pass | +| Command-output prune (needle/noise) | [jev-pruner](https://github.com/tamaratran/jev-pruner) | Manual `trimOutput` sweep (theirs); plugin eval cannot reach Jev (fail-safe original); Terminal-Bench paired pilot is integration, not a full bench | +| Pre-registered cascade vs nano/frontier/encoder | [jev-baselines-eval](https://github.com/ickma2311/jev-baselines-eval) | Both experiments **AMBIGUOUS**; cascade sign-flip at exact parity; confidence=1.0 theater; encoder 0.933/9ms with labels; serving-path ≠ model-speed; same-day errata ×3 | +| Healthcare S1+S2 (synthetic FHIR) | [explore-typesafe-ai](https://github.com/si618/explore-typesafe-ai) | Labels committed first; 60 requests / 403 judgments; **not clinically validated**; Claude wrote labels | +| Precision PDF (honest negative) | [databricks-jev-pdf-lab](https://github.com/laurentfabre/databricks-jev-pdf-lab) | No quality-equivalent Jev payoff; no OSS license selected | +| Meaning-search without embeddings | [jevgrep](https://github.com/Bentlybro/jevgrep) | 228-q stripped Flask/httpx/Django/AutoGPT: 79% top-5 vs BM25 40% / grep 20%; keyword still wins exact (BM25 top-10 96% vs 85%); not a Harbor taskset | +| RAG rerank vs generative rerank | [Jev-RAG](https://github.com/Max-sm-yc/Jev-RAG) | One-run ~30k tokens: ≥70% cost / 72% latency vs Spark *rerank*; full-context Spark still 10.60 s; costs include embeddings | +| Skills→oxlint remainder | [jev-oxlint](https://github.com/cephalization/jev-oxlint) | Phoenix: answer-key agree on every fixture; routing 0.80–0.94 vs <0.50; coarse hint not; experiment; not a hard gate | +| Native-probability calibration (analytic worlds) | [jev-arena](https://github.com/meetr1912/jev-arena) | Live *theirs* (`jev-1.13.0`, 145 noul, 2 req / 710 ms): Brier 0.0059, log loss 0.5393, ECE 0.0620; overconfident in low bins; always-0.5 Brier 0.0766. Oracle stub 0.0000. Fan-out economics. Offline default | +| Fan-out suite (heatmap / CDF / bracket) | [jev-sonar](https://github.com/meetr1912/jev-sonar); [jev-vickrey](https://github.com/meetr1912/jev-vickrey); [jev-bracket](https://github.com/meetr1912/jev-bracket) | Sonar: heatmap-as-policy; offline 75% win / Brier 0.1615; live 1-game Brier 0.1092. Vickrey: Jev never bids; live Brier 0.1391 / ECE 0.1321; oracle regret 0. Vickrey live second-price profit −163.4. Bracket: live Brier 0.2853 vs Elo 0.2322 (trailed Elo; honest) | +| Typed control plane vs DSPy / JSON Schema | [jev-dspy-control-plane](https://github.com/manikanda-kumar/jev-dspy-control-plane) | Same ontology/dataset/state/allow-list/metrics; intent/sub-intent, invalid/policy-violation, abstention/coverage, Brier/ECE, consistency, p50/p95. Offline heuristic ≠ quality | +| Tetris legal-set Choice vs Haiku | [jev-tetris-benchmark](https://github.com/planstack-ai/jev-tetris-benchmark) | Use-case demo; code enumerates ≤12 legal placements; **not a rigorous eval** | +| Domain specialist vs few-shot hosted | [Domain-jev-maker](https://github.com/help-er/Domain-jev-maker) | Independent CLINC gold (not Jev teacher). Matched-precision KL (2-decimal, zeros→0.0025): local 0.168 vs hosted 0.580 banking; r +0.933 vs +0.343. Few-shot hosted determinate McNemar n.s. (p=0.134 / 1.000). Train specialist when policy reads p | +| Cascade compare arms (native vs verbalized vs logprob) | [jav-email-cascade](https://github.com/skiingfalcon/jav-email-cascade) | 74 labelled emails; jev / gen-json / gen-logprob; shared Answer schema. Mock: gen-json confidence flat. Noul 0.5 never rounded. License null. **Not** a live Jev vs Haiku bake-off | +| ORDER BY ranking vs calibration | [jev-orderby-bench](https://github.com/yodablocks/jev-orderby-bench) | `jev-1.13.0` six gates pass. Boolean inversion 0.036; Score ordinal **0.143** vs 0.15; 53-way 0.99 tie; ECE 0.0453 / Brier 0.0524. recodelabs batch-40 inversion 0.171 **fail**. Calibration ≠ sortable | +| Class-backend economics (GLiFormer `/v1/systemone`) | [jeff](https://github.com/logan-markewich/jeff) | 1,600 items. L4 HTTP ~$2.6 vs jev ~$15.6 (~6×); A10G direct ~$0.65 (~24×); AG News 75.5% vs 90.5%; p50 151 vs 129 ms. CPU 6–20× *more* expensive. Encoder ≠ Jev replica. License null | +| Jev vs local MLX PCD vs AR JSON | [system-one-benchmark](https://github.com/mallahyari/system-one-benchmark) | LMSYS toxic-chat **n=50**. Jev-1.13.0 **84.0%** acc / Brier **0.1096** / p50 356.5 ms; PCD Qwen2.5-1.5B 52% / Brier 0.3884 / p50 227.2 ms / 1 pass O(1); AR 54% / ~30.8 passes / 98% schema errors. License null. Small n — *their* card, not a large Harbor taskset. PCD O(1) ≠ calibrated Noul | +| Evidence-packet explorer (SWE finish) | [jevex](https://github.com/jimmyhealer/jevex) (was jev-semantic-explorer) | Author-run. Claude Code 6.8→2.2 files. SWE-bench Verified n=8: **1/8 → 6/8** finish. **n=16 delta** *theirs*: 160s→**69s**, $8.74→**$3.13**, 16/16 both arms; 90s cap 1/16 vs 11/16. Packet HitFile 0.233 diagnostic | +| Meaning-grep LLM-as-judge | [jev-semgrep](https://github.com/uehaj/jev-semgrep) | 10 cases × 51-line EN/JP corpus. Precision 0.94, recall 0.98 *theirs*. Not a Harbor taskset. Dedicated fold `notes.md` §86 | +| Closed-vote CU worked example | [JevOnly](https://github.com/buluoray/JevOnly) | 11 steps / 43 Jev calls / ~340k tok / ~$0.014 / 17 s *theirs*. No planner LLM. Not a bake-off | +| Host-owned product evals | [waymode](https://github.com/mossburgh/waymode) | 24/26 public suite; 34/36 completion regression *theirs*. Bounded development evidence, not a self-driving proof | +| OMP prompt suppression (permission vs probability) | [omp-greenlight](https://github.com/SemetricLabs/omp-greenlight) | 1,013 gated calls / 10 sessions / 8.95 h. Default **40.9%** prompts removed; **0 of 94** unsafe auto-approvals on 140-row corpus. Live traffic unlabelled. Operator owns bar. Not a sandbox. ~$0.05 / 1,013 *theirs* | +| Jev question as if-statement (instrument not score) | [dinostomp](https://github.com/collapseindex/dinostomp) | `dinostomp jev`: accuracy, p(yes) cut, ECE, blank-input lean, rewording flips. Demo *theirs* 24 examples: 100% / ECE **0.062**. FINDINGS 189 (F 52 / D 99 / N 38); 99 against itself. Beside jevals, not a Harbor taskset | +| SLO routing latency cost (sync Jev vs local features) | [slo-router](https://github.com/zeeshan8281/slo-router) | Live Jev vs `slo_no_jev` on sim backends. Same routes (fast 4 / strong 4) and 100% accuracy; p95 E2E **77.93 → 490.38 ms** (~6.3×). Jev feature p50 453.58 / p95 1257.50 ms. 16/16 Jev calls; no lexical fallbacks. 3/8 task-label disagreements did not change routes. Eight-row demo is **not** a benchmark. License null. *Their* integration card | +| Effect-based shell-gate certification | [construct-auto-classifier](https://github.com/godspede/construct-auto-classifier) | Main 113 + blind 82; 5 passes; **975 decisions/model**. Jev: **0** dangerous allowed, 100% caught, 99.5% correct, $0.047/1k. Every chat model leaked 16–104 dangerous. Only Jev certified. Through the whole gate, not a Harbor taskset | +| Docs-derived instruct seed | [INSTRUCT_JEV](https://huggingface.co/datasets/ctaxnagomi/INSTRUCT_JEV) | 119 rows (47 choice / 51 noul / 21 score); 24 typed question blocks / 7 typed answers. MIT. Open-replica / jevals seed. Not a bake-off | +| Evidence-gated question packs | [jev-packs](https://github.com/dtduc-git/jev-packs) + [jevassert](https://github.com/dtduc-git/jevassert) | Nine packs `verified` on pinned `jev-1.13.0` *theirs*. **Runner LANDED** (Apache-2.0; was 404 §64). Record/replay CI: accuracy/ECE/Brier/cost/latency offline; exit 0/1/2; McNemar. First matrix 2,990 cases: Jev/Sonnet 5 accuracy tie (Δ≤0.018); Jev better calibrated 7/9; ~250× cheaper. sms-spam this-pass 0.953/0.040. `unknown` mandatory. CC0 packs. Not a Harbor taskset | +| Ranking ≠ calibration (human annotations) | [does-jev-confidence-mean-anything](https://github.com/Adilmp/does-jev-confidence-mean-anything); [jevcal](https://github.com/Adilmp/jevcal) | 8,000 judgments, `jev-1.13.0`, $0.05. AUC **~0.91**; stated **~75%** vs human **~10%**. Recalibration removes **~96% ECE**, AUC unchanged. `natural`/tightened ECE 0.156 → 0.006. jevcal: ~100 rows (94% of error). ECE gameable (constant base-rate ECE 0). License null / MIT. One domain; do not cite `threat` (n=1) | +| Hot-click CU vs per-step LLM | [ego-jev](https://github.com/jiangkoumo/ego-jev) | Alternate 3-round medians *theirs*: HN 4.9 s vs 9.7 s; wiki 5.4 s vs 10.1 s (~2×). n=3; high variance (control 7.3–22 s). **Not a benchmark.** MIT | +| Verbatim compact vs truncate vs summarize | [jev-compactor](https://github.com/edwardyen724-g/jev-compactor) | Earlier vs-Sonnet card: 64.5% / 366 ms / 4 of 4 (`notes.md` §65). Later product-arm table *theirs*: **73%** (53–76%) / **350 ms** / $0.0004 / **4 of 4** vs Anthropic 86%/16.8s/3 of 4, Codex 85%, OpenCode 85%, Gemini 61%/4 of 4. 61k session 95.4%/593ms/$0.0014. 30–250× cheaper. Two synthetic sessions, not a survey. MIT | +| Pre-send tool-result views | [dizk/jev-lens](https://github.com/dizk/jev-lens) | 500 SWE-rebench trajectories; 3,300 large results; 11.6M → 2.4M = **79%** fewer tokens. 88% command / 31% code. Post-send prune +17% cost. Harm: 2/26 later edits missed block; 0.3% dropped line quoted; 2.2% dropped identifier. Claude plugin unmeasured. MIT. Distinct from rashedInt32/jev-lens | +| tools≠use / SessionStart | [jev-carryforward](https://github.com/Dharundp6/jev-carryforward) | Plugin eval: `recall` **0/4** with tools+skill. SessionStart hook is the actual intervention. 9×3 remains a hint. MIT | +| Independent open-Jev class | [openvons](https://github.com/genai-craft/openvons) | LM 4B+head 0.916 vs 27B zshot 0.875; 8q / 22.6 ms. Vision 1/34 VRAM 36×. Voice 50 ms; 100% chatter reject. JevPick 3.2–4.8× byte-identical. Flutter 2.0–2.2 s / 11 ms for 9 q. Apache-2.0 LICENSE / GitHub SPDX NOASSERTION. Not TypeSafe | +| Physical-world S1 | [HA-Jev](https://github.com/AboveColin/HA-Jev) | `background:` triples laundry separation *theirs*. Batching 3q 712 ms vs 100q 714 ms. 30 commands $0.0017. Confidence uncalibrated. 192 mocked tests. MIT; **17★**. Not for locks/heaters | +| Same-intent VOI cache | [jevcache](https://github.com/kushals256/jevcache) | n=100 live Jev **fp=0 / precision=1 / recall=0.38 / fpr=0** vs Jaccard@0.35 fp=24 / fpr=0.48; $0.00174 *theirs*. Fail-open. MIT | +| Zeroshot vs BERT-family | [jev-zeroshot-vs-bert](https://github.com/zhuyansen/jev-zeroshot-vs-bert) | Beats DeBERTa-c on 7 sets (+0.05–+0.13; PAWS AUC +0.03; arXiv 2026 +0.30). Contaminated 0.901 vs `-c` 0.763. ≈230 / >2048 labels. Banking77 512+ feature **hurts**. DiD 0.035 vs 0.112. MIT | +| Worth-your-attention VOI | [ThinkyMiner/Winnow](https://github.com/ThinkyMiner/Winnow) | Unreviewed goldens **80%** verdict / **90%** content-type *theirs*. Distinct from kevinpita/winnow. MIT | +| Local OpenJev `/v1/decide` | [IamBusy/OpenJev](https://github.com/IamBusy/OpenJev) | v0.3 **45/60** vs v0.2 39/60; reversal 100%. Not TypeSafe drop-in. Apache-2.0. Distinct from hraness/sysone runners | +| SemIf `/v1/systemone` runoff | [semif-serve](https://github.com/dddanielliu/semif-serve) | RTX 3080 Ti Qwen3.5-4B **1164 ms** vs hosted **178 ms** *theirs*. Wire-compat ≠ replica. pyproject MIT / GitHub SPDX null | +| Conflict ≠ ignorance | [jev-typed-evaluation-collapse](https://github.com/mleyvaz/jev-typed-evaluation-collapse) | Noul 0.50–0.57 vs 0.46–0.48; named Choice p=1.0; binary red 0.67–0.85 *theirs* (v0.3). License null. NCML field note | +| BM25 vs Jev skill routing | [pi-jev-skill-bench](https://github.com/iamdin/pi-jev-skill-bench) | 43 gold; roster 50–500. Harness, not a production claim. **No live Jev numbers this pass.** MIT | +| Decision-as-memory flywheel | [DGUI_HYPERMEM-JEV](https://huggingface.co/datasets/ctaxnagomi/DGUI_HYPERMEM-JEV) | 6 rows (analyze 4 / rerank 2 / supersede 0). Sibling INSTRUCT_JEV. MIT card | +| Stop-hook attention redirect | [jev-preflight](https://github.com/muse0509/jev-preflight) | Owner-run Claude Code 2.1.267: no-key fail-open PASS; key-enabled exactly one continuation. Live API smoke: jev-1.13.0, 898/151 tokens, eight Nouls. 0.85 uncalibrated. Go MIT | +| Receipts-not-leaderboard capability map | [jev-capability-atlas](https://github.com/Zaious/jev-capability-atlas) | Hold vs break with API receipts; not a ranking. Type-safe ≠ correct (DAIR Emotion 48% / mean conf 0.819). Axis already §49; 10★ this pass. README MIT / GitHub NOASSERTION | +| Jev vs thinking-budget Qwen3.5 | [jev-frontier-100](https://github.com/softpudding/jev-frontier-100) | 100×3; Jev **77.0%**; 4B off 56.0% / 512 78.3% / 2048 **96.7%** (+12.7 to +26.7). 2B/2048 82.0% (−2.3 to +12.3). Exploratory, not preregistered. MIT. Not a ceiling | +| OOD calibration / sign by type | [jev-ood-calibration](https://github.com/scienthoon/jev-ood-calibration) | 900 synthetic + 3,721 public; ~$0.06. Public OpenBookQA ECE 0.024 / T 0.96. Synthetic all ECE **0.107 = 4.4×** floor; priority 44.7% / mean p 0.74 / T **3.40**; boolean T **0.66**. MIT. Gateway has no model version | +| Memory-lease invalidation | [invalidate](https://github.com/chopratejas/invalidate) | 157 labeled cases *theirs*: 89.2% strict / 97.5% lenient / **0 of 157** false invalidations. Six Nouls then code. Apache-2.0. MED | +| Cost-aware PR prune (pilot) | [prune-review](https://github.com/shubhangi013/prune-review) | 22 paired runs: winning-only 27.9% (post hoc); all 22 incl. 305% outlier **1.18%**; excl. outlier 15.9%. Cost not quality. Source preview. Size 365 / 1★ this pass | +| Whole-repo intent (CLI, under construction) | [jev-intent-review](https://github.com/yottayoshida/jev-intent-review) | VERIFIED/VIOLATION/UNKNOWN/NOT_APPLICABLE. Empty search ≠ proof. missed-path 7–8 req / 2–3 s; omamori #559 31 req / 14 s *theirs*. Action not written | +| Failure-finding arena (not a leaderboard) | [jevarena](https://github.com/chenmingtang830/jevarena) | Apache-2.0 TS. Public preview. JevJudge-Bench harness **not measured findings**. **≠** meetr1912/jev-arena | +| BBQ stereotype/uncertainty/cost | [jev-bbq-experiment](https://github.com/simonmesmith/jev-bbq-experiment) | 58,492 Q; Jev 1.13.0 **97.28%**; amb 99.96% / inf 94.60%; bias 0.04 / 0.34; **$0.3429 / 7.75 min** *theirs*. 12/13 amb errors stereotype-aligned. Order diagnostic 1/484. License null. Not a bias cert | +| Sentence-as-rule lint corpus | [jevlint](https://github.com/mizchi/jevlint) | 13/15 naming/comment rules **1.00/1.00** *theirs*; comment-describes-block ships unseparated. Review 2 req / $0.00013. **≠** huntedman/JevLint | +| Rust/WebGPU System One (JGLUE) | [grande](https://github.com/bokuweb/grande) | E2B zshot JNLI **0.614** ECE 0.252→**0.088** T=2.81; JCQA **0.853**. 270M **0.710/0.710**. Isolation 0.098/0.996. Packed Δmax 7e-5. License null. Softmax ≠ Noul until T | +| Clojure Laya byte parity | [laya-jolt](https://github.com/jlt-commons/laya-jolt) | Byte-identical to Python `system_one` on README quickstart *theirs*. ~1e-7 last-digit drift. Apache-2.0. Was empty skip §61 | +| ONNX ModernBERT vs live Jev | [local-jev](https://github.com/kunchenguid/local-jev) | 136 checkpoints *theirs*: done **30%** / shape **57%** / r **−0.06**; gold done 26% vs Jev 87%; 112 min vs 21 s. Confidence omitted. Not equivalence | +| Persist constraints (pi) | [pi-heed](https://github.com/Nyarlathoteppppp/pi-heed) | 79 sessions / 261 labelled: v0.8.0+Jev recall **98.5%** / false block **0.0%** / $0.000058 *theirs*. Mid-session rule change 8/13 off vs 0/13 on. Fail-open | +| Harbor SGR-judge contract (no quality headline yet) | [jev-judge-bench](https://github.com/slavadubrov/jev-judge-bench) | Frozen SLA-150. Jev vs Luna / DeepSeek-flash / glm-5.3-flash. Invalid = FN. 21 offline tests. Canaries *theirs* **not quality**: Jev OpenRouter 5/5; Luna 10/10; DeepSeek GA 10/10; DeepSeek beta 8/10; GLM 5.3 10/10; GLM 4.7 4/10 overload. $10 live Berlin in progress. Direct TypeSafe untested. Five-field/H5 untested. README MIT / GitHub SPDX NOASSERTION. **≠** chenmingtang830/jevarena **≠** fstandhartinger/jevbench | +| OpenRouter recipe samples (not benches) | [jev-cookbook](https://github.com/nexibeo/jev-cookbook) | 15 recipes; samples 16–36 handmade. Live 2026-09-19 `jev-1.13-20260917`. Recipes 01–13: 425 calls / **$0.015**; median 0.34–0.45 s; browser 5/6 *theirs*. Authors: scores show technique, **not benchmarks** | +| Hand-no-text plugin loop | [jev-use](https://github.com/shitianfang/jev-use) | Vercel `typesafe-ai/jev`, 95 calls *theirs*: p50 **220 ms** / p95 423; 12q **186 vs 2,672 ms**; 20-step 4.3 s / 0 escalated; gate **12/12** / p50 199 ms. First loop 17/20 escalate then 0/20 at margin 0.4. MIT v0.4.1. **≠** jev-ultrafast | +| Competing NAR claims (audit, not endorsement) | [openJev-verdict-2.0](https://github.com/Heman10x-NGU/openJev-verdict-2.0) | README *theirs* N=2000: acc **77.10%** / Brier **0.0636** / ECE corr **0.0144** / dist **0.1513**. Jev row is Laya-catalogued vendor baseline, not independent. **Open PR #1**: 24.7 dec/s misread as 25 ms (actual 40.5 ms; 3.5× not 28×); Laya 76.60% inside 95% CI (parity); like-for-like dist ECE 15.13% vs 21.40%, Jev 14.40% slightly lower. **≠** IamBusy/OpenJev | +| 1-token logprob local vs Jev (GUI 336) | [chakuho](https://github.com/taku-me/chakuho) | *Theirs* 2026-09-19 DGX Spark. 27B NVFP4: ordinary **245/258 (95%)** / sheets **72/78 (92%)** vs Jev Gateway 230 (89%) / 64 (82%). `__none__` gold 30: 29 vs 27 vs 8B **3**. Coverage ≠ correctness. Softmax ≠ Noul. MIT | +| Open replica engine speedup | [jevinf](https://github.com/zerodegress/jevinf) | **2.57×** one request (25.80→10.06 s) / **2.27×** dev split (84.2→37.0 s) at **100% argmax** *theirs*. Not ECE. MPS only. MIT | +| Files-to-read n=16 SWE | [jevex](https://github.com/jimmyhealer/jevex) | agy + Gemini 3.8 Flash, 1200s cap *theirs*: 160s→**69s**, $8.74→**$3.13**, 16/16 both arms. 90s cap 1/16 vs 11/16. Keep n=8 1/8→6/8 | +| Commit pre-review calibration | [commitjev](https://github.com/yodablocks/commitjev) | 13 labelled: every rule fires on its defect; **0 false on 5 clean** *theirs* (small control). Own 16 commits 3 warn / 4 review / $0.0017. Spread 0.01–0.09 | +| Laya multilingual MASSIVE / XNLI | [laya-multilingual](https://huggingface.co/convaiinnovations/laya-multilingual) | 51 langs *theirs*: acc **0.366** / ECE **0.387** / 45 of 51 ≥3× random vs English laya 0.227 / 0.733 / 23 of 51. Khmer 0.000@0.952 conf. XNLI 14-lang 0.731 vs 0.521. Ships uncalibrated ECE 0.314→0.106 after T. Apache-2.0 | +| Schema-conditioned DeBERTa scorer | [jev-schema-scorer-deberta-v3-large](https://huggingface.co/mobarmg/jev-schema-scorer-deberta-v3-large) | Hub MIT; GitHub **404**. v2 Choice **0.841** (chance 0.214) vs v1-only 0.687 *theirs*. Peaked p = ranking. Synthetic English | +| HF access this pass | open-jev-laya-bench / jev-tree-choice-cap / INSTRUCT_JEV / jevlogs-log-triage-benchmark | First three **401** (do not re-fold; no new numbers). jevlogs GitHub 404 **and** HF 401 | +| Productized public classification (vs_jev + caveats) | [classifier-dev](https://github.com/mrmps/classifier-dev) | MIT; **185★**. README *theirs*: 400 headlines **650 ms**; packing 100 = one-at-a-time; emotion ≥0.9 → **82%** / <0.5 → **29%**; multi-label F1 **0.887** / **230 ms** vs cascade **0.799** / 1.5 s; gemini-3.8-flash 87.5→90.0 / 61.8→63.7. eval *theirs*: 232 ms; AG News **87.7%** vs ling-3.0-flash **82.0%**; emotion **60.5%** vs **57.0%**; granite-4.0-h-micro F1 **0.546** vs advertised ~**0.800**. n=7 train-on-test; ~0.03 coin flip. `/benchmark` = tracked `vs-jev.json`. Not a Harbor taskset | +| Systematic-review pointer (spot check, not a validation study) | [choxos/jev-reviewer](https://github.com/choxos/jev-reviewer) | MIT; **12★**; https://jevreviewer.xera.ac. **≠** egma-ai. README *theirs* sample study (712 lines, Sep 2026): 1q 10 req / 1.2–2 s / $0.0016; 9q 17 / 2.3 s / $0.0052; 18-q template 27 / **4.6 s** / **$0.0101**. Quotes = Noul ≥ 0.5. *Not found* is an answer. Treat as spot checks | +| Prompted-JSON `/v1/systemone` bake-off (not logits; not calibrated) | [githubnext/localjev](https://github.com/githubnext/localjev) | MIT; **261★**. 1,200 req / ~23.5 min / M5 Max / oMLX 0.6.4 / Bun 1.4.0 *theirs*. Short macro: Qwen3.6 **76.7%** (AG News **90.0%**); Gemma 4 26B-A4B **75.0%** (SST-5 MAE **0.533**); DiffusionGemma **74.2%** (BoolQ **87.5%**). Qwen vs Gemma 26B = 2/120 — no definitive winner. Long-input both **69.2%**. Do not treat as calibrated. **≠** kunchenguid/local-jev. **≠** razorback16 structured-read | +| Laya packaging vs-Jev (third-party unpublished-here; post-T ≠ raw) | [NandhaKishorM/laya](https://github.com/NandhaKishorM/laya) | Apache-2.0; **710★**. README SHA `f12882b`. T4 *theirs*: 1q multilingual **32.8 ms** / 10q **72.3 ms**. Routed vs Jev 1.13.0 (Jev rows never measured here): typed-decisions **0.766** vs 0.727 (fine-tune; base 0.362/0.342 vs majority 0.461); Banking77 **0.425** vs **0.870** (77 vs 72 labels; ~3–4 tok/label); post-T ECE **0.081** vs 0.246 (raw 0.213 vs 0.144). Khmer **0.000@0.952**. Soft-acc 0.471 vs 0.580. 0.85 gating *theirs*. **≠** TypeSafe drop-in | +| External openjev census (tweet, not scores) | [@airesearch12 / Benchmark Heaven](https://x.com/airesearch12/status/2101259522933186879) | Named ~18 (system-one-open, openjev-sglang, DeBERTa open-jev, Needle 3, open-alternative-jev, Nimble 9B, SemIf, open-jev Dasein / JoshuaSP, OpenJev razorback16, mini-jev, system-one, system-one-gemma, jevlike, AlexWortega/openjev, GLiNER2, Succinct Router 14M, jev-model-router/Director/Loki). Engagement **ephemeral** (SIGNAL ~417/9/3; this pass 564/15/5). Watch [jev-models](https://benchmarkheaven.com/jev-models). **≠** jevbench v1.1. Scored sibling §78. Class-boundary: GLiNER2 + routers. Incomplete vs Laya/localjev/kev/TypeAR/openvons/… | +| Never-confidently-wrong protocol (TLA+ + chaos) | [jev-labs](https://github.com/copyleftdev/jev-labs) | 1,080 golden: 0 wrong under none/realistic/severe *theirs* (severe 314/46 escalate). Rule of three <0.28% at 95% — not a proof of zero. TLC 1,049,750 states / 0 errors. 1,490 calls `jev-1.13.0`. Synthetic, not clinical. MIT | +| Sureness metrics vs Jev `confidence` | [how-sure-is-jev](https://github.com/adarc8/how-sure-is-jev) | 60 live answers: Choice confidence = max_prob to 3 decimals. 75/25 → 0.5 vs entropy 0.19. Bands are policy. Zero-dep MIT | +| Jev-class bake-off v1.1 (historical) | [jevbench](https://github.com/fstandhartinger/jevbench) | 314 decisions. Main Score 0.6/0.2/0.2. Jev 1.13.0 **87.6** / Cap 97.8 / $0.0259/1k *theirs*. Calibration **reported, not scored**. Native vs verbalized. Partial runs not ranked. Unofficial MIT. **Superseded for the live board by v1.2** (not comparable) | +| Jev-class bake-off v1.2 (live board) | [jev-models](https://benchmarkheaven.com/jev-models) / [jevbench RESULTS-v1.2](https://github.com/fstandhartinger/jevbench/blob/main/RESULTS-v1.2.md) | Protocol `jevbench::v1.2`; scored 19 Sept 2026; 534 decisions (72/96/146/**220 hard**). Geo-mean I/C/S/K 25% each. Jev 1.13.0 **75.3** / SemIf **74.6** (−0.7) / OpenJev razorback16 67.6 *theirs*. Luna I **96.8** rank **#7**. Cal **ON** rank. Self-host latency ×2 assumption; many costs est. Option-order 72%→21%. Partial not ranked. Laya absent (gap, not named-excluded). Qwen3.8 27B Chutes TEE **≠** Archer. Unofficial MIT. README SHA `bf1e79ba`; RESULTS SHA `fdfab1a2`; HEAD `27ed3d6c`. **≠** tweet census **≠** v1.1 87.6 **≠** jev-judge-bench **≠** jevarena | +| Hourly 0842 already-folded watch (apply, don’t re-card) | `notes.md` §79 | Five HIGHs already §73–§78 (`5f44bf4` / `03fddc6` / `daa70b7` / `bcf66f1` / `db654b5`+`40a5f12`). Recipe: wire-compat ≠ logit-equiv; productize label+p + mark FALLBACK; packaging ≠ new species / script-before-p; pointer-not-generator two-pass; census ≠ scored bake-off. Skip thin (JEValuate / jevspeak / fable-jev; jev-semgrep now §86). Hard-gate Noul as PR/quality = soundness theater (totally-tim/jev-gate 0★ / claude-jev-warden 1★). Archer still Watch. Not a Harbor taskset | +| Local controller vs Live API (reflex A/B, not a bake-off) | [jev-reflex-autonomy-lab](https://github.com/khordoo/jev-reflex-autonomy-lab) | TypeScript; **7★**; license null. README SHA `130987c9`; ARCHITECTURE SHA `48da0769`; HEAD `e3297ebe`. Same physics/seed; toggle only changes where reflex decisions come from. Local = rule-based, no keys. Live = `POST /v1/systemone` `jev-latest`. 20% starting gate *theirs* does not start a mission. **≠** githubnext/localjev **≠** kunchenguid/local-jev. No ECE/taskset — Harbor-*adjacent* of backends, not a scored board. Seed = geometry ≠ async replay. `notes.md` §80 | +| OCR+AX desktop CU cost table (one screenshot, not a taskset) | [typesafe-computer-use](https://github.com/awlevin/typesafe-computer-use) | MIT; **427★**. README SHA `369f4a6a`; HEAD `cc7b5066`. *Theirs*: $0.0002 vs Opus 5 $0.032 (155×); 0.13–0.38 s vs 5.2 s; ~1.5 s vs ~5.5 s e2e. Honest caveat: dates.py rebuilds pixel-free reasoning. 0.4 / 0.5 still soft. **≠** jev-ultrafast Flights clock. Not a Harbor taskset. `notes.md` §81 | +| ASR voice-browser fixtures (not a taskset) | [jev-voice-browser](https://github.com/moritzkremb/jev-voice-browser) | JavaScript; MIT; **103★**. README SHA `fa033303`; HEAD `054db0f3`. *Theirs*: integration 27/27; Jev p50 ≈ 300 ms; ~$0.0002/call; demo ≈ $0.01. 0.5 / 0.55 / 0.6 still soft. **≠** Harbor. **≠** jev-voice-control. `notes.md` §82 | +| Wrap-as-execution (not a quality bench) | [AgentGhost](https://github.com/reddpy/AgentGhost) | TypeScript; MIT; **2★**. README SHA `44145fa9`; HEAD `ac04e4fb`. ASK throws; fail-closed. No Harbor numbers. **≠** actiongate **≠** toolgate. `notes.md` §83 | +| JP genre atlas (tweet, not scores) | [@studio_yebisu](https://x.com/studio_yebisu/status/2101065176069886152) | Apps by genre; SAM 3.1 + OpenRouter Jev noted. Engagement **ephemeral** (SIGNAL ~120k/1767/169; this pass 131,234/1,934/192). Stars research-time (typesafe-computer-use 203→427; jev-voice-browser 40→103). **≠** @airesearch12 **≠** v1.2. Not verified evals. `notes.md` §84 | +| External pedagogy (article, not scores) | [@akshay_pachaar](https://x.com/akshay_pachaar/status/2101037514945597645) | “Jev Clearly Explained.” 200×/400× TypeSafe ceiling *theirs*. schema-safe ≠ correct. Engagement **ephemeral** (SIGNAL ~183k/2095/220; this pass 233,495/2,280/235). **≠** Harbor. **≠** official docs. `notes.md` §85 | +| Meaning-grep dedicated (contrast-set, not scores) | [uehaj/jev-semgrep](https://github.com/uehaj/jev-semgrep) | JavaScript; MIT LICENSE / GitHub NOASSERTION; **51★** ephemeral (SIGNAL ★42; §61 0★). HEAD `21120e9`; README SHA `923e6a5`. Proposition ≠ embedding; all-six-refund contrast-set *theirs*; AND/OR/NOT after threshold; Semgrep.dev collision; not a gate. 0.94/0.98 keep as LLM-as-judge 10×51, not Harbor. `notes.md` §86 | +| Pre-review typed PR gate | [ci-gatekeeper-bot-jev](https://github.com/NemanjaManic/ci-gatekeeper-bot-jev) | Own-repo live Jev: 504–629 ms; secondary ~4–5 s only on human-review + elevated risk. Conservative default escalated trivial diffs. `package.json` MIT / GitHub SPDX null | + +rh-guard is a reward-hack hook, a different surface from jevgate and +from Abide (eval-integrity vs allowlist-remainder vs project soft +rules). One row each. [dinostomp](https://github.com/collapseindex/dinostomp) +is the **instrument** auditor beside that row: data/scorer/runs/claims, +plus `dinostomp jev` if-statement hygiene for a TypeSafe question +(`notes.md` §62). ECE above is wanted, not a Nimble result. +Abide replay (author-reported, not re-run; `notes.md` §47): 93 +sessions, 1,256 edits / 147 turns; independent-reviewer precision +**edit ~26% / turn ~73%** before calibrate/tune. Harbor-adjacent +measurement (frozen transcripts, phase split, independent +confirmation), not a Harbor taskset and not a jevals substitute. +Turn-phase soft rules held up better; false positives mostly fixable +in the rubric. Text/diff only. + +**Harbor-style computer-use receipt this hour:** +[solari-reflex](https://github.com/hitakshiA/solari-reflex) scores +the *task* (Stripe API / answer key), not a paragraph judge. Observe +→ decide → verified act; no screenshots. Author table vs Codex on +the same Solari machines: 60.2 s vs 194.9 s; 66 s vs 460 s; 24.2 s +vs 98.4 s (`notes.md` §48). Encoder-backend cousin: +[gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) +— same hole, local GLiNER2; their Flights demo (12.20 s visible / +13.785 s loop / ~$0.0001 API) is a **demonstration**, not a bake-off +or a vs-Jev-Ultrafast table; `DONE` is not the Harbor score +(`notes.md` §52). Specialist-form cousin, **not TypeSafe Jev:** +[Cua-S1](https://github.com/trycua/cua/tree/main/libs/cua-s1) — +option-attention among observed elements; plan ≠ execute; dry-run +default; source-only this pass. Offline utilities *name* accuracy, +abstention, coverage, wrong actions/targets, and unsafe-when-should- +abstain; **no checkpoint scores**. Tests exercise implementation, not +quality. Do not invent a vs-Jev table. Watch for `cua-s1-form-v0` +(`notes.md` §54). **Harness extract card (their PR body, not +re-run; 2026-09-19 ~00:48):** +[Stagehand #2955](https://github.com/browserbase/stagehand/pull/2955) +— gemini-3.8-flash, Browserbase, local, 25 tasks × 3: 69/75 vs +23/25 baseline (92% both). **37/75** no-LLM in **~0.5 s** vs +baseline **4.37 s** and two LLM calls; LLM-off **36/75**. Pick is a +**fast path, not a replacement.** Draft stack #2951–#2955. In-sample +thresholds on the act suite (#2953). Not Harbor. Do not merge with +solari / Flights / Cua-S1 clocks (`notes.md` §57). +**Collab-arm curriculum:** +[jev-testbench](https://github.com/ufx7/jev-testbench) — +`llm_autonomous` vs `scripted_plus_jev` vs `llm_plus_jev`; Wilson + +McNemar; Jev is not a peer arm. Bake into jevals/Harbor hygiene, do +not copy the harness. + +**Harbor on/off routing (Empirical as a *shape* and as a one-run +signal, not a measurement; 2026-09-18 ~16:48).** +[jev-gateway-bench](https://github.com/vinilana/jev-gateway-bench) +(MIT): real coding agents on chess-engine tasks; Jev routing on vs +off; hidden perft verifier the agent never sees; fresh gateway per +run. Author: first signal, not a measurement; <5 runs/mode the +summary says so. Preliminary `chess-bugfix` Codex: both 36/36 +checks; on 4 LLM req / 76,678 in / 35 s vs off 6 / 118,709 / 88 s; +Jev 4 calls ~$0.0008. A cheaper unsolved run is not a saving. A +wrongly forced tool can derail a turn. Product sibling +[jev-gateway](https://github.com/vinilana/jev-gateway) fails open if +Jev is down. Pair CI merge-gate +([latch](https://github.com/CaseReed/latch)) with this substrate +(frozen JUnit artifacts × PASS/BLOCK) and rh-guard (eval-integrity). +Do not copy npm/ports (`notes.md` §51). + +**Pre-registered independent eval (Empirical as Harbor/jevals +*practice*, including the honest negative; 2026-09-18 ~17:48).** +[jev-baselines-eval](https://github.com/ickma2311/jev-baselines-eval) +(MIT): kill/go printed by the scripts; **both AMBIGUOUS**. CLINC150 +Jev 0.870 vs nano 0.795 vs Terra 0.915. Banking77 encoder **0.933 / +9 ms** wins. Cascade Δ +0.265 at 1pp-below-frontier; **at exact +parity the sign flips** (R_jev=1.000) because confidence is exactly +1.0 on 102/200 including 6 wrong. AUROC neither direction; **no +ECE**. Latency ~2.2× of two serving paths, not 40–200×, not +model-speed. Same-day errata three rounds. Do not copy pip +(`notes.md` §55). + +**Healthcare Harbor-shaped receipt (Empirical as that named +report; not clinical validation; 2026-09-18 ~17:48).** +[explore-typesafe-ai](https://github.com/si618/explore-typesafe-ai) +— 100 synthetic Synthea patients; labels committed before any Jev +run; 60 requests to jev-1.13.0; Claude S2 blinded review. Report: +NEWS2 alone under-triaged 10/20, NEWS2+Jev 1/20; 403 judgments / +p50 329 ms / $0.0038. Claude wrote the labels. 20 cases/scenario. +**Not clinically validated.** License not in GitHub API this pass +(`notes.md` §55). + +**Honest-negative PDF lab (Empirical as a negative; no OSS +license).** +[databricks-jev-pdf-lab](https://github.com/laurentfabre/databricks-jev-pdf-lab) +— no quality-equivalent end-to-end Jev payoff. Compact tokens +changed 26/236 recommendations. Public snapshot cannot reproduce +historical accuracy. Typed output is not truth (`notes.md` §55). + +**Meaning-search Harbor-shaped card (Empirical as their stripped- +repo table; 2026-09-18 ~18:46).** +[jevgrep](https://github.com/Bentlybro/jevgrep): 228 questions on +Flask/httpx/Django/AutoGPT with docstrings and comments removed. +Right file in top 5: **79%** vs BM25 40% / grep 20%. Honest +negative: BM25 top-10 96% vs 85% when the exact wording is known. +Packed+parallel 0.9 s vs serial ~23 min on AutoGPT 4,329 files. +Frozen copies + labeled questions + comparable harnesses — not a +Harbor taskset, not Wilson/McNemar published. Do not copy +`install.sh` (`notes.md` §58). + +**Measured RAG rerank vs generative rerank (Empirical as one-run; +Hypothesis as a transfer).** +[Jev-RAG](https://github.com/Max-sm-yc/Jev-RAG): RAG+Jev+Spark +$0.00122838 / 62.3 s vs RAG+Spark-rerank+Spark $0.00421838 / +228.14 s vs Spark full-context $0.0032 / **10.60 s**. ≥70% cost +and 72% latency vs Spark *rerank*, not vs no-RAG. Costs include +embeddings. License null. Do not invent a bake-off +(`notes.md` §58). + +**jevals-shaped oxlint remainder (Empirical as Phoenix fixtures; +experiment).** +[jev-oxlint](https://github.com/cephalization/jev-oxlint): human +answer key vs Noul on every fixture (agree, wide margins); routing +sharp 0.80–0.94 vs <0.50 across 41 files; coarse hint not. Found +a real flush-only-on-success bug (noul 0.07). ~$0.002 / ~$0.015; +second run zero requests. Not a hard gate. `tenbin` owns the lint +skill (`notes.md` §58). + +**Native-probability calibration arena (Empirical as their live +card; 2026-09-18 ~19:48).** +[jev-arena](https://github.com/meetr1912/jev-arena): analytically- +known worlds; native `noul`/`choice`/`score`, not verbalized +confidence. Live `--live --trials 200 --seed 7`, `jev-1.13.0`: +**145 noul**, Brier **0.0059**, log loss 0.5393, ECE **0.0620**, +**2 requests / 710 ms**. Overconfident in low bins. Oracle stub +0.0000. Cite as *theirs*. Fan-out suite: sonar heatmap-as-policy +(offline 20-game Brier 0.1615 / 75% win; live 1-game small +sample); vickrey threshold CDF (Jev never bids; live Brier +0.1391); bracket Brier vs Elo (live **trailed Elo** 0.2853 vs +0.2322 — honest). Harbor/jevals-shaped: exact oracle, proper +scores, teeth stubs, offline default (`notes.md` §59). +**Typed control-plane bake-off shape (Empirical as metric list, +not as a quality number):** +[jev-dspy-control-plane](https://github.com/manikanda-kumar/jev-dspy-control-plane) +shares ontology/dataset/state/allow-list across OpenJEV / DSPy / +JSON Schema. Offline heuristic + contract stubs are plumbing +regression, **not** model generalization. Accuracy alone is not +enough (`notes.md` §59). +**Tetris legal-set demo (not a rigorous eval):** +[jev-tetris-benchmark](https://github.com/planstack-ai/jev-tetris-benchmark) +— code enumerates ≤12 legal placements; Jev Choice vs Haiku. +Same hole as jev-plays-games (`notes.md` §59). + +**Domain specialist vs few-shot hosted (Empirical as their +RESULTS.md; 2026-09-18 ~20:43).** +[Domain-jev-maker](https://github.com/help-er/Domain-jev-maker): +independent CLINC-150 labels, not a Jev teacher-copy. +Matched-precision KL (both rounded to two decimals, zeros → +0.0025): local 1.5B 0.168 vs hosted zero-shot 0.580 banking +(r +0.933 vs +0.343). Few-shot hosted determinate McNemar +n.s. (p=0.134 banking / p=1.000 travel). Train the specialist +when policy reads p; hosted+examples when only argmax. Do not +copy train how-to (`notes.md` §60). +**Cascade compare arms (Empirical as README + mock; not a live +Jev vs Haiku bake-off):** +[jav-email-cascade](https://github.com/skiingfalcon/jav-email-cascade) +— 74 labelled emails; jev vs gen-json vs gen-logprob on one +Answer schema. Mock: gen-json confidence essentially flat. +Noul 0.5 never rounded. License null this pass (`notes.md` §60). +**ORDER BY ranking family (Empirical as independent +measurement):** +[jev-orderby-bench](https://github.com/yodablocks/jev-orderby-bench) +— `jev-1.13.0` passes six pre-registered gates. Boolean +inversion 0.036; Score ordinal **0.143** vs 0.15 (weak link / +sort key); 53-way 0.99 tie; ECE 0.0453 / Brier 0.0524. +recodelabs batch-40 **fails** ranking (inversion 0.171). +Calibration ≠ sortable. Request shape is part of the +measurement (`notes.md` §60). +**Class-backend economics (Empirical as their RESULTS.md):** +[jeff](https://github.com/logan-markewich/jeff) — GLiFormer-400M +`/v1/systemone`. 1,600 items: L4 HTTP ~$2.6 vs jev ~$15.6 +(~6×); A10G direct ~$0.65 (~24×); AG News 75.5% vs 90.5%; p50 +151 vs 129 ms. CPU 6–20× *more* expensive. Encoder ≠ Jev +replica. License null this pass (`notes.md` §60). + +**Harbor Jev vs local MLX PCD vs AR JSON (Empirical as their +README table; 2026-09-18 ~21:39).** +[system-one-benchmark](https://github.com/mallahyari/system-one-benchmark): +`jev-1.13.0` vs Qwen2.5-1.5B 4-bit MLX PCD vs AR JSON on +lmsys/toxic-chat **n=50**. Jev **84.0%** acc, Brier +**0.1096**, precision 90.9% (1 FP), p50 356.5 ms. PCD 52% / +Brier 0.3884 / p50 227.2 ms / 1 pass. AR 54% / ~30.8 passes +/ 98% schema errors. **PCD proves O(1) speed; uncalibrated +likelihoods ≠ Noul.** License null. Clone URL still +`your-username`. Small n — *their* card, not a large Harbor +taskset. Cousin of DMB / open-jev-laya-bench / pcdServer / +jevify. Do not copy pip how-to (`notes.md` §61). +**Evidence-packet explorer (Empirical as their performance.md, +author-run):** +[jev-semantic-explorer](https://github.com/jimmyhealer/jev-semantic-explorer) +— SWE-bench Verified n=8: **1/8 → 6/8** finish (empty +output = miss); about half the model bill. Claude Code +6.8→2.2 files. Packet n=50 HitFile 0.233 vs BM25 0.159 is +**not** the product KPI. n=8 is small (`notes.md` §61). +**Meaning-grep judge test (Empirical as their report.md):** +[jev-semgrep](https://github.com/uehaj/jev-semgrep) precision +0.94 / recall 0.98 *theirs* (LLM-as-judge, cached verdicts, +10 cases × 51-line corpus). Not Harbor. Dedicated fold: +proposition ≠ embedding; contrast-set; Semgrep.dev; not a +gate (`notes.md` §61, §86). +**OMP prompt suppression (Empirical as measured traffic + +labelled corpus; 2026-09-18 ~22:38).** +[omp-greenlight](https://github.com/SemetricLabs/omp-greenlight): +1,013 gated calls / 10 sessions / 8.95 h. Default **40.9%** +prompts removed; **0 of 94** unsafe auto-approvals on a +140-row corpus. Live traffic unlabelled. Operator owns the +bar; plugin never self-tunes. Not a sandbox. ~$0.05 / 1,013 +*theirs*. Do not copy `omp plugin` (`notes.md` §62). +**Eval-instrument / Jev-as-if (Empirical as FINDINGS.md + +demo card; Harbor/jevals-adjacent hygiene).** +[dinostomp](https://github.com/collapseindex/dinostomp): +checks the instrument, not just the score. FINDINGS 189 +(F 52 / D 99 / N 38); 99 against itself. `dinostomp jev` +demo *theirs* (24 examples): 100% accuracy, ECE **0.062**, +blank 'no' at 0.81, 0/60 rewording flips. Beside jevals, +not a Harbor taskset. Do not copy pip (`notes.md` §62). +**SLO routing latency cost (Empirical as live analysis +*negative* for sync Jev; 2026-09-18 ~23:40).** +[slo-router](https://github.com/zeeshan8281/slo-router): +sim backends + real Jev. SLO no-Jev p95 **77.93 ms** vs +SLO+Jev p95 **490.38 ms**; accuracy 100% both; same routes. +Jev disagreed on 3/8 task labels and did not change routes. +Author: keep Jev off the synchronous path for this +workload. Eight-row demo is not a benchmark. Harbor-style +ablation of decision-model latency. License null. Do not +copy uvicorn (`notes.md` §63). +**Effect-gate certification (Empirical as their report; +not a Harbor taskset).** +[construct-auto-classifier](https://github.com/godspede/construct-auto-classifier): +975 decisions/model; Jev **0** dangerous allowed; every +chat model leaked 16–104. Privilege ≠ verdict. Fail-closed. +Pair with dinostomp (instrument) before treating 0/975 as +class truth. Do not copy bun (`notes.md` §63). +**Instruct seed (feedstock, not a score):** +[INSTRUCT_JEV](https://huggingface.co/datasets/ctaxnagomi/INSTRUCT_JEV) +119 rows (47/51/21); 24 typed questions / 7 typed answers. +jevals-shaped open replica. MIT. +**Evidence-gated packs (Empirical as registry tables + +landed runner; Harbor/jevals practice; 2026-09-19 ~05:46).** +[jev-packs](https://github.com/dtduc-git/jev-packs): nine +packs `verified` on pinned `jev-1.13.0` *theirs*. +[jevassert](https://github.com/dtduc-git/jevassert) **LANDED** +(Apache-2.0; was 404 §64). `check` is offline from +recordings; accuracy/ECE/Brier/cost/latency; exit 0/1/2; +McNemar. Matrix 2,990 cases: Jev/Sonnet 5 accuracy tie; +Jev better calibrated 7/9; ~250× cheaper *theirs*. +`unknown` mandatory. CC0. Distinct from INSTRUCT_JEV +(no evidence gate) and dinostomp (instrument). Do not +copy uvx (`notes.md` §64, §70). +**Harbor SGR-judge contract (Empirical as frozen protocol, +not a quality score; 2026-09-19 ~06:43).** +[jev-judge-bench](https://github.com/slavadubrov/jev-judge-bench): +SLA-150; Jev vs schema-guided LLM judges; invalid = FN; +cost/latency first-class. 21 offline tests. Canaries *theirs* +are availability, **not** F1. **No quality headline yet.** +Distinct from jevarena (failure-finding) and jevbench (Main +Score). README MIT / GitHub SPDX NOASSERTION. Do not copy +`uvx` (`notes.md` §71). +**Competing NAR claim-audit (not endorsement; 2026-09-19 +~06:43).** +[openJev-verdict-2.0](https://github.com/Heman10x-NGU/openJev-verdict-2.0): +README *theirs* 77.10%/0.0636/0.0144. Open PR #1 already +corrects throughput≠latency and Laya-parity. Like-for-like +distribution ECE vs Jev is not a win. **≠** IamBusy/OpenJev +(`notes.md` §71). +**Ranking ≠ calibration (Empirical as human-annotated +audit + tool).** +[does-jev-confidence-mean-anything](https://github.com/Adilmp/does-jev-confidence-mean-anything): +8,000 judgments vs `civil_comments`; AUC ~0.91; stated +~75% vs human ~10%; ~96% ECE removed without rank change. +[jevcal](https://github.com/Adilmp/jevcal): ~100 labelled +rows; demo 0.9 → 33% on 1,600. ECE gameable — they decide +on Brier. One domain. Do not cite `threat` (`notes.md` §64). +**Hot-click CU (Empirical as n=3 medians, not a bench; +2026-09-19 ~00:39).** +[ego-jev](https://github.com/jiangkoumo/ego-jev): HN 4.9 s vs +9.7 s; wiki 5.4 s vs 10.1 s vs per-step `kimi-k3`. High +variance. Selector-hardcoded code beats both. Do not copy +`install.sh` (`notes.md` §65). +**Verbatim compact vs summarize (Empirical as one synthetic +session).** +[jev-compactor](https://github.com/edwardyen724-g/jev-compactor): +64.5% / 366 ms vs Sonnet (`notes.md` §65); later product-arm +**73%** / 350 ms / 4 of 4 vs shipped summarizers (`notes.md` +§68). Two synthetic sessions. Do not copy npm. +**Eval integrity cluster (Empirical as their tables; +2026-09-19 ~01:47; not leaderboard theater).** +[jev-capability-atlas](https://github.com/Zaious/jev-capability-atlas) +receipts, not a ranking; type-safe ≠ correct (axis already +§49; 10★ this pass). +[jev-frontier-100](https://github.com/softpudding/jev-frontier-100): +Jev 77.0% vs Qwen3.5 4B/2048 96.7% (4B off 56.0%); +exploratory, attach the thinking budget. +[jev-ood-calibration](https://github.com/scienthoon/jev-ood-calibration): +900 tickets ECE 0.107 = 4.4× floor; priority unknowable +(44.7% / mean p 0.74 / T 3.40); sign flips by type. +Do not copy npm / Ollama (`notes.md` §66). + +**VOI cache / zeroshot displacement / skill-routing harness / +typed-evaluation collapse / local class (Empirical as their +tables; 2026-09-19 ~04:39).** +[jevcache](https://github.com/kushals256/jevcache): n=100 +live Jev **fp=0 / precision=1 / recall=0.38 / fpr=0** vs +cosine-Jaccard@0.35 fpr 0.48; $0.00174 *theirs*. Fail-open. +[jev-zeroshot-vs-bert](https://github.com/zhuyansen/jev-zeroshot-vs-bert): +Jev beats clean DeBERTa-c on all 7 sets (+0.05 to +0.13 +acc; PAWS AUC +0.03; arXiv 2026 +0.30). Contaminated NLI +AG News 0.901 vs `-c` 0.763. Label-equivalence ~230 / +>2048. lr-bge+jev hurts Banking77 at 512+ (−0.044). DiD +Jev drop 0.035 vs DeBERTa-c 0.112. Cost not logged. +[ThinkyMiner/Winnow](https://github.com/ThinkyMiner/Winnow): +unreviewed goldens **80%** verdict / **90%** content-type +*theirs*. Distinct from kevinpita/winnow. +[IamBusy/OpenJev](https://github.com/IamBusy/OpenJev): +**45/60** vs v0.2 39/60; reversal 100%. `/v1/decide` ≠ +TypeSafe. +[semif-serve](https://github.com/dddanielliu/semif-serve): +1164 vs 178 ms *theirs*. Wire-compat ≠ replica. +[pi-jev-skill-bench](https://github.com/iamdin/pi-jev-skill-bench): +43 gold; roster 50–500; **no live Jev numbers this pass**. +[jev-typed-evaluation-collapse](https://github.com/mleyvaz/jev-typed-evaluation-collapse): +Noul collapses conflict 0.50–0.57 vs ignorance 0.46–0.48; +named Choice separates p=1.0; binary Choice red 0.67–0.85 +*theirs* (manuscript v0.3). +[DGUI_HYPERMEM-JEV](https://huggingface.co/datasets/ctaxnagomi/DGUI_HYPERMEM-JEV): +6-row flywheel (analyze 4 / rerank 2 / supersede 0); +sibling INSTRUCT_JEV. +TeoMastro `bench/results/summary.md` **404 this pass** — +do not invent numbers. `notes.md` §69. Do not copy npx / +uv / plugin how-to. + +**Harbor-adjacent stdout prune (Empirical as README / evals README +behavior, not a full Terminal-Bench ranking; 2026-09-18 ~17:15).** +[jev-pruner](https://github.com/tamaratran/jev-pruner) ships +needle/noise graders and a Harbor Terminal-Bench 2.0 adapter in-repo. +Manual `trimOutput` sweep (theirs, 2026-09-18, 3 runs, `jev-latest`): +needles **24/24**; mean reduction **83% (71–92%)** on trim scenarios; +wrongly trimmed **0/12**; mean latency **240 ms**. Wider: standard +8/8 / 83%; accuracy 36/36 / 87%; real captures 10/10 / 54%; needle +matrix 9/9. `claude plugin eval` **cannot exercise pruning** (Jev +fetch refused → fail-safe original). Six-run paired Terminal-Bench +pilot is **integration, not a significance test** (full set 89 tasks +/ 178 trials; no published full-run scores this pass). Do not merge +those tables. Do not copy the Harbor launcher (`notes.md` §53). + +**Harbor-style frozen protocol vs constrained LLMs (Empirical as that +named receipt, not a ranking).** +[`nibzard/decision-model-benchmark`](https://github.com/nibzard/decision-model-benchmark) +(DMB): jev vs 8 constrained LLMs vs keyword/majority/random; five +suites; **$28.34**; raw logs. **`results/v2/v2.md` is the report of +record.** Protocol frozen before the run; negative results ship; +unknown usage is never a measured zero; later runs replace cells +whole. jev S1 banking **76.3%**, S2 spam **93.0%**, S3 **100%*** at +valid coverage **72.7%** (225 failed = 256+ Choice cap), S4 flip +**13%**, S5 admits-ignorance **49.7%** / ECE **0.246**; p50 +**264–276 ms**; S1 cost/1k **$0.07**. No class wins on quality. +Do not copy `uv`. Do not merge this Banking77 with atlas 87% or +jevals.com 79.67% (`notes.md` §49). + +**Feedstock / recompute-from-logs (not a third ranking).** +[`Jevals/jevals-data`](https://github.com/Jevals/jevals-data) +(CC-BY-4.0): release boards + per-decision JSONL + suite files for +[jevals.com](https://jevals.com). 2026-09-18 board, suite 0.1.0, 8 +systems (banking77 / helpsteer2 / pubmedqa). Formulas: +https://jevals.com/methodology/. Jev on *this* board (n=300×5): +banking77 acc **0.7967**, ECE **0.0981**, p50 **467 ms**, cost/1k +**$0.043**. Cite the release; recompute from logs; do not dump the +board as a ranking. + +**Negative: combinatorial assembly ≠ extractive keep/drop.** +[`simonmesmith/jev-arc-agi-v1-experiment`](https://github.com/simonmesmith/jev-arc-agi-v1-experiment) +— Direct Jev on ARC-AGI-1 public eval **4/400 (1%)**, 1.125% +task-weighted, ~$2.32, 10 min. Cell-wise Choice; dimensions ~90%; +rarely a complete grid. A frozen Harbor-shaped protocol that +falsifies "many small decisions add up to a puzzle." ### Bake-off mandate -Before adopting proprietary Jev vs Laya vs TypeAR vs Nimble vs Archer -vs openjev-lm, run a jevals-shaped labeled suite (or an equivalent +Before adopting proprietary Jev vs Laya vs TypeAR vs Nimble vs kev vs +blackwood-rlcd vs Archer vs openjev-lm vs a constrained LLM vs von vs +open-alternative-jev, run a jevals-shaped labeled suite (or an equivalent with this hygiene) and, for a product loop, a Harbor taskset. A design -card with no eval path is incomplete. +card with no eval path is incomplete. A green smoke test on +[jev-local](https://github.com/us/jev-local)'s **default stub** is not +that bake-off (`notes.md` §48). von's 14 MB needle at 52.6% authored144 +is not that bake-off either (`notes.md` §49). DMB is the frozen-protocol +exemplar for decision-model vs constrained-LLM vs baselines; jevals-data +is the public log feedstock. Do not promote a vendor table into a ranking. + +**Shared bake-off exemplar (Empirical as that named receipt, not a +ranking).** [`pngwn/open-jev-laya-bench`](https://huggingface.co/datasets/pngwn/open-jev-laya-bench) +(`RESULTS.md` this pass): System One Qwen3.5-4B scorer vs Laya 421M, +**26 neutral + 9 home**, **11,959** test / **3,269** cal. Scores are +accuracy + **ECE / NLL / Brier** (acc@50% coverage too). Neutral macro +acc Δ **+0.023 [+0.013, +0.032]**; home Δ **+0.229 [+0.198, +0.262]** +(intervals exclude 0). **Not TypeSafe Jev vs Laya.** Neutral prompted- +instruct on the same 4B is statistically tied with the fine-tune +(+0.003, interval includes 0). **LLM-as-judge is not the primary +System One score.** Harbor/jevals practice in the wild: held-out +`test`, temperature on `cal`, leave-one-task-out, prompted arms. +`notes.md` §46. Do not copy the scoring scripts. Archer weights are still a **Watch** — not on the Hub as of 2026-09-18 -(`notes.md` §31–§33). That bake-off is future, not Empirical. "A 9B +(`notes.md` §31–§33). That bake-off is future, not Empirical. kev is +the shipped 0.5B reconstruction on the trained decision-only path, not +that drop (`notes.md` §45). blackwood-rlcd is the open multimodal +decide head on that path **now** (CC BY-NC; Jev still leads general +text; `notes.md` §46). "A 9B LoRA is enough versus Jev" stays **Hypothesis** (`notes.md` §35). -openjev-lm is the name of that distill. Nimble is not a Jev distill +openjev-lm is the name of that distill. kev is not a Jev distill. +Nimble is not a Jev distill (model card Apache-2.0; GitHub LICENSE was 404). Meijer: marginals, not a PPL, not Kleisli (`notes.md` §34). djev-spark is a third compute graph, not the winner of this bake-off (`notes.md` §36). Do not diff --git a/CHANGELOG.md b/CHANGELOG.md index ec703d1..8e3aef6 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -99,6 +99,910 @@ Pegged against [`typesafe-ai/skills` v0.5.7](https://github.com/typesafe-ai/skil two-layer finish gate. pi-jev (not pi-jev-context). jev-plays-games option-order probe. joxide jump-by-description. laya-typed-decisions companion packaging. No wrapper. +- Hourly ~12:58 Boise fold (`research/notes.md` §44): Archer still + Watch (no architecture rewrite). Store-index fork: in-engine + ([sqlite-jev](https://github.com/mgaitan/sqlite-jev), pg-jev cousin) + vs CLI rewrite (jevql). Soft judgment inside a hard envelope + ([bitrate-advisor](https://github.com/affirmitv/bitrate-advisor); + mmalisper JOB planner +12% geomean, author-reported; join-order + Choice alone was 2× slower). Distill-to-device as a *memory* gate + (jev-gate, already §33). Encoder vs decoder open-replica receipts + (openjev-lm $0/call overnight CPU). jev-harness as Harbor-adjacent + practice (assert on action). Host adapter + ([jev-routing](https://github.com/nekowasabi/jev-routing), not MCP); + OpenClaw typed routing ([jev-claw](https://github.com/trietphan/jev-claw)). + Voice-control and JevML are README stubs. Higgsfield auto-routing is + a claim. No wrapper. +- kev (`jaredpalmer/kev`, `research/notes.md` §45): runnable Archer + reconstruction on the trained decision-only open path next to Laya / + Nimble / Watch. Qwen2.5-0.5B LoRA + pointer, Apache-2.0, `POST + /v1/systemone` drop-in. Isolation exact (packed vs separate max Δ + 3.7e-6; secret-in-sibling p=0.03 vs in-state 0.99). Held-out ECE + 0.065 (0.031 after temp scale); acc 0.799 on 1,350 ID questions. + Permute argmax flips 7.4%; IIA log-odds shift mean 0.13; boundary + forgery held. Laptop-local System One for development/eval; not a + knowledge/frontier substitute; not a Jev teacher-copy. Contrast vs + TypeAR, encoder DeBERTa, proprietary Jev. jevals/Harbor bake-off + candidate. No serve how-to. +- Hourly ~14:03 Boise fold (`research/notes.md` §46): Archer still + Watch. Open multimodal RLCD + ([blackwood-rlcd](https://huggingface.co/BlackwoodAI/blackwood-rlcd), + CC BY-NC): screenshot + marked candidates → Choice; web acc 0.907 vs + Jev 1.13 text-only 0.480; letter-shuffle 0.133 vs 0.587; ECE 0.037; + ~200 ms H100; Jev still leads general text 0.850 vs 0.786. Shared + bake-off ([open-jev-laya-bench](https://huggingface.co/datasets/pngwn/open-jev-laya-bench)): + 26+9 tasks, 11959 items; ECE/NLL/Brier; macro acc Δ +0.023 + neutral / +0.229 home; LLM-as-judge is not the score. Decision-token + QLoRA + ([Foodoo1/Qwen3-14B-RLCD-Decision-LoRA](https://huggingface.co/Foodoo1/Qwen3-14B-RLCD-Decision-LoRA)): + fraud_risk 64→95%, overall 85.2→98.8% at ~234 ms/4-field broadcast; + synthetic. jevgate frame: allowlist *proves*, Jev judges only + unlisted, fail-open. wellposed: missing `other` → confidence 1.00 + wrong; gating cannot catch it (`tenbin` owns the lint skill). + S1 reflex keeps control (jev-reflex-autonomy-lab). MED: + jev-decision-layer, jev-e2e, jevpandas. No wrapper. +- Abide (`coldteadotai/abide`, `research/notes.md` §47): productized + Jev preference lint for Claude Code / Codex / OpenCode. Soft + AGENTS.md / CLAUDE.md rules → one Score per rule on the diff (never + the conversation); hard rules stay with the linter (same layering + family as jevgate). Edit- vs turn-phase observation window; banded + confidence (≥0.8 repair / 0.5–0.8 note / <0.5 silence — their + operating point) + fail-open hooks; rubric.json quotes source + lines; calibrate/tune fix false positives in the question. Replay + of 93 sessions (1,256 edits / 147 turns) with independent review: + edit precision ~26%, turn ~73% (author-reported, before tune). + Fuller productized path of the jev-pref contract. Complementary to + rh-guard (eval-integrity vs project soft rules). Text/diff only — + not multimodal. No hook how-to. +- kev delta (`research/notes.md` §45): Hub weights + [`jaredpalmer/kev-0.5b`](https://huggingface.co/jaredpalmer/kev-0.5b); + `--run` accepts Hub ids; PEFT `task_type=FEATURE_EXTRACTION` (publish + patches legacy adapters). HIGH question-design: confront Choice + `"other"` / none-of-the-above as a wrong alternative too, vary + wording, dedicated `none_of_the_above` eval (no published rates). + Cross-link wellposed request-shape lint. No species change. No + wrapper. +- Hourly ~14:52 Boise fold (`research/notes.md` §48): Archer still + Watch. Extractive selection + offline `redecide` + ([testimonial-miner](https://github.com/AppitStudio/testimonial-miner)); + pointer-not-generator + ([jev-reviewer](https://github.com/choxos/jev-reviewer)). Local + `/v1/systemone` drop-in ([jev-local](https://github.com/us/jev-local); + default scorer is a stub until `hf`). Observe→decide→verified-act, + no screenshots ([solari-reflex](https://github.com/hitakshiA/solari-reflex); + 60.2/194.9, 66/460, 24.2/98.4 s vs Codex on Solari). Dataframe + accessor sibling ([jevframe](https://github.com/ktaletsk/jevframe); + note jevpandas). Route ≠ memory (jev-hermes). Advisory sidecar + (agent-workflow-typesafe-ai). Structure induction (dag-jev experiment). + Decision-for-control / generator-for-content (jev-agentworld-web-simulator). + Collab arms + Wilson/McNemar (jev-testbench). AST ∩ semantic (jevscan; + `tenbin` owns lint). Light Pi gate (pi-jev-approver). Laya ONNX port + ([laya-onnx](https://huggingface.co/Mattepiu/laya-onnx); do not copy + vs-Jev table). Spotcheck: SemIf 1551★; jevlike 866★; tracker + 20:12:57Z still lists Laya, not Blackwood. No wrapper. +- Hourly ~15:52 Boise fold (`research/notes.md` §49): Archer still + Watch. X discourse blocked. Boundary map / extractable-from-state + ([jev-capability-atlas](https://github.com/Zaious/jev-capability-atlas); + history suite A wrong@0.90 / B 0.07 / C right@0.97; component node; + dangerous-high ECE; DOM-as-text + fan-out). Harbor-style bake-off vs + constrained LLMs + ([DMB](https://github.com/nibzard/decision-model-benchmark) v2: jev + banking 76.3% / spam 93.0% / 256+ cap; p50 264–276 ms; $0.07/1k; no + class wins on quality). Feedstock + ([jevals-data](https://github.com/Jevals/jevals-data) CC-BY-4.0; + recompute-from-logs; 2026-09-18 board). Dual-process S1 decide / S2 + generate ([dual-process-ai](https://github.com/taro1985/dual-process-ai); + routing accuracy unmeasured). Combinatorial ≠ extractive (ARC-AGI + Direct Jev 4/400). Packed one-forward open LLM + ([open-alternative-jev](https://github.com/ikermoel/open-alternative-jev) + RACE-H 92.9% @ 4.55 q/s; not a Jev reproduction). Tiny SAN local + surface ([von](https://github.com/wfzyx/von) 14 MB; not a replica). + kev light delta **100★**. Do not merge Banking77 87% / 76.3% / + 79.67%. No wrapper. +- GLiNER2.5 extractive compaction (`research/notes.md` §50, + [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction), + Apache-2.0): architecture notes, not a plugin how-to. Pointer + keep-drop (character-offset copies) vs generator summarizers; + family with testimonial-miner / jev-reviewer. Soft retention Choice + under a hard mutation envelope (mutating tools / shell operators → + `keep_full`); low-confidence / invalid evidence fail closed to + `keep_full` — contrast many fail-open Jev gates. Same compaction + *job* as fast-jev-compaction / pi-jev-compaction; GLiNER encoder + backend; Fastino/GLiGuard sibling class. `shadowMode` default true. + Not Jev. Not multimodal. No invented metrics. +- CI merge-gate / fail-open wake VOI / S1 indexer / claim-evidence + (`research/notes.md` §51): architecture notes, not a plugin how-to. + [latch](https://github.com/CaseReed/latch) cluster-then-policy + PASS/BLOCK (pair Harbor + rh-guard). + [wakegate](https://github.com/shitianfang/wakegate) skip only if + p(wake)<0.2 (21/21 smoke). s1-graphify-indexer GLiNER extract + + escalate-S2 (10–50× unfilled). + [clear-head](https://github.com/VladyslavHontar/clear-head) + claims vs session evidence. + reification-labs/foreman description-only Phoenix scaffold (not the + super-jev loop). + [jev-gateway-bench](https://github.com/vinilana/jev-gateway-bench) + Harbor on/off one-run signal. + jev-marshal Watch/empty; jevons bounded Pi supervisor (shadow + recovery). MED: if-ai, omp-auto-mode, downloads-sorter, label-desk, + herdr-jev. Archer still Watch. No invented metrics. No wrapper. +- GLiNER2 Ultrafast observe→score→act (`research/notes.md` §52, + [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast), + MIT): architecture notes, not a browser-agent how-to. Same + observe→score-among-candidates→code-acts *job* as jev-ultrafast / + solari-reflex; local GLiNER2 (`fastino/gliner2-multi-v1`) backend, + not GLiNER2.5. No screenshots; no generated selectors; code owns + actuators. Hybrid local decide + remote fill (Mercury 2.5 default + for TYPE). `DONE` ≠ verified success. Contrast blackwood-rlcd + screenshot multimodal; laya-mind2web is DOM-index Laya (same + observed-candidate family). Fastino sibling class with + gliner25-compaction (different hole) and GLiGuard (safety schema). + Demo (theirs, not re-run): Flights 12.20 s / 13.785 s / ~$0.0001 + API — demonstration, not a bake-off. No invented metrics. +- jev-pruner evidence-preserving Bash stdout prune (`research/notes.md` + §53, [jev-pruner](https://github.com/tamaratran/jev-pruner), MIT): + architecture notes, not a plugin how-to. After Bash, Jev Noul-prunes + stdout chunks before the main LLM sees them — no summary. Hard + envelope (≤10k estimated tokens / JSON-diff-whole-doc untouched) + then soft Noul; fail-safe keep original; full archive. Marketplace + id still `fast-jev-output`. Codex is opt-in wrapper, not automatic + interception. Same evidence-preserving *family* as + fast-jev-compaction and gliner25-compaction; different *job* + (command output vs session memory) and Jev backend vs GLiNER2.5. + Manual sweep (theirs): needles 24/24; mean reduction 83% on trim + scenarios. Harbor plugin-eval cannot reach Jev. Terminal-Bench + paired pilot is integration, not a full bench. No invented metrics. +- Cua-S1 specialist System One computer-use (`research/notes.md` §54, + [cua-s1](https://github.com/trycua/cua/tree/main/libs/cua-s1), + parent MIT, ~23.3k★ this pass): architecture notes, not a Driver / + MCP / `uv` how-to. Form-oriented profile `cua-s1-form-v0`. Byte + encoder + option-attention head chooses fill/check/click/skip per + observed element; does not generate values or selectors. Plan ≠ + execute; dry-run default; `execute`/`submit` independent opt-ins; + fail-closed on unknown checkbox / fill without advertised token + `set_value`. **Not TypeSafe Jev** — parallel "System One" naming in + CUA research. Same observe→score-among-candidates→code-acts *job* + as jev-ultrafast / gliner2-ultrafast / solari-reflex / laya-mind2web; + specialist form contract, source-only this pass (no weights, no + checkpoint scores). Offline metric *names* only (accuracy, + abstention, coverage, wrong actions/targets, unsafe when should + abstain). Tests exercise implementation, not checkpoint quality. + Watch for a `cua-s1-form-v0` artifact drop. No invented metrics. +- Hourly ~17:48 Boise fold (`research/notes.md` §55): Archer still + Watch. X MCP flap; `since_id` not advanced. Architecture notes, not + a how-to. Local CUDA/PyTorch Choice/Score/Noul replica + ([jevify](https://github.com/Mintzs/jevify); uncalibrated + likelihoods ≠ Noul; no LICENSE this pass; independent of + Distillation). Decision-native RAG + ([decision-native-rag-skills](https://github.com/emergency-lee/decision-native-rag-skills); + retrieve wide → decide → evidence set; no bundled harness; no + universal benchmark). Verbatim session ledger + scored recall + ([carryforward](https://github.com/Dharundp6/jev-carryforward); + rules never judged; fail-open dump; 9×3 hint). Judgment as a + Ruby language primitive ([hunch](https://github.com/carldaws/hunch); + English-as-config; `rescue nil` fail-open at save). Healthcare + Harbor-shaped S1+S2 + ([explore-typesafe-ai](https://github.com/si618/explore-typesafe-ai); + synthetic FHIR; not clinically validated). Pre-registered + independent eval + ([jev-baselines-eval](https://github.com/ickma2311/jev-baselines-eval); + **both AMBIGUOUS**; cascade sign-flip at exact parity; + confidence=1.0 theater; encoder-with-labels wins; serving-path ≠ + model-speed; same-day errata ×3). Student-b light delta only (HF + card unchanged). MED: toolgate (pre-exec allow/block/review; Jev + not authorization), typesafe-screening-mcp (PubMed screening aid), + databricks-jev-pdf-lab (**honest negative**; no OSS license), + yannip1234/codex-jev (extractive compression family; equal + accuracy/lower cost not established), kazuhideoki/jev-search + (recursive *file* search + fzf; **not** superagents-lab web + search). No wrapper. No invented metrics. +- Hourly ~18:38 Boise 2026-09-18 / 00:38 UTC 2026-09-19 fold + (`research/notes.md` §56): Archer still Watch. Architecture + notes, not a plugin / showcase catalog. Classify-first MCP + ([jev-sift](https://github.com/kbhuw/jev-sift); batch path/url/text + → Jev without entering main agent context first; 50 / 60k / 2MB / + public-IP envelope; mocks ≠ accuracy; no LICENSE this pass; + topology A MCP, not jev-routing). Living applied-mappings atlas + ([jevable.com](https://jevable.com/); claimed 342 vs JSON-LD first + page 36; class patterns — intent columns, score-among-observed, + VOI gates, generative UI decide, robotics text-state, draft-gate + silence ≠ safer — not a 342-title hit list). Maker clocks stay + claims unless already a named receipt. No wrapper. No invented + metrics. +- Stagehand experimental Jev stack (`research/notes.md` §57, + [#2955](https://github.com/browserbase/stagehand/pull/2955) 5/5 of + #2951–#2955, all OPEN draft): architecture notes, not an SDK + how-to. Major harness productization of + observe→score-among-candidates→code-acts (cousins jev-ultrafast / + gliner2-ultrafast / cua-s1 / solari). Jev picks a11y elements; + code copies text. extract `"off"` | `"judge"` | `"pick"`. Their + card (gemini-3.8-flash, 25×3): **37/75** no-LLM ~0.5 s vs baseline + **4.37 s**; 69/75 vs 23/25 (92% both); LLM-off **36/75** — pick is + a fast path, not a replacement. Screenshot extract always LLM. + Cache-check errors never block replay. Do not merge clocks. No + invented metrics. +- Hourly ~18:46 Boise 2026-09-18 / 00:46 UTC 2026-09-19 fold + (`research/notes.md` §58): Archer still Watch. Architecture + notes, not a Convex / uv / pnpm catalog. Public judgment wall + ([ask-jev-ai](https://github.com/waynesutton/ask-jev-ai); 6 + parallel questions; policy-in-code; cost-to-1M from tokens; + license null). Meaning-search without embeddings + ([jevgrep](https://github.com/Bentlybro/jevgrep); 79% top-5 vs + BM25 40% / grep 20% on stripped repos; keyword still wins exact + strings). PR attention ≠ correctness + ([egma-ai/jev-reviewer](https://github.com/egma-ai/jev-reviewer); + **not** choxos pointer-not-generator). Skills→oxlint + ([jev-oxlint](https://github.com/cephalization/jev-oxlint); + AST prove ∩ remainder; Phoenix fixtures; not a hard gate; + `tenbin` owns lint). Session-sticky first-prompt routing + ([jev-adaptive-thinking](https://github.com/jxu-dev-c/jev-adaptive-thinking); + fail-closed fallback). Measured RAG rerank + ([Jev-RAG](https://github.com/Max-sm-yc/Jev-RAG); one-run ≥70% + cost / 72% latency vs Spark *rerank*; full-context Spark still + faster). MED: safe-sh, jev-loan-triage, TurboGuo arenas, jevbox; + hermes/mcp packs not found this pass. No wrapper. No invented + metrics. +- Hourly ~19:48 Boise 2026-09-18 / 01:48 UTC 2026-09-19 fold + (`research/notes.md` §59): Archer still Watch. Architecture + notes, not a pip / venv / Cloudflare catalog. Capability kernel + ([interlock](https://github.com/somoore/interlock); LLM ring 3 / + kernel ring 0; secrets never in the agent; Jev SENSOR; + `policy.py` BLOCK/ASK/ALLOW; type-safe ≠ correct; distinct from + toolgate). Typed control plane around DSPy + ([jev-dspy-control-plane](https://github.com/manikanda-kumar/jev-dspy-control-plane); + DSPy drafts AFTER route+action; OpenJEV / DSPy / JSON Schema + share ontology; offline heuristic ≠ quality). Native-probability + calibration arena + ([jev-arena](https://github.com/meetr1912/jev-arena); live 145 + noul Brier 0.0059 / ECE 0.0620 *theirs*; overconfident in low + bins; 2-request fan-out) plus sonar (heatmap-as-policy) / + vickrey (Jev never bids) / bracket (Brier vs Elo; live trailed + Elo). Engine owns truth / Jev owns judgment + ([game-coach](https://github.com/JoelLewis/game-coach); Wave 0 + PRD; Stockfish WASM; GPL-3.0; anti-soundness-theater with egma). + Human-confirmed port cleanup + ([port-cleanup](https://github.com/epiphany-dynamics/port-cleanup); + Jev recommends; human is the only kill trigger; identity + re-check; shields; mapped explanations). MED toolbelt: + jev-pr-labeler, jevcumber, typedecide, jevon, dsh-jev, + fast-jev-compaction-pi, jev-tetris-benchmark, modelsystem, + opencode-system-one, browser-ai, semantic-bookmark. Skip + jef-mcp (parody) and jevregist (account farming). Star spike: + SemIf 1491→1606 (this pass 1607); jevlike 851→896 (this pass + 897). No wrapper. No invented metrics. +- Hourly ~20:43 Boise 2026-09-18 / 02:43 UTC 2026-09-19 fold + (`research/notes.md` §60): Archer still Watch. Architecture + notes, not a uv / bun / Modal catalog. Domain LoRA specialist + vs few-shot hosted + ([Domain-jev-maker](https://github.com/help-er/Domain-jev-maker); + independent CLINC gold, not a Jev teacher-copy; + matched-precision KL 0.168 vs 0.580 banking; few-shot + determinate McNemar n.s.; train when downstream reads p). + Decide→policy→LLM leftover cascade + ([jav-email-cascade](https://github.com/skiingfalcon/jav-email-cascade); + jev vs gen-json vs gen-logprob; Noul 0.5 never rounded; + license null; mock gen-json flat-confidence is *their mock*). + ORDER BY ranking family + ([jev-orderby-bench](https://github.com/yodablocks/jev-orderby-bench); + six gates pass; Score ordinal 0.143 weak link; 53-way 0.99 + tie; calibration ≠ sortable; recodelabs batch-40 fails + ranking). Wire-compat GLiFormer backend + ([jeff](https://github.com/logan-markewich/jeff); typesafe-sdk + drop-in; ~$2.6 vs $15.6 L4 HTTP ~6×; A10G direct ~$0.65 ~24×; + AG News 75.5% vs 90.5%; CPU more expensive; license null; not + a Jev replica). MED: loopback gateway + ([sysone](https://github.com/hraness/sysone); hosted + local + OpenJev/NanoJev/Mini-Jev; no weights; credential from env). + No wrapper. No invented metrics. +- Hourly ~21:39 Boise 2026-09-18 / 03:39 UTC 2026-09-19 fold + (`research/notes.md` §61): Archer still Watch. Architecture + notes, not a pip / npm / bun catalog. Active-learning triage + ([jev-triage](https://github.com/ThyFriendlyFox/jev-triage); + accept / expensive teacher / human; log full distributions; + **do not distill Jev as teacher of record**, ~68% ceiling). + Evidence-packet explorer + ([jev-semantic-explorer](https://github.com/jimmyhealer/jev-semantic-explorer) + / jevex; index-once ask-many; 1/8→6/8 SWE-bench Verified + finish n=8 *theirs*; packet HitFile 0.233 diagnostic). + Meaning-grep + ([jev-semgrep](https://github.com/uehaj/jev-semgrep); AND/OR/NOT + line Nouls; JP↔EN; MIT LICENSE / GitHub NOASSERTION; 0.94/0.98 + *theirs*). Closed-vote CU + ([JevOnly](https://github.com/buluoray/JevOnly); no planner LLM; + 11 steps / 43 calls / ~$0.014 / 17 s *theirs*). Harbor Jev vs + local MLX PCD vs AR JSON + ([system-one-benchmark](https://github.com/mallahyari/system-one-benchmark); + toxic-chat n=50; Jev 84.0% / Brier 0.1096 vs PCD 52% / 0.3884; + O(1) ≠ calibrated Noul; license null). Host-owned product + ([waymode](https://github.com/mossburgh/waymode); app retains + handlers/permissions; 24/26 + 34/36 *theirs*; not a + self-driving proof). OMP/pi fail-open gates + ([omp-jev-extensions](https://github.com/luw2007/omp-jev-extensions); + `jev_acceptance_gate` + `jev_route`; contrast pi-jev-approver + fail-closed). Skip empty jev-compactor / laya-jolt. No wrapper. + No invented metrics. +- Hourly ~22:38 Boise 2026-09-18 / 04:38 UTC 2026-09-19 fold + (`research/notes.md` §62): Archer still Watch. Architecture + notes, not an `omp plugin` / pip catalog. Watch archive path + missing on this VM; receipts from live GitHub. Permission vs + probability + ([omp-greenlight](https://github.com/SemetricLabs/omp-greenlight); + 1,013 calls / 10 sessions; default **40.9%** prompts removed / + **0 of 94** unsafe auto-approvals on labelled corpus; operator + owns thresholds; plugin never self-tunes; not a sandbox; host + deny stays above). Judgment ≠ permission + ([skill-broker](https://github.com/adamjralph/skill-broker); + Hermes pre-agent outline; code owns grants; Jev never grants + access; **not a production recipe**). Eval integrity / + instrument-not-score + ([dinostomp](https://github.com/collapseindex/dinostomp); + FINDINGS 189 / 99 against itself; `dinostomp jev` if-statement + hygiene; ECE 0.062 *theirs* on 24 examples; beside jevals, not + a Harbor taskset). MED: fast-jev-opencode, jev-desktop, + jev-agent-integration, sift, JevExplore. Census: Awesomejev + 488/21644; SemIf 1641 (+13); jevlike 905 (+4); tracker likes + 41 (+1); Laya yes; Blackwood ABSENT; X MCP flapping + (`pages_archived` 0). No wrapper. No invented metrics. +- Hourly ~23:40 Boise 2026-09-18 / 05:40 UTC 2026-09-19 fold + (`research/notes.md` §63): Archer still Watch. Architecture + notes, not a uvicorn / bun / marketplace catalog. Watch + archive path missing on this VM; receipts from live GitHub + + HF. Hunches labeled. Constrained optimizer + S1 features + ([slo-router](https://github.com/zeeshan8281/slo-router); + license null; Jev task/exactness/evidence as features, never + the sole hot-path gate; fail-open local features; same + routes/accuracy; p95 **77.93 → 490.38 ms** *theirs*; eight-row + demo is not a benchmark). Privilege ≠ verdict + ([construct-auto-classifier](https://github.com/godspede/construct-auto-classifier); + Apache-2.0; effect-based shell gate; fast-allow/deny then Jev + Choice + independent risk Nouls; fail-closed; Jev **0** + dangerous / 975; every chat model leaked 16–104; operator-owned + dials). Attention filter / VOI for human review + ([jev-lens](https://github.com/rashedInt32/jev-lens) + + [jev-lens.nvim](https://github.com/rashedInt32/jev-lens.nvim); + never blocks the agent; never edits; never green unless sure). + MED: [sysone-help/sysone](https://github.com/sysone-help/sysone) + (evaluation-model-first TS SDK; **not** hraness/sysone gateway); + [INSTRUCT_JEV](https://huggingface.co/datasets/ctaxnagomi/INSTRUCT_JEV) + (119 rows; 47/51/21; jevals seed); + [swift-jev](https://github.com/ckaik/swift-jev) (LICENSE-only + this pass; not a CLI product). Census: Awesomejev 488/21644; + SemIf **1652** (+11); tracker likes **42** (+1); lastModified + unchanged; Laya yes; Blackwood ABSENT; X MCP flapping. No + wrapper. No invented metrics. +- Hourly ~00:39 Boise 2026-09-19 / 06:39 UTC fold + (`research/notes.md` §64): Archer still Watch. Architecture + notes, not a uvx / pnpm / marketplace catalog. Watch + archive path missing on this VM; receipts from live GitHub. + Hunches labeled. Measurement owns endorsement + ([jev-packs](https://github.com/dtduc-git/jev-packs); + CC0; nine packs `verified` *theirs* on pinned + `jev-1.13.0`; accuracy/ECE/cost/latency; `unknown` + mandatory; named runner jevassert **not released** / 404; + packs without evidence stay `provisional`). Jev supplies + evidence, code owns authority + ([actiongate-jev](https://github.com/omkarghugarkar007/actiongate-jev); + Apache-2.0; deterministic policy owns ALLOW|REVIEW|BLOCK; + positive score never overrides a hard security fail; + fail-closed financial/destructive/credential if Jev is + down; 500-case is label-baseline, not accuracy). Ranking ≠ + calibration + ([does-jev-confidence-mean-anything](https://github.com/Adilmp/does-jev-confidence-mean-anything) + + [jevcal](https://github.com/Adilmp/jevcal); 8,000 + human-annotated judgments; AUC **~0.91**; stated **~75%** + vs human **~10%**; two-parameter recalibration removes + **~96% ECE** without changing rank; never hard-threshold + raw p as a frequency; vendor "calibrated" often means + rank-correlation). MED: + [gqgs/laya-onnx](https://github.com/gqgs/laya-onnx) + (complete Laya→browser int8; distinct from Mattepiu); + [kunchenguid/local-jev](https://github.com/kunchenguid/local-jev) + (ModernBERT local approximation — not equivalence). Do + not re-fold sysone-help/sysone. Census: Awesomejev + 488/21644; SemIf **1660** (+8); jevlike **910** (+5); + TypeAR 9; tracker likes 42; lastModified unchanged; Laya + yes; Blackwood ABSENT; X MCP flapping. No wrapper. No + invented metrics. +- Same-hour remainder ~00:39 Boise 2026-09-19 (`research/notes.md` + §65): Archer still Watch. Do not re-fold actiongate / jev-packs + / sysone-help. Hot-click CU + ([ego-jev](https://github.com/jiangkoumo/ego-jev); MIT; indexed + viewport table → operation+target; code owns observe/execute/ + `--until`; text model only for type; HN 4.9 s vs 9.7 s / wiki + 5.4 s vs 10.1 s *theirs* n=3, high variance, not a bench). + Jev judges relevance, code decides structure + ([jev-compactor](https://github.com/edwardyen724-g/jev-compactor); + MIT; was empty skip §61; never rewrite; regex floor; compaction + fail-open if Jev down, safety fail-closed; 64.5% / 366 ms / + $0.0004 / 0 invented paths / 4 of 4 facts vs Sonnet summary + 96.2% / 1 invented path, one session). Local rules first, + never auto-train on the model's own hides + ([x-reply-filter](https://github.com/zhuyansen/x-reply-filter); + MIT; `rules.js` then batched Nouls; confirm-queue). OpenCode + port already §62: fast-jev-opencode. Census as §64. No wrapper. + No invented metrics. +- Hourly ~01:47 Boise 2026-09-19 (`research/notes.md` §66): Archer + still Watch. Three clusters: **control-plane combinators** + ([decision-combinators](https://github.com/voidning/decision-combinators); + Then/Gate/Vote/Cascade/Weighted; not literal AND/OR; not chat + turns) + **skill VOI** + ([skillranker](https://github.com/Dicklesworthstone/skillranker); + 52★; two-pass + none-of-these; hook **fail-open** — corrects + §7 fail-closed); **eval integrity without leaderboard theater** + ([jev-capability-atlas](https://github.com/Zaious/jev-capability-atlas) + 10★ receipts, type-safe ≠ correct, axis already §49; + [jev-frontier-100](https://github.com/softpudding/jev-frontier-100) + Jev 77.0% vs Qwen3.5 4B/2048 96.7% / 4B off 56.0%, exploratory; + [jev-ood-calibration](https://github.com/scienthoon/jev-ood-calibration) + 900 tickets ECE 0.107 = 4.4× floor, Choice/Score T~3.3 vs + boolean T 0.66, unknowable priority mean p 0.74); **gate + doctrine clone** + ([turnstile](https://github.com/zyphr-labs/turnstile); Apache-2.0; + policy first, Jev remainder, replay; missing Jev → Review); + **MLX one-pass replica economics** + ([jevmlx](https://github.com/bnsd55/jevmlx); 28★; softmax ≠ + Noul; no local leaderboard yet). MED: invalidate (0 of 157 + false invalidations), jev-intent-review (under construction), + prune-review (22-run cost 1.18% with 305% outlier). Census not + re-derived. No wrapper. No invented metrics. +- Hourly ~02:38 Boise 2026-09-19 (`research/notes.md` §67): Archer + still Watch. Do not re-fold the 01:47 list except sibling + contrast. **TLA+ compose with judgment** + ([jev-labs](https://github.com/copyleftdev/jev-labs); MIT; + never confidently wrong; 1,080 golden 0 wrong *theirs* under + chaos, escalate 5%→18% severe; TLC 1,049,750 states / 0 + errors; synthetic, not clinical). **Advance/coverage ledger** + ([seal](https://github.com/Reasonofmoon/seal); MIT; no seal, + no advance; coverage.path auto|code|human|escalate; mint ≠ + product brain). **skill-broker sibling** (outline already + §62; grants in code vs turnstile runtime vs skillranker + advisory). **Sureness** + ([how-sure-is-jev](https://github.com/adarc8/how-sure-is-jev); + MIT; Choice confidence = max_prob; 75/25 → 0.5 vs entropy + 0.19). **JevBench v1.1** + ([jevbench](https://github.com/fstandhartinger/jevbench); + MIT; unofficial; Main Score 0.6/0.2/0.2; Jev 1.13.0 **87.6**; + calibration reported not scored). **CI typed gate** + ([ci-gatekeeper-bot-jev](https://github.com/NemanjaManic/ci-gatekeeper-bot-jev); + package.json MIT / GitHub SPDX null; 504–629 ms *theirs*). + **Codex MCP adapter** + ([jev-in-codex](https://github.com/teempai/jev-in-codex); + MIT; ranking unbenchmarked; lexical fallback). Census: + SemIf **1683** (+11); jevlike **923** (+5); tracker likes + **43** (+1); Awesomejev 488/21644 unchanged. No wrapper. No + invented metrics. +- Hourly ~03:38 Boise 2026-09-19 (`research/notes.md` §68): Archer + still Watch. Do not re-fold the 02:38 list except sibling + contrast. **Judgment as attention redirect, not a merge + blocker** + ([jev-preflight](https://github.com/muse0509/jev-preflight); + Go MIT; eight risk axes; assist=one reinspect; fail-open; + uncalibrated 0.85; owner-run Claude Code 2.1.267: no-key + fail-open PASS, key-enabled exactly one continuation). + **Landed-script trust / headless≠auto-approve** + ([construct-auto-classifier](https://github.com/godspede/construct-auto-classifier) + delta; cert still Jev **0** dangerous / 975; $0.047/1k). + **Jev judges relevance; code decides structure** + ([jev-compactor](https://github.com/edwardyen724-g/jev-compactor) + product-arm **73%** / 350 ms / 4 of 4 *theirs*; 30–250× + cheaper than shipped summarizers; §65 64.5% is vs-Sonnet). + **Compress-before-first-send** + ([dizk/jev-lens](https://github.com/dizk/jev-lens); MIT; + 79% fewer tokens / 500 SWE-rebench; post-send prune +17% + cost; distinct from rashedInt32/jev-lens). **tools≠use** + ([jev-carryforward](https://github.com/Dharundp6/jev-carryforward) + 0/4 recall; SessionStart > hoping). **Observational + memory** + ([pi-observational-memory-jev](https://github.com/willfish/pi-observational-memory-jev); + keep/kind verbatim; model-free compact). **Independent + open-Jev class** + ([openvons](https://github.com/genai-craft/openvons); + Apache-2.0 LICENSE / GitHub SPDX NOASSERTION; 7★; JevPick + 3.2–4.8×; `/v1/systemone` wire-compat ≠ replica). + **Physical-world S1** + ([HA-Jev](https://github.com/AboveColin/HA-Jev); MIT; + **17★**; sensors from typed answers; not for + locks/heaters). **Judgment outside the store** + ([jevql](https://github.com/kylemclaren/jevql); CLI + judges; vanilla Postgres never sees `jev()`). Short + consumer bullet: [sift](https://github.com/bohutang/sift) + ~$0.00003/post. Census: Awesomejev **561** (+73, agent + tooling 87→107); SemIf **1704**. No wrapper. No invented + metrics. +- Hourly ~04:39 Boise 2026-09-19 (`research/notes.md` §69): Archer + still Watch. Do not re-fold the 03:38 list except sibling + contrast / combinators rename. **Digital-design combinators** + ([jev-combinators](https://github.com/voidning/jev-combinators) + is the rename of decision-combinators; extended Router / + Loop / Retry / Fallback / Memory; metaphor ≠ literal AND/OR). + **VOI cache admission** + ([jevcache](https://github.com/kushals256/jevcache); MIT; + same-intent skip LLM; n=100 *theirs* 0 FP / precision 1 / + recall 0.38 / fpr 0 vs Jaccard@0.35 fpr 0.48; fail-open). + **Harbor skill-routing harness** + ([pi-jev-skill-bench](https://github.com/iamdin/pi-jev-skill-bench) + + [pi-jev-skill-suggestion](https://github.com/iamdin/pi-jev-skill-suggestion); + BM25 vs Jev at roster 50–500; 43 gold; no live numbers this + pass; no-key no-op; tool mode is tools≠use cousin). + **Zeroshot vs BERT displacement** + ([jev-zeroshot-vs-bert](https://github.com/zhuyansen/jev-zeroshot-vs-bert); + +0.05–+0.13 vs DeBERTa-c; contamination 0.901 vs `-c` 0.763; + ≈230 / >2048 labels; DiD 0.035 vs 0.112 *theirs*). + **Typed escalate/continue/abort baton** + ([jev-handoff](https://github.com/shitianfang/jev-handoff); + MIT; inverted loop; gate never grants; fail-open; Vercel + drops confidence). **Worth-your-attention VOI** + ([ThinkyMiner/Winnow](https://github.com/ThinkyMiner/Winnow); + MIT; 80%/90% *theirs*; **≠** kevinpita/winnow). + **Jev WHETHER / Python HOW / LLM WHAT** + ([hermes-jev-router](https://github.com/rsdkrasen/hermes-jev-router); + license null; community plugin; skip-next needs core patch). + **Conflict ≠ ignorance** + ([jev-typed-evaluation-collapse](https://github.com/mleyvaz/jev-typed-evaluation-collapse); + Noul collapses; named Choice p=1.0; binary red 0.67–0.85 + *theirs*). **Playwright executes, Jev chooses** + ([browser-jev](https://github.com/DowLucas/browser-jev); + license null; sample-from-distribution). **Local class** + ([OpenJev](https://github.com/IamBusy/OpenJev) Apache-2.0 + `/v1/decide` 45/60 *theirs*, not TypeSafe drop-in, ≠ + hraness/sysone runners; + [semif-serve](https://github.com/dddanielliu/semif-serve) + 1164 vs 178 ms; runoff ≠ softmax; wire-compat ≠ replica). + Toolbelt notes: jev-security-scan / jev-decisions / TeoMastro + (summary.md 404 this pass); **rh-guard owns reward-hack**. + Flywheel: + [DGUI_HYPERMEM-JEV](https://huggingface.co/datasets/ctaxnagomi/DGUI_HYPERMEM-JEV) + 6-row schema. Census: Awesomejev **flat 561/27007**; tracker + likes **43→45**; SemIf **1714** (+10); jevlike **926** (+3). + No wrapper. No invented metrics. +- Hourly ~05:46 Boise 2026-09-19 (`research/notes.md` §70): Archer + still Watch. Do not re-fold §50–§69 HIGH except sibling + contrast / jevassert landing / prune-review, intent-review, + laya-jolt, local-jev deltas. **Record/replay CI LANDED** + ([jevassert](https://github.com/dtduc-git/jevassert); + Apache-2.0; accuracy/ECE/Brier/cost/latency offline from + recordings; exit 0/1/2; McNemar; Action `@v0`). + **Evidence-gated packs now have a runner** + ([jev-packs](https://github.com/dtduc-git/jev-packs); + size 0→458; 2,990-case matrix *theirs*: Jev/Sonnet 5 + accuracy tie Δ≤0.018, Jev better calibrated 7/9, ~250× + cheaper; sms-spam this-pass 0.953/ECE 0.040). + **Failure-finding arena** + ([jevarena](https://github.com/chenmingtang830/jevarena); + Apache-2.0; **≠** meetr1912/jev-arena; harness not findings). + **BBQ stereotype/uncertainty/cost** + ([jev-bbq-experiment](https://github.com/simonmesmith/jev-bbq-experiment); + license null; 58,492; 97.28%; bias 0.04/0.34; $0.3429 / + 7.75 min *theirs*; not a general bias cert). + **Decider ≠ executor** + ([jeffrey](https://github.com/thomasbrueggemann/jeffrey); + MIT; Jev next-tool/progress/risk/done; LLM fills args; + pick ≠ fill). **Sentence-as-rule lint** + ([jevlint](https://github.com/mizchi/jevlint); MIT; + ast-grep × `ask:`; 13/15 1.00/1.00 *theirs*; **≠** + huntedman/JevLint). **VOI hunk prune** + ([prune-review](https://github.com/shubhangi013/prune-review); + 22-run 1.18% with 305% outlier; ~20% target; cost not + quality). **Whole-repo intent** + ([jev-intent-review](https://github.com/yottayoshida/jev-intent-review); + VERIFIED/VIOLATION/UNKNOWN; empty search ≠ proof). + **GLiNER2 System One spec** + ([Jev_from_GLiNER2](https://github.com/Eran-BA/Jev_from_GLiNER2); + spec-only; ≠ jeff). **Open replica substrates** + ([grande](https://github.com/bokuweb/grande) JGLUE 0.614/ + 0.853 + 270M 0.710/0.710 *theirs*; + [laya-jolt](https://github.com/jlt-commons/laya-jolt) + byte parity; [JEV-CPU](https://github.com/leesk212/JEV-CPU) + PoC, Meanblock 404; [local-jev](https://github.com/kunchenguid/local-jev) + done 30%/shape 57%). **Persist constraints** + ([pi-heed](https://github.com/Nyarlathoteppppp/pi-heed); + 98.5%/0 false block *theirs*). Toolbelt note: + actiongate slogan already §64. MED: + [system-one-responsible-ai](https://github.com/david-j-lustig/system-one-responsible-ai) + size-0 framing stub. Census not re-derived. No wrapper. + No invented metrics. +- Hourly ~06:43 Boise 2026-09-19 (`research/notes.md` §71): Archer + still Watch. Do not re-fold §50–§70 HIGH except sibling + contrast. **Harbor SGR-judge contract** + ([jev-judge-bench](https://github.com/slavadubrov/jev-judge-bench); + README MIT / GitHub SPDX NOASSERTION; frozen SLA-150; Jev vs + Luna / DeepSeek-flash / glm-5.3-flash; invalid = FN; + 21 offline tests; canaries ≠ quality; **no quality headline + yet**; **≠** jevarena / jevbench). **Empty skip** + ([jev-context-pruner](https://github.com/IPECTER/jev-context-pruner); + 409 empty). **Hand no-text steps** + ([jev-use](https://github.com/shitianfang/jev-use); MIT + v0.4.1; p50 220 ms; 186 vs 2,672 ms; gate 12/12; Vercel + drops confidence → margin 0.4; first loop 17/20 then 0/20 + *theirs*; **≠** jev-ultrafast). **Pi System-One control + plane** ([pi-jev-control](https://github.com/goodruizhan/pi-jev-control); + license null; v0.3.0 private; GUI never force-click; + compaction never writes session). **Never free-generates** + ([jev-gpt](https://github.com/florian-hoenicke/jev-gpt); + license null; ~400 calls / 75 s / 2¢ *theirs*). + **OpenRouter recipe atlas** + ([jev-cookbook](https://github.com/nexibeo/jev-cookbook); + MIT; 1★; 16–36 samples not benches; 425 calls / $0.015; + browser 5/6 *theirs*). **Personal-history feed** + ([jevfeed](https://github.com/fengyiqicoder/jevfeed); MIT; + no social graph; one request per batch of ten). + **Competing NAR claim-audit, not endorsement** + ([openJev-verdict-2.0](https://github.com/Heman10x-NGU/openJev-verdict-2.0); + README Apache-2.0 / GitHub SPDX NOASSERTION; 77.10%/0.0636/ + 0.0144 *theirs* unverified; **open PR #1**: throughput≠ + latency, Laya parity, like-for-like ECE; **≠** + IamBusy/OpenJev). Census not re-derived. No wrapper. No + invented metrics. +- Hourly ~07:49 Boise 2026-09-19 (`research/notes.md` §72): Archer + still Watch. Do not re-fold §50–§71 HIGH except sibling + contrast. **1-token logprob endpoint ≠ Noul** + ([chakuho](https://github.com/taku-me/chakuho); MIT; + coverage ≠ correctness; GUI 336 *theirs* 27B 95%/92% + vs Jev 89%/82%; `__none__` 97% vs 8B 10%). **Open + replica engine** ([jevinf](https://github.com/zerodegress/jevinf); + MIT; 2.57×/2.27× 100% argmax; MPS only). **Unofficial + Elixir SDK ≠ OTP peer** + ([typesafe-elixir-sdk](https://github.com/phiat/typesafe-elixir-sdk); + MIT; 1★; ≠ dannote/jev). **jevex rename + n=16 VOI** + ([jevex](https://github.com/jimmyhealer/jevex); 160s→69s + / $8.74→$3.13 / 16/16 *theirs*; keep n=8 1/8→6/8). + **Commit attention≠verdict** + ([commitjev](https://github.com/yodablocks/commitjev); + MIT; middle band never rounded; 0 false on 5 clean + *theirs*). **Hermes plugin is Agnes not TypeSafe** + ([hermes-plugin-jev](https://github.com/Mrmimee/hermes-plugin-jev)). + **Pi compact ≠ compaction** + ([pi-jev-compact](https://github.com/dev-willbird1936/pi-jev-compact); + MIT). **Empty skip** + ([jev-runway](https://github.com/IPECTER/jev-runway); + LICENSE-only). **Decision-native inbox** + ([mailordinal](https://github.com/Milo318/mailordinal); + MIT). **Unofficial jev-cli not ready** + ([jev-cli](https://github.com/shaharia-lab/jev-cli); + 0.0.0; ≠ jevql). **Laya multilingual** + ([laya-multilingual](https://huggingface.co/convaiinnovations/laya-multilingual); + MASSIVE 0.366/0.387; Khmer 0.000@0.952; ships + uncalibrated). **Schema-scorer Hub** (GitHub 404; v2 + Choice 0.841; peaked ranking). HF 401 this pass on + open-jev-laya-bench / jev-tree-choice-cap / + INSTRUCT_JEV; jevlogs 404+401. Census not re-derived. + No wrapper. No invented metrics. +- User-provided signal ~08:37 Boise 2026-09-19 + (`research/notes.md` §73): **Skip Archer.** + **Productized System One HTTP** + ([classifier-dev](https://github.com/mrmps/classifier-dev); + MIT; **185★**; https://classifier.dev). Label + + calibrated confidence as the public contract; batch + `{id,text}[]` ~1000; Jev primary, LLM fallback only. + 400 headlines **650 ms** *theirs*. **Escalate-under- + threshold:** smart re-asks single-label <0.7; + multi-label ignores (re-judge worse, 23 s). Emotion + ≥0.9 → 82% / <0.5 → 29%; gemini-3.8-flash 87.5→90.0 / + 61.8→63.7 *theirs*. Multi-label F1 **0.887** / **230 ms** + vs cascade **0.799** / 1.5 s (eval 232 ms; AG News + **87.7%** vs 82.0%). **Measurement-first:** `/benchmark` + from tracked JSON; read eval/README (n=7 train-on-test; + ~0.03 coin flip). **Silent FALLBACK:** granite F1 + **0.546** vs advertised ~**0.800** *theirs*; rh-guard + owns the gate. Life/business (spam/inbox/feedback), + not SWE-only. Distinct from ask-jev-ai wall. No wrapper. + No invented metrics. +- User-provided signal ~08:48 Boise 2026-09-19 + (`research/notes.md` §74): **Skip Archer.** **Delta of + §48.** Pointer-not-generator at evidence-synthesis + scale + ([choxos/jev-reviewer](https://github.com/choxos/jev-reviewer); + MIT; **12★**; https://jevreviewer.xera.ac). **≠** + [egma-ai/jev-reviewer](https://github.com/egma-ai/jev-reviewer). + Two-pass Choice (which line) + Noul (does this line + itself answer); quotes = Noul ≥ 0.5 *theirs*. *Not + found* / *Unclear* first-class. Human tick is the + product (checked never overwritten). 18-q template + **4.6 s / $0.0101** *theirs* (spot check, not a + validation study). Cochrane / PRISMA / RoB, not + SWE-only. No wrapper. No invented metrics. +- User-provided signal ~08:56 Boise 2026-09-19 + (`research/notes.md` §75): **Skip Archer.** + **Wire-compat ≠ logit-equiv** + ([githubnext/localjev](https://github.com/githubnext/localjev); + MIT; **261★**; GitHub Next). **≠** + [kunchenguid/local-jev](https://github.com/kunchenguid/local-jev). + Bun `POST /v1/systemone` on DiffusionGemma via Chat + Completions; TypeSafe SDK drop-in. Prompted JSON → + validate/retry → normalize + entropy confidence — not + razorback16 structured-read logits. Harbor-shaped + bake-off *theirs*: 1,200 req; Qwen3.6 short macro + **76.7%**; Gemma 4 26B-A4B **75.0%**; DiffusionGemma + **74.2%**; no definitive winner (2/120); do not treat + as calibrated. LM Studio still cannot load + DiffusionGemma. Do not copy bun / `.env`. No wrapper. + No invented metrics. +- User-provided signal ~09:07 Boise 2026-09-19 + (`research/notes.md` §76): **Skip Archer.** **Laya + packaging, not a new species** + ([NandhaKishorM/laya](https://github.com/NandhaKishorM/laya); + Apache-2.0; **710★**). PyPI + `Router` over Hub + [`laya`](https://huggingface.co/convaiinnovations/laya) / + [`laya-multilingual`](https://huggingface.co/convaiinnovations/laya-multilingual) / + [`laya-typed-decisions`](https://huggingface.co/convaiinnovations/laya-typed-decisions). + **≠** TypeSafe `/v1/systemone`. **≠** githubnext/localjev. + T4 *theirs*: 1q **32.8 ms** (~7.8× vs Jev p50 + 236–276 ms). Post-T ECE **0.081** vs Jev **0.246**; + raw ECE still trails (0.213 vs 0.144). Banking77 + **0.425** vs Jev **0.870** (77 vs 72; ~3–4 tok/label). + typed-decisions **0.766** is a fine-tune (base + 0.362/0.342 vs majority 0.461). Soft-acc 0.471 vs + 0.580. Khmer **0.000@0.952** — Router because gating + cannot catch. 0.85 still soft. Jev rows third-party + unpublished-here. Do not copy pip / preload. No + wrapper. No invented metrics. +- User-provided signal ~09:14 Boise 2026-09-19 + (`research/notes.md` §77): **Skip Archer.** **External + openjev census ≠ scored bake-off** + ([@airesearch12](https://x.com/airesearch12/status/2101259522933186879); + Florian S / Benchmark Heaven). Named ~18 (system-one-open, + openjev-sglang, DeBERTa open-jev, Needle 3, + open-alternative-jev, Nimble 9B, SemIf, open-jev Dasein / + JoshuaSP, OpenJev razorback16, mini-jev, system-one, + system-one-gemma, jevlike, AlexWortega/openjev, GLiNER2, + Succinct Router 14M, jev-model-router/Director/Loki). + GLiNER2 + routers are **class-boundary**. Incomplete vs + Laya / localjev / kev / TypeAR / openvons. Engagement + **ephemeral** (SIGNAL ~417/9/3; this pass 564/15/5). **≠** + jevbench v1.1. Watch + [jev-models](https://benchmarkheaven.com/jev-models); do + not paste live ranks. Do not copy Stripe. No wrapper. No + invented metrics. +- User-provided signal ~09:24 Boise 2026-09-19 + (`research/notes.md` §78): **Skip Archer.** **JevBench v1.2 + scored board** + ([benchmarkheaven.com/jev-models](https://benchmarkheaven.com/jev-models); + harness [fstandhartinger/jevbench](https://github.com/fstandhartinger/jevbench) + MIT; 0★; HEAD `27ed3d6c`). Protocol `jevbench::v1.2`; scored + 19 Sept 2026; 534 decisions (hard 220 = 30% of Intelligence). + Official Score = geometric mean of I/C/S/K at 25% each. Jev + 1.13.0 **75.3**; SemIf (Qwen3.5-4B) **74.6** (−0.7); OpenJev + DiffusionGemma (razorback16) 67.6 *theirs*. Luna Intelligence + **96.8** rank **#7** on cost. Calibration **on** the rank + (delta from v1.1). Weighting is a product design. Option-order + 72%→21%. Self-host latency ×2 is an assumption; many costs + est. Laya absent (gap, not named-excluded). GLiNER2 mapping + issues; apps out. Qwen3.8 27B Chutes TEE **≠** Archer. **≠** + tweet census §77 **≠** v1.1 87.6. Do not copy Stripe / CLI. + No wrapper. No invented metrics. +- Hourly System One watch ~08:42 Boise 2026-09-19 + (`research/notes.md` §79): **Skip Archer.** Named HIGHs + **already folded** (§73–§78) — extract **how-to-apply**, + not a hit list: wire-compat ≠ logit-equiv (prompted JSON + ≠ structured logit); productize label+p and mark + `FALLBACK`; packaging ≠ new species / script-before-p / + 0.85 still soft; pointer-not-generator (two-pass; *Not + found*; human tick); external census ≠ scored bake-off / + geo-mean weights are a design. **Skip thin noise** + (JEValuate / jevspeak / fable-jev; jev-semgrep already + §61). Hard-gating a Noul as a PR/quality gate is + soundness theater + ([totally-tim/jev-gate](https://github.com/totally-tim/jev-gate) + 0★ ≠ jev-gateway; [claude-jev-warden](https://github.com/connectedGraph/claude-jev-warden) + 1★). Qwen3.8 27B ≠ Archer. No wrapper. No invented + metrics. +- User-provided HIGH ~09:50 Boise 2026-09-19 + (`research/notes.md` §80): **Skip Archer.** Delta of + §46, not a new species. + [khordoo/jev-reflex-autonomy-lab](https://github.com/khordoo/jev-reflex-autonomy-lab) + (TypeScript; **7★**; license null; README SHA + `130987c9`; ARCHITECTURE SHA `48da0769`; HEAD + `e3297ebe`). S1 never stalls waiting; S2 is one-use + advisory and never flies. Purple confidence = + **consumed** S2 (purple bar = arrival; red = fail). + Local controller is rule-based **≠** + githubnext/localjev **≠** kunchenguid/local-jev. + Live API `POST /v1/systemone` `jev-latest`; 20% + starting gate *theirs* still soft and does not start + a mission. Seed = geometry ≠ async replay. No pixels + to either provider; confidence ≠ selected + probability; S2 never grants. README GLM 5.3 vs + ARCHITECTURE muse-spark-1.3-contributor — quote both + *theirs*. Experimental viz, not a flight controller. + Do not copy npm / `.dev.vars`. No wrapper. No + invented metrics. +- User-provided HIGH ~09:51 Boise 2026-09-19 + (`research/notes.md` §81): **Skip Archer.** Productized + observe→score-among-candidates→code-acts on a Mac, + not a new species. + [awlevin/typesafe-computer-use](https://github.com/awlevin/typesafe-computer-use) + (Python; MIT; **427★**; README SHA `369f4a6a`; HEAD + `cc7b5066`). OCR+AX → numbered items → TypeSafe + Choices (`kind`/`item`/`site`/`offscreen`) → + deterministic click/type. **Never ships a screenshot + to frontier for the *decision***; the one-shot + **answer** writer may receive the capture (reader + packet, not the Choice). Writer only for free text. + Overlapping options = false low confidence. AX bonus + never sole (Spotify 0 *theirs*). Post-type Noul 0.5 + and `--min-confidence` 0.4 still soft. $0.0002 vs + Opus $0.032 (155×) *theirs* on **one screenshot**, + not a Harbor taskset. Honest caveat: dates.py rebuilds + pixel-free reasoning. **≠** jev-ultrafast **≠** + cua-s1 **≠** jev-macos-loop **≠** camoufox. Do not + copy `uv sync` / `.env`. No wrapper. No invented + metrics. +- User-provided HIGH ~10:01 Boise 2026-09-19 + (`research/notes.md` §82): **Skip Archer.** Productized + ASR observe→score-among-candidates→code-acts in + headed Chromium, not a new species and not omni. + [moritzkremb/jev-voice-browser](https://github.com/moritzkremb/jev-voice-browser) + (JavaScript; MIT; **103★**; README SHA `fa033303`; + HEAD `054db0f3`). Web Speech partials → one 9–11- + question Jev request (~250–350 ms *theirs*) → + policy. Pointer-not-generator for spans. Closed-set + may act on a partial; free-text waits. Spoken + confirm is convenience, not auth. Numbered overlay, + no second model. Integration 27/27 / ~$0.0002/call + *theirs* fixtures, not a Harbor taskset. 0.5 / 0.55 + / 0.6 still soft. **≠** jev-voice-control **≠** + nikolas-j **≠** Aj1905 **≠** typesafe-computer-use + OCR. Do not copy `npm` / `.env` / `run.sh`. No + wrapper. No invented metrics. Expand the §39 tweet; + do not re-card it. +- User-provided HIGH ~10:20 Boise 2026-09-19 + (`research/notes.md` §83–§84): **Skip Archer.** Two + signals, one fold. + [reddpy/AgentGhost](https://github.com/reddpy/AgentGhost) + (TypeScript; MIT; **2★**; README SHA `44145fa9`; + HEAD `ac04e4fb`). Intent-aware ALLOW/ASK/DENY + wrap-as-execution: the wrap *is* the tool function; + rules first; ASK throws; `failMode: closed`. Judge + is a slot. Provider-hosted tools out of reach. + **≠** jwen5419807/agentghost **≠** vventirozos + **≠** actiongate **≠** toolgate **≠** jev-use. + rh-guard owns the gate cousin. Do not copy `npm` / + `.env` / `AUTO_APPROVE`. [@studio_yebisu JP genre + atlas](https://x.com/studio_yebisu/status/2101065176069886152) + (2026-09-18T21:45:48Z). Apps by hole, not a scored + bake-off. Stars research-time (typesafe-computer-use + 203→**427**; jev-voice-browser 40→**103**). Not + verified evals. Engagement ephemeral (this pass + 131,234 / 1,934 / 192). SAM 3.1 already §39. + OpenRouter Jev no-waitlist is WATCH, not a recipe. + **≠** @airesearch12 class census **≠** v1.2 board. + Do not dump the 30 repos. No wrapper. No invented + metrics. +- User-provided HIGH ~10:25 Boise 2026-09-19 + (`research/notes.md` §85): **Skip Archer.** External + pedagogy / how-to-apply, not a new species. + [@akshay_pachaar “Jev Clearly Explained”](https://x.com/akshay_pachaar/status/2101037514945597645) + (article https://x.com/i/article/2100940576741093376; + 2026-09-18T19:55:53Z). LLM hammer for bounded + decisions; code owns branches; parallel questions; + thresholds in code; **schema-safe ≠ correct**; + placements = routing / tool-risk / verify with LLM; + shadow-mode; questions-as-code. **200× / 400×** and + 70–500 ms / $0.042/MTok are TypeSafe **ceiling** + claims *theirs*, not Harbor. Text-only; not looking + at the screen. **≠** official docs **≠** Flavio + Copes **≠** LangChain harness **≠** AgentGhost. + Engagement ephemeral (this pass 233,495 / 2,280 / + 235). Do not copy the Python samples. No wrapper. + No invented metrics. +- User-provided HIGH ~10:30 Boise 2026-09-19 + (`research/notes.md` §86): **Skip Archer.** Dedicated + fold of [uehaj/jev-semgrep](https://github.com/uehaj/jev-semgrep) + (light-noted §61). Grep by meaning via Jev Noul; + proposition ≠ embedding; contrast-set (all six + about a refund; only customer-*asking* pass); + AND/OR/NOT are boolean ops on *thresholded* bits + (do not multiply p; ≠ jev-combinators metaphor). + Cross-lingual; no index; EN safer near threshold. + Semgrep.dev SAST name collision. **Not a gate** + (ranking fail-open; rh-guard skip). LICENSE MIT / + GitHub NOASSERTION. HEAD `21120e9`; README SHA + `923e6a5`. Stars ephemeral (0 → SIGNAL ★42 → **51** + this pass). 0.94/0.98 LLM-as-judge 10×51 *theirs*, + not Harbor. **≠** jevgrep **≠** jev-sift **≠** + jevex **≠** semgrep.dev. Do not copy npm / `npx` / + `.env` / marketplace. No wrapper. No invented + metrics. - Effect-oriented loops (`notes.md` §28, `mappings.md` §19): Ward's ZIO client keeps Jev as the outer Choice and the handler as the effect. Not Effect.ts. GLiNER author: GLiNER2 "like jev" is GLiGuard diff --git a/README.md b/README.md index 47d4592..9a1e43f 100644 --- a/README.md +++ b/README.md @@ -27,12 +27,24 @@ never launder a Noul as a proof. - `.agents/skills/augustus/SKILL.md` — working protocol + decision-design card - `.agents/skills/augustus/references/mental-models.md` — cross-domain frames (EU, abstention, VOI, MCDA, SDT, search/control, Leveson, - NATM/snap-fit/Norman); not SWE-only + NATM/snap-fit/Norman); not SWE-only. Extractable-from-state boundary + map (self-contained vs needs outside knowledge) - `.agents/skills/augustus/references/judgment-class.md` — the class (Jev - exemplar, not monopoly): open heads (Laya, encoder DeBERTa, LoRA - distill), constrained-AR (TypeAR, pcdServer), announced decision-model (Watch), - GLiNER/GLiClass species (locate vs categorize vs local multi-head), - listwise vs decision objectives, vision scoring, when-to-use axes, + exemplar, not monopoly): open heads (Laya, kev, encoder DeBERTa, LoRA + distill, domain specialist on independent gold), constrained-AR (TypeAR, pcdServer), announced decision-model (Watch), + open multimodal RLCD (blackwood-rlcd; not Archer), Laya ONNX port, + contract-compatible local `/v1/systemone` (stub until hf scorer; also kev pointer / von tiny SAN — not replicas; **jevify** CUDA/PyTorch packed-logprob cousin — uncalibrated likelihoods ≠ Noul; **jeff** GLiFormer-400M encoder drop-in — not a Jev replica; **sysone** loopback gateway routes hosted + local, not a model; **githubnext/localjev** prompted JSON ≠ structured-read logits — ≠ kunchenguid/local-jev), + GLiNER/GLiClass species (locate vs categorize vs local multi-head; + GLiNER2.5 extractive compaction as a named job, not a new species; + GLiNER code-graph indexer + escalate-S2, 10–50× unfilled; + GLiNER2 observe→score-among-candidates computer-use as a *different* + named job, not GLiNER2.5; typesafe-computer-use OCR+AX desktop hosted Jev + of the same hole, never screenshot-to-frontier for the decision), + **OpenJev** local `/v1/decide` (not TypeSafe drop-in; distinct from + hraness/sysone OpenJev runners), **semif-serve** SemIf `/v1/systemone` + runoff (wire-compat ≠ replica), + listwise vs decision objectives, vision scoring, when-to-use axes + (including decision-model vs constrained LLM), agent-architecture portents - `.agents/skills/augustus/references/formal-methods.md` — judgment vs proof ownership; Alloy Analyzer vs Apalache (finder ≠ BMC ≠ @@ -45,18 +57,226 @@ never launder a Noul as a proof. alias of the FM pillar - `.agents/skills/augustus/references/mixed-architecture.md` — default placement: judgment-class model + LLM + code; preference lint; provider - (Jev default / other family with self-eval) + (Jev default / other family with self-eval); dual-process S1 decide / S2 + generate; component node; DOM-as-text + fan-out; shadow-mode compaction rollout; + fail-open wake vs fail-closed merge-gate; Harbor on/off routing; + hybrid local decide + remote fill; `DONE` ≠ verified success; + evidence-preserving stdout prune (hard envelope then Noul); + specialist S1 computer-use (Cua-S1 form-v0; plan ≠ execute; not TypeSafe Jev); + judgment as a language primitive (hunch); decision-native RAG + (retrieve wide → decide → evidence set); classify-first MCP + (jev-sift); draft-gate heartbeat; living class-pattern atlas; + public judgment wall; PR attention ≠ correctness; session-sticky + first-prompt route; capability kernel (secrets never in agent; + Jev SENSOR); typed control plane around DSPy; engine owns truth / + Jev owns judgment; human-confirmed kill; decide→policy→LLM leftover + cascade; wire-compat encoder backend; loopback gateway; + closed-vote computer-use (no planner LLM); host-owned handlers × + System One; active-learning triage (do not distill Jev as teacher); + evidence-packet explorer; meaning-grep AND/OR/NOT; OMP prompt + suppression (permission vs probability; operator owns the bar); + judgment ≠ permission (skill-broker outline, not a recipe); + constrained optimizer + S1 features (slo-router; never sole + hot-path gate); effect-based shell gate (privilege ≠ verdict); + attention filter / VOI (jev-lens; never blocks; never green unless sure); + measurement owns endorsement (jev-packs evidence-gated + jevassert record/replay CI); + Jev supplies evidence / code owns authority (actiongate-jev); + ranking ≠ calibration (never hard-threshold raw p as frequency); + hot-click CU (ego-jev; indexed table; S1 on click path); + Jev judges relevance / code decides structure (jev-compactor); + local rules first / never auto-train on own hides (x-reply-filter); + control-plane combinators (not chat turns); skill VOI / abstention + (skillranker hook fail-open); receipts not leaderboard (atlas); + OOD / AUC ≠ ECE (sign flips by type); thinking-budget bake-off + (frontier-100); turnstile evidence≠authority + replay; MLX + one-pass replica economics (jevmlx; softmax ≠ Noul); + never confidently wrong / TLA+ compose (jev-labs); + no seal no advance / coverage ledger (seal; mint ≠ product + brain); sureness bands (how-sure-is-jev; max_prob is generous); + JevBench Harbor practice (calibration not in Main Score); + CI typed gate before expensive review (ci-gatekeeper); + Codex MCP host adapter (jev-in-codex); + judgment as attention redirect (jev-preflight; not a merge blocker); + compress-before-first-send (dizk/jev-lens; 79% fewer tokens); + tools≠use / SessionStart over hoping (carryforward 0/4); + observational memory (pi-om keep/kind verbatim); + open-Jev class (openvons; JevPick; wire-compat ≠ replica); + physical-world S1 (HA-Jev; not for locks); + judgment outside the store (jevql CLI); + landed-script trust / headless≠auto-approve (construct); + digital-design combinators (jev-combinators rename + extended five); + VOI cache admission (jevcache 0 FP/100); + worth-your-attention VOI (ThinkyMiner/Winnow ≠ kevinpita/winnow); + Jev WHETHER / Python HOW / LLM WHAT (hermes-jev-router); + typed escalate/continue/abort baton (jev-handoff; gate never grants); + Playwright executes, Jev chooses (browser-jev); + OpenJev `/v1/decide` ≠ drop-in + SemIf runoff wire; + conflict ≠ ignorance (named Choice escape); + decision-as-memory flywheel (DGUI_HYPERMEM-JEV); + record/replay CI (jevassert LANDED); + measurement owns endorsement now has a runner (jev-packs + 2,990-case matrix; calibration+cost first-class); + failure-finding arena (jevarena ≠ jev-arena); + BBQ stereotype/uncertainty/cost (not a bias cert); + decider≠executor (jeffrey; pick ≠ fill); + sentence-as-rule lint (mizchi/jevlint ≠ huntedman/JevLint); + VOI hunk prune (prune-review ~20% cost target); + whole-repo intent VERIFIED/VIOLATION/UNKNOWN; + GLiNER2 System One spec ≠ replica; + open replica substrates (grande / laya-jolt / JEV-CPU / + local-jev measured not equivalent); + githubnext/localjev prompted JSON ≠ structured-read logits; + persist constraints across compaction (pi-heed); + Harbor SGR-judge contract (jev-judge-bench; canaries ≠ quality; + no headline yet; ≠ jevarena/jevbench); + hand no-text steps (jev-use; Vercel drops confidence); + Pi System-One control plane (pi-jev-control; GUI never force-click); + never free-generates (jev-gpt tree of Choices); + OpenRouter recipe atlas (jev-cookbook; samples not benches); + personal-history feed (jevfeed; no social graph); + competing NAR claim-audit (openJev-verdict-2.0; PR #1; ≠ OpenJev); + empty compaction-proxy skip (IPECTER); + 1-token logprob endpoint ≠ Noul (chakuho; coverage ≠ correctness); + open replica engine (jevinf; argmax-parity ≠ ECE); + unofficial Elixir SDK ≠ OTP peer; + jevex n=16 files-to-read VOI; + commit pre-review attention≠verdict (middle band); + Hermes plugin is Agnes not TypeSafe; + pi-jev-compact ≠ pi-jev-compaction; + decision-native inbox (mailordinal); + unofficial jev-cli not ready (≠ jevql); + laya-multilingual English checkpoint confident-wrong OOD; + schema-scorer peaked ranking ≠ calibration; + productized System One HTTP (classifier.dev; label+confidence; batch ~1000); + escalate-under-threshold (smart single-label <0.7; multi-label ignores); + silent FALLBACK (granite 0.546 vs advertised 0.800; rh-guard owns the gate); + systematic-review pointer (choxos/jev-reviewer ≠ egma-ai; two-pass Choice+Noul; human tick is the product); + wire-compat ≠ logit-equiv (githubnext/localjev ≠ kunchenguid/local-jev; prompted JSON ≠ structured-read logits); + Laya packaging ≠ new species (NandhaKishorM/laya; Router script-before-p; post-T ECE ≠ raw ECE; 0.85 still soft); + external openjev census ≠ scored bake-off (@airesearch12; GLiNER2+routers class-boundary; incomplete vs watch; Harbor honesty watch); + JevBench v1.2 geometric-mean I/C/S/K (cal ON rank; Jev 75.3 / SemIf 74.6 *theirs*; Luna I=96.8 rank #7; option-order 72→21; ×2/est. Harbor honesty; Laya absent gap; Qwen3.8 27B ≠ Archer); + hourly 0842 already-folded recipe (wire≠logit · FALLBACK · packaging honesty · pointer-not-generator · leaderboard VOI; skip thin noise; hard-gate Noul as PR/quality = soundness theater); + S1 never stalls / S2 one-use advisory (khordoo/jev-reflex-autonomy-lab delta; purple = consumed; Local controller ≠ githubnext/localjev; seed = geometry; 20% still soft; no pixels; S2 never grants); + OCR+AX desktop observe→score→act (typesafe-computer-use; never screenshot-to-frontier for the decision; overlapping options = doubt; 155× *theirs* one screenshot; 0.4/0.5 still soft; **≠** jev-ultrafast **≠** cua-s1); + ASR voice-browser observe→score→act (jev-voice-browser; never waveform-to-Jev; partial-speech VOI; spoken confirm ≠ auth; numbered overlay; 27/27 *theirs* fixtures; **≠** jev-voice-control **≠** nikolas-j **≠** OCR desktop); + wrap-as-execution ALLOW/ASK/DENY (AgentGhost; wrap *is* the tool function; rules first; ASK throws; fail-closed; **≠** actiongate **≠** jev-use; rh-guard owns the gate cousin); + JP genre atlas (@studio_yebisu; stars research-time ≠ eval; **≠** class census **≠** v1.2); + external pedagogy / how-to-apply (@akshay_pachaar “Jev Clearly Explained”; LLM hammer; schema-safe ≠ correct; 200×/400× TypeSafe ceiling; shadow + questions-as-code; **≠** official docs **≠** Flavio); + meaning-grep dedicated (jev-semgrep; proposition ≠ embedding; AND/OR/NOT after threshold; Semgrep.dev collision; not a gate; contrast-set refund) - `.agents/skills/augustus/references/applied-mappings.md` — context sieve, - exact-text keep/drop, env triage, moderation/ranking, skill routing + exact-text keep/drop (extractive / pointer-not-generator; char-offset compaction; observed a11y/DOM controls; Bash stdout prune; verbatim session ledger / carryforward 0/4 tools≠use; classify-first MCP / jev-sift; Stagehand extract pick-and-copy; jevcumber meaning-as-spec; closed-vote JevOnly; host-owned waymode; jev-compactor framework-agnostic compact+gate 73% product-arm; dizk/jev-lens pre-send views; pi-om observational keep/kind), env triage (OpenSmoke + latch merge-gate; ci-gatekeeper pre-review typed gate; jev-preflight Stop-hook attention redirect, not a merge blocker), moderation/ranking (decision-native RAG evidence set; living class-pattern atlas; meaning-search without embeddings / jevgrep; meaning-grep jev-semgrep; evidence-packet jevex; measured RAG rerank vs generative rerank; sift ~$0.00003/post; ThinkyMiner/Winnow worth-your-attention VOI ≠ kevinpita/winnow), skill routing (route ≠ memory; session-sticky first-prompt lock; OMP/pi fail-open jev_route; OMP prompt suppression / omp-greenlight; skill-broker outline — Jev never grants access; slo-router constrained optimizer + S1 features; skillranker VOI / hook fail-open; jev-in-codex Codex MCP adapter; pi-jev-skill-bench Harbor roster-size harness; pi-jev-skill-suggestion strip-roster), capability kernel / human-confirmed gate (interlock vs toolgate; port-cleanup; permission vs probability; spoken confirm ≠ auth / jev-voice-browser; wrap-as-execution ALLOW/ASK/DENY / AgentGhost — wrap *is* execution; ASK throws; fail-closed; rh-guard owns the gate cousin; construct-auto-classifier privilege ≠ verdict + landed-script / headless≠auto-approve; actiongate-jev — Jev supplies evidence, code owns authority; turnstile — policy first, replay; seal — no seal no advance / coverage ledger; jev-labs — never confidently wrong; jev-handoff typed baton — gate never grants), decide→policy→LLM leftover cascade (jav-email-cascade), closed-vote computer-use (applied-mappings §9; ego-jev hot-click cousin; browser-jev Playwright executes Jev chooses; jeffrey decider≠executor; jev-use hand no-text steps; jev-gpt never free-generates; typesafe-computer-use OCR+AX desktop cousin; jev-voice-browser ASR voice-browser cousin), VOI hunk prune / whole-repo intent (prune-review / jev-intent-review), persist constraints (pi-heed; Jev never writes policy), Pi control plane (pi-jev-control), personal-history ranking (jevfeed; ≠ Winnow), OpenRouter recipe atlas (jev-cookbook), empty compaction-proxy skip (IPECTER), Pi verbatim summarizer replacement (pi-jev-compact ≠ pi-jev-compaction), decision-native inbox (mailordinal), commit pre-review (commitjev), jevex n=16 files-to-read, productized classification API (classifier.dev; spam/inbox/feedback), systematic-review pointer (choxos/jev-reviewer ≠ egma-ai; two-pass; human check never overwritten), prompted-JSON local `/v1/systemone` (githubnext/localjev ≠ kunchenguid/local-jev; wire-compat ≠ logit-equiv), + Laya packaging (NandhaKishorM/laya Router over Hub ckpts; not a TypeSafe drop-in), + external openjev census (@airesearch12 / Benchmark Heaven; tweet ≠ v1.1 ≠ live ranks), + JevBench v1.2 scored board (geo-mean I/C/S/K; cal ON; 534 decisions; ≠ v1.1 87.6), + hourly 0842 apply-the-five (already §73–§78; skip thin; soundness-theater PR gate), + continuous-control S1/S2 (khordoo delta §80; escalate without stall; Local ≠ localjev), + OCR+AX desktop CU (typesafe-computer-use §81; exclusive actions; split kind/item/site; writer/decider; 155× *theirs* one screenshot), + ASR voice-browser CU (jev-voice-browser §82; partial-speech VOI; pointer spans; spoken confirm ≠ auth; 27/27 *theirs* fixtures), + wrap-as-execution ALLOW/ASK/DENY (AgentGhost §83; wrap *is* execution; ASK throws; fail-closed; rh-guard owns the gate cousin), + JP genre atlas (@studio_yebisu §84; stars research-time; not verified evals; **≠** class census **≠** v1.2), + external pedagogy (@akshay_pachaar §85; schema-safe ≠ correct; 200×/400× TypeSafe ceiling; questions-as-code), + meaning-grep dedicated (jev-semgrep §86; proposition ≠ embedding; boolean composition after threshold; Semgrep.dev collision; not a gate) - `.agents/skills/augustus/references/faq.md` — "just classification", - stack replacement, Jev vs open head vs encoder vs LoRA vs constrained AR, - GLiNER vs GLiClass vs CLIP, LLM-as-judge, not-another-how-to + stack replacement, Jev vs open head vs encoder vs LoRA vs constrained AR vs kev vs blackwood, + wait-for-Archer, missing-other confident-wrong, soft project rules vs linter (Abide), extractive/pointer-not-generator, compaction summarize vs pointer, encoder vs Jev compaction, fail-closed keep_full, shadow-mode rollout, fail-open vs fail-closed wake vs CI gate, observe→score→act backend-agnostic, hybrid local decide + remote fill, DONE ≠ verified success, stdout prune vs session compaction, Cua-S1 vs TypeSafe Jev, plan ≠ execute / dry-run, local drop-in vs stub scorer, route ≠ memory, when-it-holds / extractable-from-state, decision-model vs constrained LLM, dual-process S1/S2, S2 never flies / Local controller ≠ localjev / purple = consumed, combinatorial grid ≠ extractive, GLiNER vs GLiClass vs CLIP, LLM-as-judge, in-engine vs CLI store, + hard envelope (bitrate / planner), not-another-how-to, + uncalibrated local likelihoods ≠ Noul, decision-native RAG, classify-first MCP, living applied-mappings atlas / class patterns, draft-gate silence ≠ safer, robotics text-state vs pixels, Stagehand extract pick-and-copy / fast-path not replacement, public judgment wall / six parallel questions, meaning-search without embeddings, attention≠correctness PR review, skills→oxlint not a hard gate, session-sticky fail-closed routing, measured RAG rerank vs generative rerank, cascade + sign-flip / calibration theater, Precision PDF honest negative, + type-safe ≠ correct / Jev is SENSOR not policy, Ax/DSPy knobs vs + typed control plane, native vs verbalized confidence, engine owns + truth / Jev owns judgment, train specialist vs few-shot hosted, + Noul 0.5 cannot-tell never rounded, calibration ≠ sortable, + local `/v1/systemone` ≠ Jev (GLiFormer / gateway), + do not distill Jev as teacher of record, PCD O(1) ≠ calibrated Noul, + closed-vote CU vs Stagehand pick, OMP/pi fail-open vs pi-jev-approver, + permission vs probability / operator owns the safety bar, judgment ≠ + permission / Jev never grants access, eval integrity / instrument not + score, Jev not sole hot-path gate / constrained optimizer + S1 features, + privilege ≠ verdict / effect contracts, attention filter not permission, + measurement owns endorsement / evidence-gated packs, Jev supplies + evidence / code owns authority, ranking ≠ calibration / never + hard-threshold raw p as frequency, Jev `done` ≠ browser success, + never auto-train on the model's own hides, pointer compact ≠ + LLM summarize, combinators not a new model, AUC ≠ ECE / sign + by type, thinking-budget bake-off, local MLX one-pass ≠ Noul, never confidently wrong / TLA+ compose, no seal no advance, sureness vs max_prob, JevBench calibration not in Main Score, CI typed gate before expensive review, Codex MCP adapter, two jev-lens products, Stop-hook not merge blocker, tools≠use, openvons not TypeSafe, Noul not for locks/heaters, two Winnow products, OpenJev `/v1/decide` not drop-in, SemIf runoff ≠ replica, combinators rename + extended five, Noul conflict≠ignorance, jevcache fail-open, typed baton never grants, jevassert record/replay CI, jevarena≠jev-arena, BBQ not a bias cert, jeffrey pick≠fill, jevlint≠JevLint, local-jev not equivalent, constraints survive compaction (pi-heed), jev-judge-bench≠jevarena≠jevbench (no quality headline yet), jev-use≠ultrafast / Vercel drops confidence, pi-jev-control GUI never force-click, jev-gpt never generates, cookbook samples not benches, jevfeed no social graph, openJev-verdict claims ≠ OpenJev / not endorsement, 1-token logprob ≠ Noul / coverage ≠ correctness, jevinf replica ≠ TypeSafe, elixir-sdk ≠ dannote/jev, jevex n=16 rename, commitjev middle band, hermes-plugin-jev is Agnes, pi-jev-compact ≠ pi-jev-compaction, IPECTER runway empty, mailordinal inbox, jev-cli not ready ≠ jevql, laya-multilingual confident-wrong OOD, schema-scorer peaked ranking, HF 401 / GitHub 404 Hub-only, classifier.dev productized HTTP / escalate-under-threshold / silent FALLBACK / vs_jev tracked JSON, choxos/jev-reviewer ≠ egma-ai / two-pass Choice+Noul / not-found / human tick is the product, githubnext/localjev ≠ kunchenguid/local-jev / wire-compat ≠ logit-equiv / prompted JSON ≠ structured read / 1200-req bake-off caveats, NandhaKishorM/laya packaging ≠ Hub-only / Router script-before-p / post-T ECE ≠ raw / Banking77 token-budget / 0.85 still soft / vs-Jev unpublished-here, @airesearch12 census ≠ jevbench v1.1 / GLiNER2+routers class-boundary / incomplete vs Laya-localjev-kev / likes ephemeral, JevBench v1.2 geo-mean I/C/S/K / cal ON rank / 75.3 vs 87.6 not a drop / Luna I=97 rank #7 / option-order 72→21 / ×2 latency assumption / Laya absent gap / Qwen3.8 27B ≠ Archer, hourly 0842 already-folded / apply-the-five / skip thin noise / hard-gate Noul as PR gate is soundness theater, screenshot-to-Jev for CU / 155× Harbor score (typesafe-computer-use: no, and no), waveform-to-Jev / 27/27 Harbor score (jev-voice-browser: no, and no; spoken confirm ≠ auth), AgentGhost sidecar / ASK skip (no, and no; wrap *is* execution; ASK throws; fail-closed), JP genre atlas bake-off / live ★ (studio_yebisu: no, and no; stars research-time; not verified evals), Akshay how-to / 200× Harbor (no, and no; TypeSafe ceiling; schema-safe ≠ correct), jev-semgrep Semgrep.dev / embedding tricks / a gate (no, no, and no; proposition ≠ embedding; boolean after threshold; ranking fail-open) + - `.agents/skills/augustus/references/mappings.md` — classical-method mappings with boundaries, counterexamples, acceptance tests (including Hypothesis cards §6–§19 — promote only with a test that ran) - `.agents/skills/augustus/references/validation.md` — design gate, eval recipes, Jev-for-skills (routing, self-monitoring, testing, modularity, - frontmatter), and Eval & hill-climb (jevals hygiene + Harbor taskset) + frontmatter), and Eval & hill-climb (jevals hygiene + Harbor taskset; + open-jev-laya-bench as ECE/NLL/Brier bake-off exemplar; DMB as + Harbor-style frozen protocol vs constrained LLMs; jevals-data as + CC-BY-4.0 recompute-from-logs feedstock; Abide replay as + Harbor-adjacent soft-rule measurement; solari-reflex Harbor-style + computer-use; gliner2-ultrafast encoder-backend cousin (`DONE` ≠ + success; demo is not a bake-off); Cua-S1 specialist form source-only + (metric names, no checkpoint scores; not TypeSafe Jev); Stagehand + extract pick-and-copy 37/75 no-LLM ~0.5s vs 4.37s (*their* card; + pick ≠ replacement; draft #2951–#2955); jevgrep 79% top-5 vs BM25 + / grep on stripped repos; Jev-RAG one-run vs Spark rerank + (full-context Spark still faster); jev-oxlint Phoenix answer-key; + native-probability calibration arena (jev-arena live Brier 0.0059 / + ECE 0.0620 *theirs*); typed control-plane bake-off shape + (jev-dspy-control-plane; offline stubs ≠ quality); jev-testbench collab arms; ARC-AGI Direct Jev as + combinatorial-≠-extractive negative; jev-gateway-bench Harbor on/off + routing one-run signal; jev-pruner Harbor needle/noise + Terminal-Bench + integration pilot, not a full bench; jev-baselines-eval pre-registered + **AMBIGUOUS** + cascade sign-flip; explore-typesafe-ai synthetic FHIR + Harbor-shaped, not clinically validated; databricks-jev-pdf-lab honest + negative; Domain-jev-maker specialist vs few-shot (KL/r/McNemar); + jav-email-cascade compare arms; jev-orderby-bench ORDER BY gates + (calibration ≠ sortable); jeff GLiFormer cost/accuracy; + system-one-benchmark Jev vs MLX PCD vs AR JSON n=50 (Brier 0.1096 vs + 0.3884); jevex 1/8→6/8 SWE finish n=8; jev-semgrep 0.94/0.98 (10×51 *theirs*; dedicated §86; **51★** ephemeral); + omp-greenlight 1,013/10 default 40.9% / 0 of 94; dinostomp jev-as-if + ECE 0.062 *theirs* / FINDINGS 189; slo-router p95 77.93→490.38 same + routes; construct-auto-classifier Jev 0 dangerous / 975; INSTRUCT_JEV + 119-row instruct seed; jev-packs nine verified packs on pinned + jev-1.13 + **jevassert LANDED** (2,990-case matrix; Jev/Sonnet 5 + accuracy tie, Jev better calibrated 7/9, ~250× cheaper; + sms-spam 0.953/0.040); BBQ 58,492 / 97.28% / $0.3429 *theirs*; + jevlint 13/15 1.00/1.00; grande JGLUE 0.614/0.853; local-jev + done 30%/shape 57%; pi-heed 98.5%/0 false block; does-jev-confidence 8,000 judgments + AUC ~0.91 / stated ~75% vs human ~10% / ~96% ECE removed; + ego-jev n=3 medians ~2× vs per-step LLM; jev-compactor 64.5%/ + 366ms/0 invented paths vs Sonnet summary, one session; + jev-frontier-100 Jev 77.0% vs Qwen3.5 4B/2048 96.7% + (exploratory); jev-ood-calibration 900 tickets ECE 0.107 = + 4.4× floor / sign flips by type; jev-labs 1,080 golden 0 + wrong under chaos (escalate; not a proof of zero); + jevbench v1.1 Jev 1.13.0 Main 87.6 (calibration not scored); + how-sure-is-jev Choice confidence = max_prob; ci-gatekeeper + 504–629 ms own-repo; dizk/jev-lens 79% / 500 trajectories; + jev-compactor product-arm 73%/350ms/4 of 4; carryforward + 0/4 recall; openvons JevPick 3.2–4.8×; HA-Jev 17★ not for + locks; jev-preflight fail-open 8 axes; jevcache 0 FP/100; + zeroshot-vs-bert +0.05–+0.13 / DiD; ThinkyMiner/Winnow + 80%/90%; OpenJev 45/60; semif-serve 1164 vs 178 ms; + typed-evaluation-collapse Noul vs named Choice; + jev-judge-bench SLA-150 contract / canaries ≠ quality / no headline + yet; jev-use 220 ms p50 / 12/12 / Vercel 0.4 *theirs*; jev-cookbook + 425/$0.015 samples not benches; openJev-verdict-2.0 77.10%/0.0636/ + 0.0144 *theirs* unverified + PR #1 claim-audit; + chakuho GUI 336 27B 95%/92% vs Jev 89%/82% *theirs*; + jevinf 2.57×/2.27× 100% argmax; jevex n=16 160s→69s / + $8.74→$3.13; commitjev 0 false on 5 clean *theirs*; + laya-multilingual MASSIVE 0.366/0.387 vs 0.227/0.733; + schema-scorer v2 Choice 0.841; HF 401 this pass; + classifier.dev 400/650 ms; F1 0.887 / 230 ms vs 0.799; + AG News 87.7%; granite 0.546 vs advertised 0.800 *theirs*; + NandhaKishorM/laya T4 32.8 ms / post-T ECE 0.081 vs + Jev 0.246; Banking77 0.425 vs 0.870; 0.766 fine-tune + *theirs*; + @airesearch12 census tweet (engagement ephemeral; not a scored bake-off); + JevBench v1.2 Jev 75.3 / SemIf 74.6 *theirs*; Luna I=96.8 rank #7; cal ON; ≠ v1.1 87.6; + hourly 0842 recipe already §73–§78 / skip thin / soundness-theater PR gate; + khordoo/jev-reflex-autonomy-lab Local-vs-Live A/B, not a scored bake-off; + AgentGhost wrap-as-execution, not a quality bench (MIT **2★**; ASK throws); + @studio_yebisu JP genre atlas tweet (engagement ephemeral; stars research-time; not verified evals); + @akshay_pachaar “Jev Clearly Explained” (engagement ephemeral; 200×/400× TypeSafe ceiling; schema-safe ≠ correct); + uehaj/jev-semgrep meaning-grep dedicated (51★ ephemeral; 0.94/0.98 *theirs* 10×51; Semgrep.dev collision; not a gate)) - `.agents/skills/augustus/references/boundary-audit.md` — existing-system insertion: fit test, opportunity map, smallest boundary, red flags - `.agents/skills/augustus/scripts/evaluate_decisions.py` — offline evaluator diff --git a/docs/ecosystem.md b/docs/ecosystem.md index 3637daa..8ef1673 100644 --- a/docs/ecosystem.md +++ b/docs/ecosystem.md @@ -17,18 +17,31 @@ weekdays. Jev is the densest public corpus, not the class monopoly. - **lhemerly/mcts-agent** — batched Noul prune + Choice PUCT priors + Score/10 leaf value (no random rollouts). - **carlaiau/jev-reranking** — independent TREC DL2019 benchmark: zero-shot Jev best MAP 0.4748, nDCG@10 0.683 vs monoBERT 0.718; $0.76 per 41k pairs. - **superagents-lab/jev-search** — federated web search: Jev understands intent (query/sources/time-range), lanes fan out concurrently, Jev ranks results; merge by URL + engine agreement + rank. +- **kazuhideoki/jev-search** — recursive *file* search + fzf. Not the federated web product. `notes.md` §55. +- **kbhuw/jev-sift** — classify-first MCP: batch path/url/text → Jev before the main agent reads. Topology A, not a host adapter. `notes.md` §56. +- **Bentlybro/jevgrep** — meaning-search CLI+MCP without embeddings (`jgrep`). Packed parallel relevance; 79% top-5 vs BM25 40% / grep 20% on docstring-stripped repos. Keyword still wins exact strings. `notes.md` §58. +- **Max-sm-yc/Jev-RAG** — one-run RAG+Jev rerank vs Muse Spark rerank (≥70% cost / 72% latency); full-context Spark still faster. `notes.md` §58. ### Languages & runtimes - **probably-lang (southpolesteve)** — a programming language whose **loop conditions are Jev feelings**: `while draft feels "like a LinkedIn influencer post" { … }`. Judgment-state recordings give deterministic replay. +- **carldaws/hunch** — Ruby library, not a new language: `almost_certain?` / `pick` / `rate` as chance/Choice/Score. English-as-config. Validations `rescue nil` fail-open at save. `notes.md` §55. - **dannote/jev** — Elixir/OTP: Jev as a peer process; answers are messages; "clause order is the routing, thresholds are guards"; network-free tests. - **jamesward/zio-typesafe-ai** — Effect-oriented (ZIO) client: Jev is the outer Choice; the handler runs the effect. Not Effect.ts and not the Jev HTTP contract. Hypothesis: `mappings.md` §19 (`notes.md` §28). ### Agent harnesses & self-supervision - **Kevthetech143/super-jev** — domain-independent loop: observe → questions → decide → **permit (independent of confidence)** → execute (idempotency key) → verify → JSONL replay. -- **AntonioCoppe/jev-harness** — policy + confidence gate + shadow mode + offline eval CLI; 24-row filter 48.9s (Claude CLI) vs 1.3s Jev. Selective abstention. +- **AntonioCoppe/jev-harness** — policy + confidence gate + shadow mode + offline eval CLI asserting on the **action**; 24-row filter 48.9s (Claude CLI) vs 1.3s Jev. Harbor/jevals-adjacent practice. `notes.md` §33, §44. +- **khordoo/jev-reflex-autonomy-lab** — S1 Jev reflex keeps control; optional S2 planner is one-use advice on low confidence. **Delta:** escalate without stalling; purple telemetry = consumed not arrived; Local controller (rule-based) ≠ githubnext/localjev; 20% starting gate still soft; seed = geometry not replay; no pixels; experimental viz, not a flight controller. `notes.md` §46, §80. +- **perixtar/jev-e2e** — NL cases; Jev selects observed controls; Playwright independently checks. PASS/FAIL/BLOCKED. Alpha. `notes.md` §46. +- **Wany-i/jev-decision-layer** — business decision tool; caller names the judgment; `gate` is part of the result. Unofficial. `notes.md` §46. +- **yalindogusahin/jevpandas** — pandas semantic index; noul/choice/score; LICENSE absent this pass. `notes.md` §46. Accessor sibling: **ktaletsk/jevframe** (PyPI; pandas and Polars `.jev`; full `p__`). `notes.md` §48. - **Friedjof/jev-mobile** — durable Android worker + Mobile MCP; Jev sees prevalidated candidates only. `notes.md` §33. - **jcpsimmons/jev-macos-loop** — Apple-silicon computer-use; local OmniParser/OCR/AX; text-only Jev. Finder demo independently verified. +- **awlevin/typesafe-computer-use** — productized macOS OCR+AX → hosted TypeSafe Choices → deterministic click/type. Never ships a screenshot for the *decision*; writer only for free text; overlapping options = doubt; split kind/item/site; 155× *theirs* one screenshot, not a Harbor taskset. **≠** jev-ultrafast **≠** cua-s1 **≠** jev-macos-loop **≠** camoufox. `notes.md` §81. +- **moritzkremb/jev-voice-browser** — productized ASR → Playwright observe→score→act. Partial transcripts → one 9–11-question Jev request; pointer spans; `is_command` / `complete` / `destructive` gates; numbered overlay, no second model. Spoken confirm ≠ auth. 27/27 fixtures *theirs*, not a Harbor taskset. **≠** jev-voice-control **≠** nikolas-j **≠** typesafe-computer-use. `notes.md` §82. +- **reddpy/AgentGhost** — wrap-as-execution ALLOW/ASK/DENY. The wrap *is* the tool function; rules first; ASK throws; fail-closed. Judge swappable. MIT **2★**. **≠** jwen5419807/agentghost **≠** vventirozos **≠** actiongate **≠** jev-use. rh-guard owns the gate cousin. `notes.md` §83. - **rajdhakad9826/routeKit** — Jev estimates task requirements; policy engine selects the LLM. Jev does not pick the model. +- **jxu-dev-c/jev-adaptive-thinking** — session-sticky first-prompt Jev classification; fail-closed lock to `gpt-5.6-sol`. License null. `notes.md` §58. - **Dicklesworthstone/skillranker** — hook ranks the skill catalog from live context with a calibration loop. - **GodsBoy/jev-agent-skill-router** — 94.4% vs 70.8% lexical routing on 72 requests. - **matthewdonsemail-lab/open-typesafe-camoufox** — browser agent at ~$0.0002/step: 11-way action Choice, free text only when needed. @@ -42,31 +55,39 @@ Per-keystroke launchers (104ms median, sequence-tagged staleness), firehose mode - **arnabgho/rlcd-lite** — GRPO + Brier proper-scoring-rule reward → calibrated decisions; binary reward doesn't calibrate. - **stephanj/parallelConstraintDecoding** — whole JSON schema of booleans/enums in two forward passes (prefill → parallel masked fields). - **Foadsf/jev-for-engineers**, **AbdelStark/jev-benchmarks**, **BrendanH18/jev-lab** — measurement discipline and cost/latency visibility. -- **dayhaysoos/jevals** — local MIT workbench: labeled cases (Noul / Choice / Score), compare runs, WebMCP + agent skill. Empirical acceptance-test surface for Hypothesis mapping cards; complements `evaluate_decisions.py`. Not affiliated with TypeSafe. Pointer: `research/notes.md` §24. Hygiene and the Harbor substrate: `validation.md` Eval & hill-climb (`notes.md` §40). +- **dayhaysoos/jevals** — local MIT workbench: labeled cases (Noul / Choice / Score), compare runs, WebMCP + agent skill. Empirical acceptance-test surface for Hypothesis mapping cards; complements `evaluate_decisions.py`. Not affiliated with TypeSafe. Pointer: `research/notes.md` §24. Hygiene and the Harbor substrate: `validation.md` Eval & hill-climb (`notes.md` §40). Shared bake-off exemplar: [`pngwn/open-jev-laya-bench`](https://huggingface.co/datasets/pngwn/open-jev-laya-bench) (ECE/NLL/Brier; LLM-as-judge is not the score; `notes.md` §46). Harbor-style frozen protocol vs constrained LLMs: [`nibzard/decision-model-benchmark`](https://github.com/nibzard/decision-model-benchmark) (DMB v2; jev banking 76.3% / spam 93.0% / 256+ cap; p50 264–276 ms; $0.07/1k; `notes.md` §49). Feedstock: [`Jevals/jevals-data`](https://github.com/Jevals/jevals-data) (CC-BY-4.0 boards + JSONL; recompute-from-logs; 2026-09-18 board). Boundary map (not a leaderboard): [`Zaious/jev-capability-atlas`](https://github.com/Zaious/jev-capability-atlas). Combinatorial negative: [`simonmesmith/jev-arc-agi-v1-experiment`](https://github.com/simonmesmith/jev-arc-agi-v1-experiment) (Direct Jev 4/400). Pre-registered independent eval (honest negative): [`ickma2311/jev-baselines-eval`](https://github.com/ickma2311/jev-baselines-eval) (both AMBIGUOUS; cascade sign-flip; `notes.md` §55). Healthcare Harbor-shaped: [`si618/explore-typesafe-ai`](https://github.com/si618/explore-typesafe-ai) (synthetic FHIR; not clinically validated). - **jeiel85/jevscope** — local-first visual debugger + JSONL regression for Choice/Score/Noul; compare two definitions; policy buckets are JevScope-derived. Sits next to jevals. Pointer: `research/notes.md` §25. ### Local / open heads & GLi\* species -- **GLiNER / GLiNER2.5 / GLiClass** — species map: locate spans vs categorize the sequence vs local multi-head (fastino-ai GLiNER2.5 CPU-first). Peer of Jev, not a footnote. `references/judgment-class.md`. Author primary source: GLiNER2 "like jev" is schema-conditioned categorize (GLiGuard), not a Noul (`notes.md` §28). 36× Browser Use claim is a tweet (`notes.md` §25). +- **GLiNER / GLiNER2.5 / GLiClass** — species map: locate spans vs categorize the sequence vs local multi-head (fastino-ai GLiNER2.5 CPU-first). Peer of Jev, not a footnote. `references/judgment-class.md`. Author primary source: GLiNER2 "like jev" is schema-conditioned categorize (GLiGuard), not a Noul (`notes.md` §28). 36× Browser Use claim is a tweet (`notes.md` §25). Named jobs: [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) — extractive retention Choice + char-offset copies; not a summarizer; not Jev (`notes.md` §50). [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) — GLiNER2 `fastino/gliner2-multi-v1` scores observed a11y/DOM controls; not GLiNER2.5; not multimodal (`notes.md` §52). - **GLiGuard** (fastino-ai) — 0.3B GLiNER2 encoder, checkpoint `fastino/gliguard-LLMGuardrails-300M`. One bidirectional pass over a safety schema. Same interface shape as batched questions; different objective. Not a Jev weight clone. `judgment-class.md`; `notes.md` §30. -- **DECRUX9812/openjev-lm** — Qwen2.5-0.5B+LoRA distilled from hosted Jev answers; 65/70 = 92.9% on 70 hand-labelled rows (one annotator, one domain, one seed) overnight on 6 vCPU. Its 98.1% on fresh rows is teacher *agreement*, not gold. -- **convaiinnovations/laya** — open Choice/Score/Noul head, text-only, 512 tok. Companion packaging this hour: [`laya-typed-decisions`](https://huggingface.co/convaiinnovations/laya-typed-decisions) (421.3M, acc 0.766 / Brier 0.066 unverified). `notes.md` §18, §42. +- **DECRUX9812/openjev-lm** — Qwen2.5-0.5B+LoRA distilled from hosted Jev answers; 65/70 = 92.9% on 70 hand-labelled rows (one annotator, one domain, one seed) overnight on 6 vCPU, $0/call. Its 98.1% on fresh rows is teacher *agreement*, not gold. Receipts pattern: `notes.md` §25, §44. +- **convaiinnovations/laya** — open Choice/Score/Noul head, text-only, 512 tok. Companion packaging: [`laya-typed-decisions`](https://huggingface.co/convaiinnovations/laya-typed-decisions) (421.3M, acc 0.766 / Brier 0.066 unverified). Shared bake-off: [`pngwn/open-jev-laya-bench`](https://huggingface.co/datasets/pngwn/open-jev-laya-bench) (26+9 tasks, 11959 items; ECE/NLL/Brier; not TypeSafe Jev vs Laya). ONNX replica: [`Mattepiu/laya-onnx`](https://huggingface.co/Mattepiu/laya-onnx) (~15 ms CPU; do not copy vs-Jev table). **GitHub/PyPI face this hour:** [`NandhaKishorM/laya`](https://github.com/NandhaKishorM/laya) (Apache-2.0; **710★**; Router; vs-Jev unpublished-here; `notes.md` §76). `notes.md` §18, §42, §46, §48, §72, §76. +- **jaredpalmer/kev** — Qwen2.5-0.5B LoRA + pointer readout; Apache-2.0; Hub [`jaredpalmer/kev-0.5b`](https://huggingface.co/jaredpalmer/kev-0.5b) plus GitHub release tarball. Runnable Archer reconstruction (`POST /v1/systemone`). Public gold, not a Jev teacher. Isolation exact; ID ECE 0.065 (0.031 after T); acc 0.799 / 1,350. NOTA training must confront `"other"` as a wrong alternative (`notes.md` §45 delta). **100★** this pass (light activity delta; `notes.md` §49). Not a knowledge/frontier substitute. +- **ikermoel/open-alternative-jev** — packed one-forward logprob System One on open LLMs (HF + vLLM). Apache-2.0. **Not a Jev reproduction.** RACE-H 92.9% @ 4.55 q/s on Qwen3.6-27B 8-bit; interference 6–9%. Space demo. `notes.md` §49. +- **wfzyx/von** — 14 MB Needle SAN; local `POST /v1/systemone`; sub-15 ms CPU claim / ~38 ms embed. authored144 needle 52.6% — **not a calibrated Jev replica**. Distinguish from jev-local stub and kev pointer. Do not copy vs-Jev table. `notes.md` §49. +- **Mintzs/jevify** — CUDA/PyTorch packed-logprob cousin on Qwen2.5-1.5B (`ora_decision_engine`). Uncalibrated likelihoods ≠ Noul. Independent of Distillation. No LICENSE this pass. `notes.md` §55. +- **BlackwoodAI/blackwood-rlcd** — open multimodal RLCD (image-text-to-text), Jev-compatible shim, CC BY-NC 4.0. Screenshot + marked candidates → Choice. Card: web acc 0.907 vs Jev 1.13 text-only 0.480; letter-shuffle 0.133 vs 0.587; ECE 0.037; ~200 ms H100. Jev still leads general text 0.850 vs 0.786. Not Archer Watch. `notes.md` §46. +- **Foodoo1/Qwen3-14B-RLCD-Decision-LoRA** — decision-token QLoRA on Qwen3-14B under parallel constrained decoding. Held-out 200-case / 4-field: fraud_risk 64→95%, overall 85.2→98.8% at ~234 ms. Synthetic; not a financial product. `notes.md` §46. - **zmtomorrow/TypeAR** — constrained autoregressive decoding surface: typed fields on a pretrained open model, no retraining. Not a proper-scoring decision head. `research/notes.md` §32. - **stephanj/pcdServer** — native Parallel Constrained Decoder (C++20, llama.cpp GGUF, Apple+Linux). TypeAR-class serving: 2–256 enums, 1–63 parallel fields; softmax over allowed values is not a Noul. `notes.md` §42. - **com-kotobalabs/open-jev-deberta-v3-large** — encoder open-jev, DeBERTa-v3-large 434M, apache-2.0, public gold (not a Jev teacher). In-domain ECE 0.022; OOD acc 0.854→0.690. `notes.md` §33. - **Mikhail/mini-jev-runs** — 27.9k schema-driven decisions; one forward pass; answer from next-token logits; no token generated. Calibration / constrained-decoding gold. `notes.md` §33. - **kokuren/jp-sns-jev7-estimator** — JP SNS seven-axis ONNX distill; teacher scores, not calibrated probabilities; `threat` F1@0.5 = 0. Domain-local categorize. -- **Archer Hume open decision-model** — **Watch.** Qwen3.8 27B dense, 265k, multimodal no audio; one forward pass locally once AR is removed. Driver: AU healthcare data-residency. No Hub weights this pass. `notes.md` §31–§33. +- **Archer Hume open decision-model** — **Watch.** Qwen3.8 27B dense, 265k, multimodal no audio; one forward pass locally once AR is removed. Driver: AU healthcare data-residency. No Hub weights this pass. `notes.md` §31–§33. Runnable *architecture* productization (0.5B, not that drop): **jaredpalmer/kev**. `notes.md` §45. - **bespokelabsai/nimble** — open recipe: contrastive hard labels, not a Jev distill. Model card Apache-2.0 LoRA `bespokelabs/Bespoke-Nimble-9B` on Qwen3.5-9B (repo license absent). Their 324-row holdout is a named receipt, not a ranking. `research/notes.md` §35. - **mmastrac/djev-spark** — DiffusionGemma 26B-A4B NVFP4, Jev-shaped decisions, images as an extension. Third compute graph. Interface claim, not a win over a decision head. `research/notes.md` §36. -- **Perception then judgment** — SAM 3.1 (masks and tracks) or ASR (a transcript) are upstream producers, not the perceive species. System One on that state is decide. Composition, not native omni. Information dies at the interface. `research/notes.md` §39. +- **Perception then judgment** — SAM 3.1 (masks and tracks) or ASR (a transcript) are upstream producers, not the perceive species. System One on that state is decide. Composition, not native omni. Information dies at the interface. Open multimodal *decide* that ships now: blackwood-rlcd (not Archer). `research/notes.md` §39, §46. ### Structural prove ∩ remainder -- **thevibeworks/jevgate** — Proven / Refused / Unknown; cannot block; 0/59 unsafe unasked held-out. Allowlist ∩ System One. +- **thevibeworks/jevgate** — Proven / Refused / Unknown; cannot block; 0/59 unsafe unasked held-out. Allowlist **proves** read-only verbs; Jev judges only unlisted. Allowlist ∩ System One. +- **coldteadotai/abide** — same family, different remainder: linter proves lintable rules; Jev Scores residual soft AGENTS.md rules; fail-open, banded. `notes.md` §47. +- **suraj-phanindra/wellposed** — lint the Jev request before it comes back confidently wrong. Missing `other` → confidence 1.00 on a wrong Choice; gating cannot catch it. `tenbin` owns the lint skill. `notes.md` §46. - **misbahsy/doc-router** — page OCR router: 155→87 billed, 1.74× $ on 19 docs / 155 pages. Same sandwich. ### Agent harnesses extras (this hour) - **kevinpita/pi-jev-context** — reversible Pi context sieve: hide, do not delete; `/jev off` restores. Cousin of winnow/jevprune. -- **vava-nessa/pi-jev-compaction** (and `tamaratran/fast-jev-compaction`) — verbatim drop, never summarize. Pair with jev-gate-student-b for local memory-gating. +- **vava-nessa/pi-jev-compaction** (and `tamaratran/fast-jev-compaction`) — verbatim drop, never summarize. Pair with jev-gate-student-b for local memory-gating. Same *job* as [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) (GLiNER2.5 encoder backend; `notes.md` §50). Stdout cousin [jev-pruner](https://github.com/tamaratran/jev-pruner) — same family, prune Bash before the LLM, not session memory (`notes.md` §53). - **reachjalil/jev-tree** — authored taxonomy so each Choice stays under 255; truncate silently drops the tail (`jev-tree-choice-cap`). - **HacksonClark / SREGym-Lite** — Jev ranks next tests/evidence; does not diagnose; 20/50→24/50 with 2 regressions. `notes.md` §33. - **ddfeyes/jev-mode** — bulk triage/tag/route off frontier context; synthetic 1,000: −77.8% tokens; accuracy is parity. `notes.md` §42. @@ -77,6 +98,12 @@ Per-keystroke launchers (104ms median, sequence-tagged staleness), firehose mode - **TheoOliveira/pi-jev** — Pi semantic tool/skill routing. Not kevinpita/pi-jev-context. - **vtrivedy/jev-plays-games** — Choice over legal moves; probabilities ≠ win odds. - **ant4g0nist/joxide** — zoxide index, Jev shortlist, local paths only. +- **mgaitan/sqlite-jev** — batched NL judgments as a SQLite loadable extension (`jev_rows`). In-engine sibling of jevql's CLI rewrite; inspired by pg-jev. Semantic full scan, not an index. `notes.md` §44. +- **affirmitv/bitrate-advisor** — live ABR: Jev proposes, deterministic policy clamps (never bolder). Missing the model returns policy. `notes.md` §44. +- **nekowasabi/jev-routing** — Go host adapter for Claude Code / Codex / Grok Build. Not MCP, not npx. `notes.md` §44. +- **trietphan/jev-claw** — OpenClaw typed routing: Jev classifies, `decide()` in code. `notes.md` §44. +- **chris-wozniczek/jev-voice-control** — Speech → Jev → macOS actions. README-only this pass. Hypothesis. `notes.md` §44. Productized cousin: **moritzkremb/jev-voice-browser** (§82). Do not collapse. +- **gamesonrblx/JevML** — claimed PCA/MCMC/diffusion/NCA primitives + a picker. README-only. Hypothesis. `notes.md` §44. ### Skills & tooling - **typesafe-ai/skills** — official skill (contracts/patterns). @@ -84,6 +111,7 @@ Per-keystroke launchers (104ms median, sequence-tagged staleness), firehose mode - **ax-llm/ax** — native `typesafe` provider; **typesafeainate/dspy-typesafeify** — DSPy decorator PoC. - **riff (scale-venture-partners)** — hybrid static+semantic linter with ruff-style JEV codes and per-finding calibrated p; 14 calls ≈ $0.0004. - **huntedman/JevLint** — file-level convention Nouls (magic-strings, descriptive-names); write→check→fix; no line-level/auto-fix. Sibling of jev-pref. Independent. Pointer: `notes.md` §26. +- **coldteadotai/abide** — productized Jev preference lint (Claude Code / Codex / OpenCode). Soft instruction-file rules as one Score per rule on the diff; linter owns hard rules; bands + fail-open; compile/calibrate/tune/replay. Replay 93 sessions: edit precision ~26% / turn ~73% (independent review, before tune). Fuller path of jev-pref; complementary to rh-guard. Text/diff only. `notes.md` §47. - **super-jev / probably / jev-search** above — see their cards in `archive/findings.md`. ### Mixed architecture (2026-09-18T14 discourse + topic:jev) @@ -92,14 +120,378 @@ Default placement, not a new product class: Jev judges, an LLM writes, code owns control. Movers that sharpened the card: `git-jev-stage` (exact hunk Choice), `jevprune` (per-line relevance with an always-keep set), `llama-index-jev` (rerank fails open / select fails closed), `jev-pref` -(AGENTS.md as criteria), `lizard-agent` (no LLM when nothing needs writing), -`jevql` (judgment as SQL `WHERE`), OpenSmoke (Jev over every step, LLM only +(AGENTS.md as criteria; Abide is the productized path, `notes.md` §47), `lizard-agent` (no LLM when nothing needs writing), +`jevql` (judgment as SQL `WHERE`; CLI so Postgres never sees `jev()`), +`sqlite-jev` (in-engine SQLite extension; same hole), OpenSmoke (Jev over every step, LLM only on flags). Neighbor skills `tenbin` and `decision-first` are *not* Augustus clones — they own lint/eval and try-Jev-first habit. Entropy allocator (**Hypothesis**): cheap typed scorers for low- and medium-entropy decisions; a frontier write only for high-entropy synthesis (`judgment-class.md`; `research/notes.md` §38). +### Hourly ~14:52 Boise (extractive / local surface / speed layer) + +Patterns, not a catalog. `notes.md` §48. TypeSafe Jev is the exemplar in +the READMEs, not a monopoly. + +- **AppitStudio/testimonial-miner** — extractive selection + multi-question broadcast + offline `redecide`. Model never writes the quote. +- **choxos/jev-reviewer** — pointer-not-generator: line ids; verbatim copy with place; *not found* is an answer. **Delta §74:** **12★**; two-pass Choice+Noul; human tick never overwritten; https://jevreviewer.xera.ac. **≠** egma-ai. +- **egma-ai/jev-reviewer** — Jev assigns PR **attention** P0/P1/P2; OpenAI writes behavior deltas. Attention ≠ correctness. Not the choxos pointer product. `notes.md` §58. +- **us/jev-local** — contract-compatible `POST /v1/systemone`. Default scorer is a **stub** until `JEVLOCAL_SCORER=hf`. +- **hitakshiA/solari-reflex** — observe → decide → verified act; no screenshots. Author table vs Codex on Solari ~3–7× wall. Encoder-backend cousin: gliner2-ultrafast (`notes.md` §52). Specialist-form cousin: cua-s1 (`notes.md` §54). Harness cousin: Stagehand experimental Jev stack (`notes.md` §57). +- **ktaletsk/jevframe** — pandas/Polars `.jev` accessor; full `p__`; sibling of jevpandas. +- **de-niji/jev-hermes** — route ≠ memory: cheap intent gate skips memory tours. +- **ngallodev-software/agent-workflow-typesafe-ai** — advisory sidecar receipts; never changes host routing (Apache-2.0). +- **Joymfl/dag-jev** — structure induction over a bag (experiment; empty README; no metrics). +- **knowlet/jev-agentworld-web-simulator** — decision for control, generator for content; SQLite world. +- **ufx7/jev-testbench** — collab arms (`llm_autonomous` / `scripted_plus_jev` / `llm_plus_jev`); Wilson / McNemar. +- **alexykn/jevscan** — Tree-sitter ∩ typed questions. `tenbin` owns the lint skill. +- **cephalization/jev-oxlint** — skills→oxlint remainder after AST/precheck; Phoenix fixtures; experiment; not a hard gate. `notes.md` §58. +- **phin-tech/pi-jev-approver** — Pi shell gate; fail-closed without a key. Light rh-guard-adjacent note. +- **Mattepiu/laya-onnx** — Laya ONNX port (~15 ms CPU). Do not copy the vs-Jev table. Distinct later replica: [`gqgs/laya-onnx`](https://github.com/gqgs/laya-onnx) (complete Laya→browser int8; conversion smoke; `notes.md` §64). +- **kunchenguid/local-jev** — ModernBERT `/v1/systemone` approximation; **not** behavioral equivalence. Distinct from jev-local stub and jeff. `notes.md` §64. + +Spotcheck this pass (not a fold): SemIf **1551★** (+60 vs awesome claim 1491); jevlike **866★**. Awesomejev 488/21644 not re-derived (public snapshot still 410 / 10,093). Tracker lastModified **2026-09-18T20:12:57Z**; Laya listed; Blackwood not. Archer still Watch. + +### Hourly ~15:52 Boise (boundary map / Harbor bake-off / dual-process) + +Patterns, not a catalog. `notes.md` §49. TypeSafe Jev is the exemplar, not a monopoly. Archer still Watch. X discourse blocked this hour. + +- **Zaious/jev-capability-atlas** — when-it-holds map with API receipts. Axis: extractable from fed state vs needs outside knowledge. History suite table (A wrong@0.90 / B near-flat 0.07 / C right@0.97). Component node ≠ internals-as-FSM. Dangerous-high ECE (DAIR Emotion). Browser-use = DOM-as-text + fan-out, not vision. +- **nibzard/decision-model-benchmark (DMB)** — frozen protocol: jev vs 8 constrained LLMs vs baselines. v2 report of record. jev banking 76.3% / spam 93.0% / 256+ cap; p50 264–276 ms; $0.07/1k. No class wins on quality. +- **Jevals/jevals-data** — CC-BY-4.0 boards + JSONL. Recompute-from-logs. 2026-09-18 board (do not merge Banking77 with DMB/atlas). +- **taro1985/dual-process-ai** — Kahneman S1 decide / S2 generate. Routing accuracy unmeasured. Keyword fallback ≠ S1. +- **simonmesmith/jev-arc-agi-v1-experiment** — Direct Jev 4/400 (1%). Combinatorial ≠ extractive. +- **ikermoel/open-alternative-jev** — packed one-forward; RACE-H 92.9% @ 4.55 q/s. Not a Jev reproduction. +- **wfzyx/von** — 14 MB SAN local drop-in. Distinguishes from jev-local stub / kev pointer. Do not copy vs-Jev table. +- **jaredpalmer/kev** — light delta: **100★** this pass. No species rewrite. + +### Hourly ~16:22 Boise (GLiNER2.5 extractive compaction) + +Architecture notes, not a plugin catalog. `notes.md` §50. TypeSafe Jev is the exemplar, not a monopoly. **Not Jev. Not multimodal.** Archer still Watch. + +- **m-newhauser/gliner25-compaction** — local GLiNER2.5 (`fastino/gliner2.5-base-v1`) retention Choice (`keep_full` / `keep_evidence` / `keep_call_only` / `drop`) + exact character-offset copies. Pointer family with testimonial-miner / jev-reviewer. Mutating tools / shell operators → `keep_full` in code. Fail-closed to `keep_full` (contrast many fail-open Jev gates). Same compaction *job* as fast-jev-compaction / pi-jev-compaction; encoder backend; Fastino/GLiGuard sibling class. `shadowMode` default true. Experimental; characters not tokens; no published retention-quality rates. Apache-2.0. + +### Hourly ~16:48 Boise (CI merge-gate / fail-open wake / S1 indexer / claim-evidence) + +Architecture notes, not a plugin catalog. `notes.md` §51. TypeSafe Jev is the exemplar, not a monopoly. Archer still Watch. + +- **CaseReed/latch** — merge-gate: cluster in code, Jev labels cause, policy owns Gate PASS (infra) vs BLOCK (real). Judge never says ignore alone. Playwright reporter fail-open; `--gate` is a separate step. Pair Harbor + rh-guard. MIT. +- **shitianfang/wakegate** — fail-open VOI wake/resume (Horvitz). Skip only if Jev answers and p(wake)<0.2. 21/21 smoke (same author wrote scenarios+question). Contrast pi-jev-approver fail-closed / jevgate cannot-block. MIT. +- **GreyssonEnterprises/s1-graphify-indexer** (+ `s1-indexer`) — GLiNER default code-graph; escalate LLM only if backend loaded and low conf. 10–50× **unfilled**. Query does not invent edges. License not on GitHub this pass. +- **VladyslavHontar/clear-head** — Stop hook: claims vs session evidence. Keyword retriever; `JEV_FIRM` 0.6 never blocks below. MIT. 1★. +- **reification-labs/foreman** — description-only Phoenix scaffold (S1 specialists + S2 coordinator). Not the super-jev "foreman" loop. No Jev dep. Do not invent an Elixir API. +- **vinilana/jev-gateway-bench** — Harbor on/off routing; hidden chess perft; one-run signal (36/36 both; 4 vs 6 LLM req). Product sibling `jev-gateway` fail-open if Jev down. MIT. +- **LightningK0ala/jev-marshal** — Watch / empty repo. Policy-as-judgment PR cousin of Abide / if-ai. +- **LilDojd/jevons** — bounded Pi supervisor; shadow recovery; never generates commands. MIT. +- MED: if-ai (plain-English PR checks, fail-closed on error); omp-auto-mode (safe/unsafe/ask); jev-downloads-sorter (device-loop Choice); jev-label-desk (description-only); herdr-jev (~260 ms triage + triad; no-key heuristic). + +### Hourly ~16:56 Boise (GLiNER2 Ultrafast observe→score→act) + +Architecture notes, not a browser-agent catalog. `notes.md` §52. TypeSafe Jev is the exemplar, not a monopoly. **Not Jev. Not GLiNER2.5. Not multimodal.** Archer still Watch. + +- **sahibzada-allahyar/gliner2-ultrafast** — local GLiNER2 (`fastino/gliner2-multi-v1`) scores observed a11y/DOM controls. Adaptation of jev-ultrafast. No screenshots; no generated selectors; code owns actuators. Hybrid local decide + remote Mercury 2.5 fill. `DONE` ≠ verified success. Same *job* as jev-ultrafast / solari-reflex; encoder backend. Contrast blackwood-rlcd (screenshot + marked letters). Cousin: ShaunSpark/laya-mind2web-browser-agent (Laya over DOM indices). Fastino sibling class with gliner25-compaction (different hole) and GLiGuard (safety schema). Demo (theirs, not re-run): Flights 12.20 s visible / 13.785 s loop / ~$0.0001 API — demonstration, not a bake-off. MIT. + +### Hourly ~17:15 Boise (jev-pruner evidence-preserving Bash stdout prune) + +Architecture notes, not a plugin catalog. `notes.md` §53. TypeSafe Jev is the exemplar, not the monopoly. **Not a summarizer. Not session compaction. Not GLiNER.** Archer still Watch. + +- **tamaratran/jev-pruner** — after Bash, Jev Noul-prunes stdout chunks before the main LLM sees them. No summary. Hard envelope (≤10k estimated tokens; JSON/XML/YAML/diff/binary; whole-document commands) then soft Noul. Fail-safe keep original; full archive. Marketplace id still `fast-jev-output`. Codex is opt-in wrapper, not automatic interception. Same evidence-preserving *family* as fast-jev-compaction / gliner25-compaction; different *job* (command output vs session memory). Manual sweep (theirs): needles 24/24; mean reduction 83% on trim scenarios. Harbor plugin-eval cannot reach Jev. Terminal-Bench paired pilot is integration, not a full bench. MIT. + +### Hourly ~17:21 Boise (Cua-S1 specialist form System One, source-only) + +Architecture notes, not a Driver / MCP catalog. `notes.md` §54. TypeSafe Jev is the exemplar, not the monopoly. **Not TypeSafe Jev. Not GLiNER. Not a general CUA. Not multimodal pixels-in.** Archer still Watch. Weights Watch. + +- **trycua/cua `libs/cua-s1`** — specialist System One computer-use research. Profile `cua-s1-form-v0` (form-oriented UI). Parent MIT; ~23.3k★ this pass. Byte-level `tinyx` encoder + option-attention head: per observed element fill (from extracted `Label: value`) / check / click / skip. Does not generate values or selectors. Code owns execution order. Plan ≠ execute; dry-run default; `execute` and `submit` independent opt-ins; submit at most one high-confidence Button labeled exactly `Submit` / `Submit Form`; fail-closed on missing checkbox state; fill execution fails closed unless the runtime advertises token-based `set_value`. Source-only: no weights, no checkpoint scores. Offline metric *names* only (accuracy, abstention, coverage, wrong actions/targets, unsafe when should abstain). Tests exercise implementation, not checkpoint quality. Same observe→score-among-candidates→code-acts *job* as jev-ultrafast / gliner2-ultrafast / solari-reflex / laya-mind2web; parallel "System One" name in CUA, not a TypeSafe contract. Watch for a `cua-s1-form-v0` artifact drop. + +### Hourly ~17:48 Boise (CUDA replica, decision-native RAG, verbatim recall, Ruby primitive, FHIR Harbor, AMBIGUOUS baselines) + +Architecture notes, not a CUDA/venv / gem / mcp / uv catalog. `notes.md` §55. TypeSafe Jev is the exemplar, not the monopoly. Archer still Watch. X MCP flap; `since_id` not advanced. + +- **Mintzs/jevify** — CUDA/PyTorch parallel Choice/Score/Noul *shape* on Qwen2.5-1.5B (`ora_decision_engine`). CUDA graphs, branch kernels, literal-label scoring. **Uncalibrated likelihoods ≠ Noul.** Independent of Distillation. Default refund workflow is not a validated policy. No LICENSE this pass. +- **emergency-lee/decision-native-rag-skills** — MIT. Retrieve wide → decide → evidence set → conflict resolve → reason only over kept evidence. Provider-agnostic. No bundled harness. No universal benchmark. Core Augustus RAG mental model. Classify-first MCP cousin: **kbhuw/jev-sift** (`notes.md` §56). +- **Dharundp6/jev-carryforward** — MIT, 1★, npm `carryforward`. Verbatim session ledger; Jev scores recall; rules never judged; fail-open dump. 9×3 hint, not proof. +- **carldaws/hunch** — MIT. Ruby `chance`/`pick`/`rate`; English-as-config; `rescue nil` fail-open at save. Cousin of probably-lang (library, not a new language). +- **si618/explore-typesafe-ai** — FHIR S1 (Jev) + Claude S2 on 100 synthetic Synthea patients. Labels first. 60 requests / 403 judgments. **Not clinically validated.** License not in API this pass. +- **ickma2311/jev-baselines-eval** — MIT. Pre-registered vs nano/frontier/encoder. **Both AMBIGUOUS.** Cascade sign-flip at exact parity; confidence=1.0 theater; encoder 0.933/9ms with labels; serving-path ≠ model-speed; same-day errata ×3. jevals/Harbor practice exemplar. +- **SargeDev/jev-gate-student-b** — light delta only; HF card unchanged (MAE 0.187 / Pearson 0.791 / 90% n=60; fail-open; teacher-copy). +- MED: **fdemir/toolgate** (pre-exec allow/block/review; Jev not authorization; 72-case synthetic); **masa-med-ai/typesafe-screening-mcp** (PubMed include/maybe/exclude; 326 hits ~17s ~$0.014; screening aid); **laurentfabre/databricks-jev-pdf-lab** (honest negative; no OSS license); **yannip1234/codex-jev** (extractive Codex compression; 185→44 estimated tokens is an integration demo; equal accuracy/lower cost not established); **kazuhideoki/jev-search** (recursive *file* search + fzf; **not** superagents-lab federated web search). + +### Hourly ~18:38 Boise 2026-09-18 / 00:38 UTC 2026-09-19 (classify-first MCP + living applied-mappings atlas) + +Architecture notes, not a plugin / showcase catalog. `notes.md` §56. TypeSafe Jev is the exemplar, not the monopoly. Archer still Watch. No invented metrics. No 342-title dump. + +- **kbhuw/jev-sift** — classify first, read selectively. Batch path / public URL / inline text → Jev relevance or 1–8 typed questions. Content to Jev without entering main agent context first. Envelope (theirs): 50 items, 60k char, 2 MB / 20 s, public-IP only, no JS/cookies/login, PDFs unsupported. Uncertain/errors/truncation ≠ irrelevant. Transport tests (mocks) ≠ accuracy. No LICENSE this pass. Same retrieve-wide → decide → evidence-set family as decision-native-rag-skills. Topology A MCP; **not** nekowasabi/jev-routing (host adapter). +- **jevable.com** — living applied-mappings atlas. Claimed **342** curated projects; JSON-LD first page **36**. Categories: Agents, Browser extensions, Creative tools, Data & research, Developer tools, Experiments, Finance, Games, Marketing, Productivity, Robotics. No public API this pass. Class patterns: intent columns, score-among-observed, VOI gates, generative UI decide, robotics text-state, draft-gate silence ≠ safer. Maker clocks stay claims unless already a named receipt. + +### Hourly ~18:48 Boise 2026-09-18 / 00:48 UTC 2026-09-19 (Stagehand experimental Jev pick-and-copy) + +Architecture notes, not an SDK catalog. `notes.md` §57. TypeSafe Jev is the exemplar, not the monopoly. Archer still Watch. Draft stack. No invented metrics. + +- **browserbase/stagehand #2951–#2955** (MIT parent; all OPEN draft; author miguelg719). 5/5 user link: [#2955](https://github.com/browserbase/stagehand/pull/2955) extract completion **judge** + **pick-and-copy**. Jev picks a11y elements; code copies text. `extract` `"off"` | `"judge"` | `"pick"`. Both modes send page/extracted content to TypeSafe. Schema/gate/screenshot-always-LLM in code; LLM fallback. Their card (gemini-3.8-flash, Browserbase, local, 25×3): 69/75 vs 23/25 (92% both); **37/75** no-LLM ~0.5 s vs baseline **4.37 s** / two LLM calls; LLM-off **36/75** — pick is a fast path, not a replacement. Stack: #2951 editable ids (outline byte-for-byte unchanged); #2952 client + pick library (`best`+`strict`); #2953 act tree; #2954 observe + cache-check (errors never block replay). Same observe→score-among-candidates→code-acts *job* as jev-ultrafast / gliner2-ultrafast / cua-s1 / solari, inside a major harness. Do not merge clocks. Do not copy `experimentalJevAct`. + +### Hourly ~18:46 Boise 2026-09-18 / 00:46 UTC 2026-09-19 (public wall, meaning-search, attention≠correctness, skills→oxlint, session-sticky route, measured RAG rerank) + +Architecture notes, not a Convex / uv / pnpm / dylib catalog. `notes.md` §58. TypeSafe Jev is the exemplar, not the monopoly. Archer still Watch. No invented metrics. Do not re-fold §50–§57. + +- **waynesutton/ask-jev-ai** — public realtime judgment wall; 6 parallel questions/ask; policy-in-code (`convex/questions.ts`); safety p≥0.6 blocked; no-key allowlist (UI says Jev offline). Cost $0.000032–$0.000041/ask from TypeSafe token counts → $32–$41/1M. Live askjev.ai. License null this pass. Productized System One primitive surface. +- **Bentlybro/jevgrep** — MIT. Meaning-search CLI+MCP (`jgrep`) without embeddings. Packed parallel Jev relevance; two-stage outline→zoom. 228 questions on docstring-stripped repos: **79% top-5** vs BM25 40% / grep 20%. Keyword still wins exact strings (BM25 top-10 96% vs 85%). Packed+parallel 0.9 s vs serial ~23 min AutoGPT 4,329 files. Distinct from kazuhideoki / superagents-lab / jev-sift. +- **egma-ai/jev-reviewer** — MIT. Jev assigns PR attention P0/P1/P2; OpenAI writes behavior deltas. Attention ≠ correctness (anti-soundness-theater). **Not** choxos/jev-reviewer. Local CLI; does not publish PR comments. Incomplete never P2. Demo: real Jev + labeled prepared explanation copy; live OpenAI pending funded API. +- **cephalization/jev-oxlint** — experiment; nothing published; license null. Skills→oxlint: AST/precheck in code; guidance whole-file in state; survey/calibrate/propose. Phoenix: answer-key agree on every fixture; found flush-only-on-success (noul 0.07); routing sharp; coarse hint not. Not a hard gate. `tenbin` owns the lint skill. +- **jxu-dev-c/jev-adaptive-thinking** — Go CLIProxyAPI plugin; license null. Session-sticky first-prompt classification; later turns never reclassify; fail-closed lock to `gpt-5.6-sol`. Same family as routeKit. Live testing left to the deployer. +- **Max-sm-yc/Jev-RAG** — license null. One-run: ≥70% cost / 72% latency vs Muse Spark *rerank*; full-context Spark still faster (10.60 s). Costs include embeddings. +- MED: **EpicEric/safe-sh** (AGPL-3.0; static shell analysis, not pre-exec auth); **ravikadam/jev-loan-triage** (17 typed questions; policy in `loan.js`); **TurboGuo/jev-fedspeech** + **jev-dating** (Jev vs chat arenas; prior empty search was a query miss); **g-h-miles/jevbox** (MIT; drum grooves). hermes/mcp packs: **no new pack this pass** (hermes-jev-north-star / jev-hermes already folded). + +### Hourly ~19:48 Boise 2026-09-18 / 01:48 UTC 2026-09-19 (capability kernel, typed DSPy control plane, calibration arena, engine-owns-truth, human-confirmed kill) + +Architecture notes, not a pip / venv / Cloudflare catalog. `notes.md` §59. TypeSafe Jev is the exemplar, not the monopoly. Archer still Watch. No invented metrics. Skip parody (`SPFreedom/jef-mcp`) and account farming (`rayelzz/jevregist`). Star spike this day: SemIf 1491→1606 (this pass 1607★); vinnylarouge/jevlike 851→896 (this pass 897★). + +- **somoore/interlock** — MIT, Python. Capability kernel: LLM ring 3 / Interlock ring 0. Secrets never enter the agent (canaries/placeholders). Closed action space. Jev is SENSOR; `policy.py` decides BLOCK/ASK/ALLOW. Anti-pattern: launch-week firewalls that ask "dangerous?" after the LLM already decided with real secrets in scope. Type-safe ≠ correct; irreversible behind threshold AND human. Distinct from toolgate. 38-case local-judge set, not a blind paper. +- **manikanda-kumar/jev-dspy-control-plane** — MIT, Python. Benchmark-first: constrained classifier as typed control plane around DSPy. Classifier → ontology → security override → confidence → state machine → tool allow-list; DSPy drafts AFTER route+action. OpenJEV / DSPy / JSON Schema share ontology. Offline heuristic ≠ quality. Accuracy alone is not enough. +- **meetr1912/jev-arena** (+ **jev-sonar** / **jev-vickrey** / **jev-bracket**) — MIT, Python. Native-probability calibration on analytic worlds (not verbalized confidence). Live *theirs*: 145 noul, Brier 0.0059, ECE 0.0620, 2 requests / 710 ms; overconfident in low bins. Sonar: heatmap-as-policy. Vickrey: Jev never bids; code monotonizes CDF. Bracket: Brier vs Elo; live trailed Elo (honest). +- **JoelLewis/game-coach** — GPL-3.0, TypeScript. Wave 0 PRD: Stockfish owns truth; Jev owns judgment; templates + capped writing model own words. Anti-soundness-theater with egma attention≠correctness. ~$0.012/chess game (theirs). +- **epiphany-dynamics/port-cleanup** — MIT, Swift. Jev recommends; human is the only kill trigger; identity re-check; shields override; mapped explanations not raw model prose. conf ≥ 0.8 for kill recs. Gate UX + rh-guard cousin. +- MED toolbelt: **1jehuang/jev-pr-labeler** (conceptual scope, not line counts); **RubyBrewsday/jevcumber** (.feature only; pointer among observed controls); **shkumbinhasani/typedecide** (class SDK; not on npm); **douglance/jevon** (CLI+MCP; key not in agent config); **buberlo/dsh-jev** (DSH plugin; can only gate, never widen); **zaycruz/fast-jev-compaction-pi** (pi port of fast-jev-compaction); **planstack-ai/jev-tetris-benchmark** (legal set in code; not a rigorous eval); **fabricioctelles/modelsystem** (modelsystem.one catalog; 1★; not affiliated); **emirbartu/opencode-system-one** (fail-open plugin; license null; 1★); **phanngoc/browser-ai** (Go CDP; design done, implementation tracked); **acorn181/semantic-bookmark** (user-authored semantic rules). + +### Hourly ~20:43 Boise 2026-09-18 / 02:43 UTC 2026-09-19 (domain specialist vs few-shot hosted, decide→policy leftover cascade, ORDER BY calibration≠sortable, GLiFormer wire-compat backend) + +Architecture notes, not a uv / bun / Modal catalog. `notes.md` §60. TypeSafe Jev is the exemplar, not the monopoly. Archer still Watch. No invented metrics. Do not re-fold §50–§59. + +- **help-er/Domain-jev-maker** — MIT, Python. Domain LoRA on independent CLINC-150 gold (not a Jev teacher-copy). Matched-precision KL: local 0.168 vs hosted zero-shot 0.580 banking; r +0.933 vs +0.343. Few-shot hosted determinate McNemar n.s. Train specialist when downstream reads p; hosted+examples when only argmax. +- **skiingfalcon/jav-email-cascade** — Python; license null. Decide→policy→LLM leftover. jev vs gen-json vs gen-logprob on one Answer schema. Noul 0.5 never rounded. Mock gen-json confidence flat is *their mock*. Distinct from dual-process-ai (routing unmeasured). +- **yodablocks/jev-orderby-bench** — MIT, Python. Independent ORDER BY measurement. Six gates pass. Score ordinal 0.143 weak link; 53-way 0.99 tie; calibration ≠ sortable. recodelabs batch-40 fails ranking. Not a fourth DuckDB extension. +- **logan-markewich/jeff** — Python; license null. GLiFormer-400M `/v1/systemone` typesafe-sdk drop-in. ~$2.6 vs $15.6 L4 HTTP (~6×); A10G direct ~$0.65 (~24×); AG News 75.5% vs 90.5%. CPU more expensive. Encoder ≠ Jev replica. +- MED: **hraness/sysone** — MIT, TypeScript. Loopback gateway; routes hosted Jev + local OpenJev/NanoJev/Mini-Jev; does not run weights; credential from env never config. + +### Hourly ~21:39 Boise 2026-09-18 / 03:39 UTC 2026-09-19 (active-learning triage / don't distill Jev as teacher, evidence-packet explorer, meaning-grep, closed-vote CU, Jev vs MLX PCD Harbor, host-owned waymode, OMP/pi fail-open gates) + +Architecture notes, not a pip / npm / bun catalog. `notes.md` §61. TypeSafe Jev is the exemplar, not the monopoly. Archer still Watch. No invented metrics. Do not re-fold §50–§60. Skip empty `edwardyen724-g/jev-compactor` and `jlt-commons/laya-jolt`. + +- **ThyFriendlyFox/jev-triage** — MIT, Python. Active-learning: high conf accept / middling expensive teacher / low-or-boundary human. Logs full distributions. **Do not distill Jev as teacher of record** (~68% ceiling). Real outcomes stay the targets. Distinguish Domain-jev-maker (independent gold) vs openjev-lm (teacher-copy). +- **jimmyhealer/jev-semantic-explorer** — MIT, Python (jevex). Index-once ask-many; citable evidence packets. Claude Code 6.8→2.2 files. SWE-bench Verified n=8: **1/8 → 6/8** finish (empty = miss, author-run). Packet n=50 HitFile 0.233 vs BM25 0.159 is diagnostic, not the product KPI. +- **uehaj/jev-semgrep** — JavaScript; LICENSE MIT (GitHub NOASSERTION). Zero-dep Node; AND/OR/NOT over *thresholded* line Nouls; proposition ≠ embedding; contrast-set refund; Semgrep.dev collision; **not a gate**. Distinct from jevgrep. Precision 0.94 / recall 0.98 *theirs*. Dedicated fold `notes.md` §86 (**51★** ephemeral). +- **buluoray/JevOnly** — Apache-2.0, Python. Closed-vote-only: code builds options, Jev only picks; **no planner LLM**. Fact register + verify/undo. 11 steps / 43 calls / ~$0.014 / 17 s *theirs*. Distinct from Stagehand LLM fallback. +- **mallahyari/system-one-benchmark** — Python; license null. Harbor-shaped Jev vs local MLX PCD (Qwen2.5-1.5B) vs AR JSON on LMSYS toxic-chat n=50. Jev **84.0%** / Brier **0.1096**; PCD 52% / 0.3884 / O(1). PCD speed ≠ calibrated Noul. Small n. +- **mossburgh/waymode** — MIT, TypeScript. App retains handlers/permissions/validation/state; Jev over live typed actions. 24/26 + 34/36 *theirs* — bounded development evidence, not a self-driving proof. Not on npm. +- **luw2007/omp-jev-extensions** — MIT, TypeScript. OMP/pi `jev_acceptance_gate` + `jev_route`. **Fail-open** if Jev missing (`confidence: 0`). Contrast pi-jev-approver fail-closed. + +### Hourly ~22:38 Boise 2026-09-18 / 04:38 UTC 2026-09-19 (permission vs probability, judgment ≠ permission outline, eval instrument-not-score) + +Architecture notes, not an `omp plugin` / pip catalog. `notes.md` §62. TypeSafe Jev is the exemplar, not the monopoly. Archer still Watch. No invented metrics. Do not re-fold §50–§61. Watch archive path missing this VM; receipts from live GitHub. + +- **SemetricLabs/omp-greenlight** — MIT, Python. OMP plugin grades gated tool calls and suppresses the approval prompt when Jev says allow. **1,013 calls / 10 sessions / 8.95 h.** Default **40.9%** prompts removed; **0 of 94** unsafe auto-approvals on a 140-row labelled corpus (live traffic unlabelled). Operator owns thresholds; plugin never self-tunes the safety bar. Not a sandbox; host `bash.patterns: deny` fires first. Agent prose withheld (0→3 corpus misses). Never shadows a built-in tool. Composes with waymode / omp-jev-extensions. Distinct from specpi-jev-guard / toolgate / interlock. +- **adamjralph/skill-broker** — language/license null. **Project-outline only** (`PROJECT-OUTLINE.md`). Hermes pre-agent: code owns catalog/policy/grants; Jev scores relevance/confidence and **never grants access**. Jev down → foundation-only; never broaden access. Replayable route evidence. **Not a production recipe.** Distinct from jev-hermes and shipped §5 routers. +- **collapseindex/dinostomp** — Python; README Apache-2.0 (GitHub NOASSERTION); 5★. Eval verification layer: checks the instrument, not just the score. FINDINGS.md 189 (F 52 / D 99 / N 38); **99 against itself**. `dinostomp jev` tests a Jev question like an if-statement (accuracy / p(yes) cut / ECE / blank lean / rewording). Demo *theirs* 24 examples: 100% / ECE **0.062**. Beside jevals, not a Harbor taskset. Anti-soundness-theater cousin of rh-guard / egma / game-coach. +- MED: **nrdz-labs/fast-jev-opencode** (MIT, TypeScript; OpenCode V2 context-hook port of fast-jev-compaction; fail-open; 1★); **yikangy873-gif/jev-desktop** (MIT, JS; Codex Computer Use action selection); **MrDiamondBallz/jev-agent-integration** (MIT, Python; provider-neutral Hermes skill/plugin); **bohutang/sift** (MIT, JS; X feed semantic labels/hide); **CorieW/JevExplore** (TypeScript; license null; bounded web action-space discovery). + +### Hourly ~23:40 Boise 2026-09-18 / 05:40 UTC 2026-09-19 (constrained optimizer + S1 features, privilege ≠ verdict, attention/VOI never-block) + +Architecture notes, not a uvicorn / bun / marketplace catalog. `notes.md` §63. TypeSafe Jev is the exemplar, not the monopoly. Archer still Watch. No invented metrics. Do not re-fold §50–§62. Watch archive path missing this VM; receipts from live GitHub + HF. Hunches labeled. + +- **zeeshan8281/slo-router** — Python; license null. Jev supplies bounded task / exactness / external-evidence features; a constrained controller picks the cheapest backend meeting quality + latency SLO floors. Fail-open to local deterministic features. Exactness raises the quality floor; never overrides capability. Measured *theirs*: same routes/accuracy as the local path; p95 E2E **77.93 → 490.38 ms** (~6.3×). 3/8 task-label disagreements did not change routes. Eight-row demo is not a benchmark. **Hunch:** System One on the feature side of a constrained optimizer, never the sole hot-path gate. Harbor-style latency measurement is mandatory before claiming “Jev routing.” +- **godspede/construct-auto-classifier** — Apache-2.0, TypeScript. Effect-based shell safety gate (OpenCode / Antigravity). Fast-allow/deny <1 ms, then Jev Choice + nine independent risk Nouls. Operator-owned `minConfidence` / `riskThreshold`. Fail-closed. **Privilege ≠ verdict** (`sudo status` can be safe). Certification *theirs*: **975** decisions/model; Jev **0** dangerous allowed; every chat model leaked 16–104. Pair with dinostomp (instrument) and omp-greenlight (operator-owned dial). **Hunch:** contracts on effects, not surface tokens. +- **rashedInt32/jev-lens** (+ **jev-lens.nvim**) — MIT, JavaScript / Lua. Claude Code stop-hook: calibrated “do I need to look / which files / strip debris?” **Never blocks** the agent, never edits files, never says green unless sure. nvim popup is display-only (no API key). Distinct from jev-gates (stops writes). **Hunch:** attention filter / VOI for human review, not a permission gate. Complements skill-broker and omp-greenlight. +- MED: **sysone-help/sysone** (MIT TS; evaluation-model-first SDK; cancellable; never auto-retry; first adapter Jev via Vercel AI Gateway; **not** hraness/sysone loopback gateway); **ctaxnagomi/INSTRUCT_JEV** (HF; MIT; 119 rows, 47/51/21 choice/noul/score; jevals seed); **ckaik/swift-jev** (MIT; LICENSE-only this pass; not a CLI product). + +### Hourly ~00:39 Boise 2026-09-19 / 06:39 UTC (measurement owns endorsement, Jev supplies evidence / code owns authority, ranking ≠ calibration) + +Architecture notes, not a uvx / pnpm / marketplace catalog. `notes.md` §64. TypeSafe Jev is the exemplar, not the monopoly. Archer still Watch. No invented metrics. Do not re-fold §50–§63. Watch archive path missing this VM; receipts from live GitHub. Hunches labeled. Do not re-fold sysone-help/sysone. + +- **dtduc-git/jev-packs** — Python; CC0-1.0. Evidence-gated registry: `pack.yaml` + `cases.jsonl` + `evidence.md`. A pack is only `verified` after accuracy / ECE / cost / latency on pinned `jev-1.13.0`. `unknown` mandatory on every Choice/Score. Backend-neutral. Named suite jevassert → jev-packs → jev-table; **jevassert and jev-table 404 this pass** (runner not released; packs structurally validated by CI). Nine packs `verified` *theirs* (single-run): citation-support 800 / 0.919 / 0.022; rag-answerability 840 / 0.908 / 0.021; moderation 1350 / 0.906 / 0.027; rag-passage-relevance 800 / 0.899 / 0.035; entity-merge 840 / 0.857 / 0.017; support-triage 1350 / 0.887 / 0.059; sms-spam 150 / 0.967 / 0.053; boolq-yes-no 150 / 0.887 / 0.063; banking-intent 150 / 0.840 / 0.090. Dataset-derived keep upstream licenses. Distinct from INSTRUCT_JEV (no evidence gate) and dinostomp (instrument). **Hunch:** Harbor/jevals pattern — measurement owns endorsement; packs without evidence stay `provisional`. +- **omkarghugarkar007/actiongate-jev** — TypeScript Apache-2.0. Runtime authorization: deterministic policy / RBAC / schemas own ALLOW | REVIEW | BLOCK; TypeSafe Jev via OpenRouter is semantic evidence only. Slogan: **"Jev supplies evidence. Code owns authority."** A positive model score never overrides a deterministic security failure. Six narrow questions, never one "is this safe?" Financial / destructive / credential fail-closed if Jev is down. 500-case eval is **label-baseline integrity, not accuracy**. Early MVP. Distinct from toolgate / interlock / construct / greenlight. **Hunch:** canonical anti-soundness-theater counterexample — decision models as sensors, not sole hard gates. Wrap cousin (do not collapse): **reddpy/AgentGhost** — wrap *is* execution; ASK throws; fail-closed (`notes.md` §83). +- **Adilmp/does-jev-confidence-mean-anything** (+ **Adilmp/jevcal**) — Python; license null / MIT. Calibration audit: **8,000** judgments vs `civil_comments` humans on `jev-1.13.0`; $0.05. Ranking strong (AUC **~0.91**; rank corr 0.96) but probabilities systematically shifted toward "yes": when Jev said **~75%**, humans flagged **~10%**. Two-parameter recalibration removes **~96% of ECE** without changing rank (`natural`/tightened ECE 0.156 → 0.006). Accuracy is a trap (`insult`@0.5 61.0% vs always-no 67.8% while AUC 0.83). Vendor "calibrated" is rank-correlation, not frequency units. Companion jevcal: ~100 labelled rows (94% of error at 100). ECE gameable (constant base-rate ECE 0) — they decide on **Brier**. One domain; do not cite `threat` (n=1). **Hunch:** never hard-threshold raw decision-model p as a frequency without domain recalibration. +- MED: **gqgs/laya-onnx** (complete Laya→browser ONNX int8 496.8 MiB; conversion smoke; distinct from Mattepiu/laya-onnx); **kunchenguid/local-jev** (ModernBERT local `/v1/systemone` approximation — not behavioral equivalence; distinct from jev-local stub and jeff). + +### Hourly ~00:39 Boise 2026-09-19 remainder (hot-click CU, verbatim compact+gate, local-rules-then-remainder) + +Architecture notes, not an install.sh / pnpm / wrangler catalog. `notes.md` §65. Same hour as §64. TypeSafe Jev is the exemplar, not the monopoly. Archer still Watch. No invented metrics. Do not re-fold actiongate / jev-packs / sysone-help. + +- **jiangkoumo/ego-jev** — JavaScript MIT. Drive ego-lite with Jev: indexed viewport table → operation+target in one request. Code owns observe/execute/stale-ref/loop/`--until`. Text model only when typing is needed; never guess a fill. Jev `done` ≠ business success. Measured *theirs*: HN **4.9 s vs 9.7 s**, wiki **5.4 s vs 10.1 s** (~2× vs per-step `kimi-k3`; n=3; high variance; not a benchmark). Cousin of jev-ultrafast. Distinct from JevOnly / waymode / Stagehand. +- **edwardyen724-g/jev-compactor** — TypeScript MIT; **1★**. Was empty skip §61. **"Jev judges relevance. Code decides structure."** Never rewrite. Regex floor in code. Compaction fail-open if Jev down; safety fail-closed on pending destructive/exfil. One synthetic 12.7k-token session *theirs*: **64.5%** / **366 ms** / **$0.0004** / **0** hallucinated paths / **4 of 4** facts vs truncate 53%/1 of 4 vs Sonnet summary 96.2%/6.1 s/1 invented path. Claude Code shorter path: fast-jev-compaction. OpenCode fail-open port already §62: fast-jev-opencode. +- **zhuyansen/x-reply-filter** — JavaScript MIT. Chrome MV3. Local `rules.js` first, then batched four Nouls. Collapse not delete. Auto-hides sit in a confirm queue — **never auto-train on the model's own hides**. E2E *theirs*: 3 samples → 0.90/0.93 vs 0.08/0.10. Cousin of bohutang/sift. Cheap hold-before-show cookbook. + +### Hourly ~01:47 Boise 2026-09-19 / 07:47 UTC (control-plane combinators, receipts-not-leaderboard, skill VOI, OOD/AUC≠ECE, frontier-100, turnstile, jevmlx) + +Architecture notes, not an npm / cargo / pip / bun catalog. `notes.md` §66. TypeSafe Jev is the exemplar, not the monopoly. Archer still Watch. No invented metrics. Do not re-fold §50–§65 HIGH except one-line. Atlas axis already §49; skillranker existence already §7; jevmlx reconstruction already §1. Hunches labeled. + +- **voidning/decision-combinators** — TypeScript; README MIT / GitHub license null. Then / Gate / Vote / Cascade / Weighted over Choice/Score/Noul. Analogized as logic gates; **not** literal AND/OR (those stay in code). No measurements. **Hunch:** System One as a control plane, not chat turns. +- **Zaious/jev-capability-atlas** — **10★** this pass. Receipts-not-leaderboard hold/break map. Type-safe ≠ correct (schema-valid ≠ picked-right). Extractable-from-state axis already §49 — do not rehash the history suite. Unofficial. +- **Dicklesworthstone/skillranker** — Rust; **52★**. Two-pass + none-of-these. Claude hook **fail-open** (quiet exit 0) — corrects §7 fail-closed. VOI over a skill library. Distinct from skill-broker (grants). +- **scienthoon/jev-ood-calibration** — MIT. 900 synthetic tickets + 3 public benches; ~$0.06. Public OpenBookQA ECE 0.024 / T 0.96. Synthetic ECE **0.107 = 4.4×** floor; priority (unknowable org rule) 44.7% / mean p **0.74** / T **3.40**; boolean T **0.66**. Sign flips by type. Do not threshold `confidence`. Complements does-jev-confidence. +- **softpudding/jev-frontier-100** — MIT. 100×3. Jev **77.0%**; Qwen3.5 4B off **56.0%** / 512 **78.3%** / 2048 **96.7%**. Exploratory, not preregistered. Attach the thinking budget. Not a ceiling. +- **zyphr-labs/turnstile** — Apache-2.0; experimental alpha; no npm. Deterministic policy first; Jev remainder; receipts + replay. Missing Jev → Review. Jev never grants what policy denied. Actiongate-class clone. +- **bnsd55/jevmlx** — MIT; **28★**. MLX one-pass schema→JSON+probs. Softmax ≠ Noul. No local leaderboard yet. Distinct from system-one-benchmark Harbor table. +- MED: **chopratejas/invalidate** (Apache-2.0; 5★; 157 cases 89.2%/97.5%/0 false invalidations; memory leases); **yottayoshida/jev-intent-review** (under construction; empty search ≠ proof); **shubhangi013/prune-review** (source preview; 22-run cost 1.18% with 305% outlier). + +### Hourly ~02:38 Boise 2026-09-19 / 08:38 UTC (TLA+ compose, SEAL coverage ledger, skill-broker sibling, sureness, JevBench v1.1, CI typed gate, Codex MCP) + +Architecture notes, not a cargo / pip / npm / action.yml catalog. `notes.md` §67. TypeSafe Jev is the exemplar, not the monopoly. Archer still Watch. No invented metrics. Do not re-fold the 01:47 list except sibling contrast. Hunches labeled. + +- **copyleftdev/jev-labs** — Python MIT. TLA+ consensus kernel around live Jev. **Never confidently wrong.** 1,080 golden: 0 wrong under none/realistic/severe *theirs* (severe 314 correct / 46 escalate). Rule of three <0.28% — not a proof of zero. TLC 1,049,750 states / 0 errors. Synthetic pharmacy, not clinical. Inverse of soundness theater: escalate instead of hard-gate. +- **Reasonofmoon/seal** — Python MIT. **No seal, no advance.** Jev answers questions; SEAL answers whether the world may change. coverage.path ∈ {auto|code|human|escalate} visible. Mint ≠ product brain. Zero runtime deps. +- **adamjralph/skill-broker** — outline already §62. Sibling this hour: grants in code vs turnstile runtime authorize vs skillranker advisory VOI. Still not a production recipe. Language/license null. +- **adarc8/how-sure-is-jev** — Python MIT; zero-dep. max_prob/margin/entropy/gini/perplexity → CERTAIN|…|CLUELESS. Choice confidence = max_prob (most generous). Pair with ood-calibration. +- **fstandhartinger/jevbench** — Python MIT; unofficial. v1.1 Capability/Speed/Cost → Main Score. Jev 1.13.0 **87.6** *theirs*. Calibration **reported, not scored**. Native vs verbalized. Partial runs not ranked. Harbor/jevals practice, not a vendor eval. +- **NemanjaManic/ci-gatekeeper-bot-jev** — package.json MIT / GitHub SPDX null. Four typed questions → auto-approve|human-review|block before expensive review. Own-repo Jev **504–629 ms**. Conservative default escalated trivial diffs. Cousin of latch, not flaky-vs-real. +- **teempai/jev-in-codex** — TypeScript MIT. Codex MCP: jev_select_capability / jev_search / jev_triage. Ranking unbenchmarked. Lexical fallback. Distinct from jev-routing (not MCP). + +### Hourly ~03:38 Boise 2026-09-19 / 09:38 UTC (attention redirect, pre-send views, tools≠use, observational memory, open-Jev class, physical-world S1) + +Architecture notes, not a cargo / pip / npm / plugin catalog. `notes.md` §68. TypeSafe Jev is the exemplar, not the monopoly. Archer still Watch. No invented metrics. Do not re-fold the 02:38 list except sibling contrast. Hunches labeled. + +- **muse0509/jev-preflight** — Go MIT; 1★; v0.1.0 Public Beta. Claude Code Stop-hook: eight risk axes in one request; assist = one reinspect then finish; **fail-open**; uncalibrated 0.85. Not a merge blocker (contrast latch / ci-gatekeeper / construct). Owner-run Claude Code 2.1.267: no-key fail-open PASS; key-enabled exactly one continuation *theirs*. Distinct from rashedInt32/jev-lens (human Stop filter). +- **godspede/construct-auto-classifier** — delta. Cert unchanged: Jev **0** dangerous / 975; $0.047/1k; chat leaked. **Landed-script trust** (byte-identical to remote default branch; trusts whoever controls that remote). **Headless ≠ auto-approve** (deny-and-report, not pending/auto-approve). Two gates cannot share one OpenCode prompt. +- **edwardyen724-g/jev-compactor** — product-arm bench *theirs*: **73%** (53–76%) / **350 ms** / $0.0004 / **4 of 4** vs Anthropic 86%/16.8s/3 of 4, Codex 85%, OpenCode 85%, Gemini 61%/4 of 4. 30–250× cheaper. Foreman safety in the same ~300 ms pass. §65 64.5%/366ms is vs-Sonnet on the same session. "Jev judges relevance. Code decides structure." +- **dizk/jev-lens** — TypeScript MIT. Pre-send view selection (outline/focus/testlog/…). 500 SWE-rebench trajectories: **79%** fewer tokens (11.6M → 2.4M). Compress **before** first send — post-send prune +17% cost (cache). **Not** rashedInt32/jev-lens. Claude plugin unmeasured. +- **Dharundp6/jev-carryforward** — delta. Plugin eval: `recall` **0/4** with tools+skill. **tools≠use.** SessionStart hook > hoping the model reaches for memory. 9×3 remains a hint. +- **willfish/pi-observational-memory-jev** — TypeScript MIT. Pi `/om`: Jev keep/kind only; verbatim ledger; model-free compact; kind-keyed durable topics. Failed Jev does not drain the buffer. Same anti-summary thesis as fast-jev-compaction / jev-compactor. Do not install beside amosblomqvist `/om`. +- **genai-craft/openvons** — Python; Apache-2.0 LICENSE / GitHub SPDX NOASSERTION; **7★**. Independent open-Jev class (LM/vision/voice finite-choice+prob; NOTA; execute/confirm/reject). Unrelated to TypeSafe. `/v1/systemone` **wire-compat, not a replica**. JevPick 3.2–4.8× byte-identical *theirs*. Flutter on-device. +- **AboveColin/HA-Jev** — Python MIT; **17★**. First real card (was a gallery stub). Sensors from typed answers; confidence gating; Jev-gates-LLM examples. **Not for locks/heaters/smoke.** `background:` triples laundry separation *theirs*. Confidence uncalibrated. Physical-world System One. +- **kylemclaren/jevql** — architecture note. CLI judges; vanilla Postgres never sees `jev()`. **Judgment outside the store** vs pg-jev / sqlite-jev in-engine. +- **bohutang/sift** — short use-case only. ~$0.00003/post *theirs*. Substance/Humor/Chit-chat/Promo/Junk + AI-written. Minimal consumer categorization surface. + +### Hourly ~04:39 Boise 2026-09-19 / 10:39 UTC (digital-design combinators, VOI cache, skill-routing Harbor harness, zeroshot displacement, typed handoff) + +Architecture notes, not an npm / npx / bun / plugin catalog. `notes.md` §69. TypeSafe Jev is the exemplar, not the monopoly. Archer still Watch. No invented metrics. Do not re-fold the 03:38 list except sibling contrast / combinators rename. Hunches labeled. + +- **voidning/jev-combinators** — TypeScript; package MIT / GitHub SPDX null. **Rename** of decision-combinators (same `created_at`). Core five + **extended** Router / Loop / Retry / Fallback / Memory. Digital-design slogan (transistors / logic gates / chip) is a *metaphor* for soft classifiers; **not** literal AND/OR. npm `jev-combinators` 0.1.0. No measurements. Works with TypeSafe + any `/v1/systemone`. +- **kushals256/jevcache** — TypeScript MIT. OpenAI-compatible proxy: Jev admits same-intent cache hits; skip the expensive LLM. Fail-open. Live eval n=100 *theirs*: Jev **0 FP / precision 1 / recall 0.38 / fpr 0** vs Jaccard@0.35 fpr 0.48; $0.00174. Not in v0: streaming HITs. +- **iamdin/pi-jev-skill-bench** + **pi-jev-skill-suggestion** — TypeScript MIT. Harbor/jevals comparative harness: BM25 vs Jev at roster 50–500; 43 gold. **No live Jev numbers this pass.** Suggestion: strip roster; two-stage 0.30 / 0.40; no-key no-op; tool mode is tools≠use cousin. +- **zhuyansen/jev-zeroshot-vs-bert** — Python MIT. Jev beats clean DeBERTa-c on 7 sets (+0.05–+0.13; PAWS AUC +0.03; arXiv 2026 +0.30) *theirs*. Contaminated 0.901 vs `-c` 0.763. Label-equivalence ~230 / >2048. Banking77 512+ feature hurts. DiD 0.035 vs 0.112. +- **shitianfang/jev-handoff** — TypeScript MIT; alpha v0.1. MCP baton: typed escalate/continue/abort. Gate `allow` never grants. Fail-open. Inverted loop. Vercel drops confidence. Same author as wakegate. +- **ThinkyMiner/Winnow** — TypeScript MIT. Chrome worth-your-attention VOI. **Distinct from kevinpita/winnow.** read/skim/save/skip from typed answers; 80%/90% *theirs*. Not on Chrome Web Store. +- **rsdkrasen/hermes-jev-router** — Python; license null. **Jev WHETHER / Python HOW / LLM WHAT.** Compact original chunks; skip next main-model (needs core patch). Fail-open. Community plugin, not vendor. +- **mleyvaz/jev-typed-evaluation-collapse** — Python; license null. NCML field note v0.3 *theirs*: Noul collapses conflict vs ignorance; named Choice separates p=1.0; binary Choice lexically biased. +- **DowLucas/browser-jev** — TypeScript; license null. Playwright executes, Jev chooses. Sample from the distribution not argmax. Fail only high conf **and** high severity. +- **IamBusy/OpenJev** — Python Apache-2.0. Local 0.6B LoRA+scalar head. `/v1/decide` **not** TypeSafe drop-in. 45/60 *theirs*. Distinct from hraness/sysone OpenJev runners. **dddanielliu/semif-serve** — SemIf behind `/v1/systemone`; 1164 vs 178 ms *theirs*; wire-compat ≠ replica. +- Toolbelt notes: **win4r/jev-security-scan** (MIT; not a cert), **bojansandhaus/jev-decisions** (MIT; 1★; reviews never stop commands), **TeoMastro/jev-vs-llm-guardrails-intent-router** (license null; summary.md 404 this pass). **rh-guard owns reward-hack.** +- **ctaxnagomi/DGUI_HYPERMEM-JEV** — HF MIT; 6-row flywheel (analyze 4 / rerank 2 / supersede 0). Sibling INSTRUCT_JEV. + +Census this hour (user-provided): Awesomejev **flat 561/27007**; tracker likes **43→45**, lastModified unchanged; SemIf **1714** (+10); jevlike **926** (+3). Archer still NOT landed. + +### Hourly ~05:46 Boise 2026-09-19 / 11:55 UTC (record/replay CI, BBQ, decider≠executor, sentence-as-rule lint, open replica substrates) + +Architecture notes, not an npm / uvx / cargo / plugin catalog. `notes.md` §70. TypeSafe Jev is the exemplar, not the monopoly. Archer still Watch. No invented metrics. Do not re-fold §50–§69 HIGH except sibling contrast / jevassert landing / prune-review, intent-review, laya-jolt, local-jev deltas. Hunches labeled. Calibration+cost as first-class gates. Open replicas diversify substrates under one contract — not TypeSafe clones. + +- **dtduc-git/jevassert** — Python Apache-2.0. **LANDED** (was 404 §64). Record/replay CI: accuracy + ECE/Brier + cost/latency **offline from recordings**. PyPI `jevassert`; GH Action `@v0`. Pack SPEC v0 with jev-packs. `unknown` mandatory. Backend-neutral. Exit 0/1/2. McNemar. Do not copy `uvx`. +- **dtduc-git/jev-packs** — **delta**. CC0-1.0; size 0→458. Now pairs with the landed runner. First full matrix *theirs* (2,990 cases): Jev and Sonnet 5 statistical tie on accuracy (Δ≤0.018); Jev better calibrated 7/9; ~250× cheaper ($0.000014–0.000031 vs ~$0.0036). sms-spam this-pass **0.953 / ECE 0.040**. +- **chenmingtang830/jevarena** — TypeScript Apache-2.0. Open BYOK arena; preview jevarena-lab.vercel.app. **Failure-finding, not crowning winners.** Python harness remains JevJudge-Bench. Runnable harness, **not measured findings**. **≠** meetr1912/jev-arena. Always qualify the owner. +- **simonmesmith/jev-bbq-experiment** — R; license null. Full BBQ 58,492; Jev 1.13.0 **56,900 / 97.28%**; amb 99.96% / inf 94.60%; bias **0.04 / 0.34**; **$0.3429 / 7.75 min** *theirs*. 12 of 13 amb errors stereotype-aligned; inf errors mostly unknown (1,487/1,579). Order diagnostic 1/484. **Not a general bias cert.** Dataset CC BY 4.0 BBQ. +- **thomasbrueggemann/jeffrey** — TypeScript MIT. **Decider ≠ executor:** Jev next-tool/progress/risk/done; LLM only fills args. Loop Jev→tool→Jev. Risk≥0.5 pause. Stuck ladder (2 Jev / 0 steps). Pick ≠ fill. Mapping §9 still rejects the fused planner-writer. +- **mizchi/jevlint** — TypeScript MIT. **ast-grep subjects × sentence `ask:` scored by Jev.** Matcher silent-fail vs Jev loud. 13/15 naming/comment rules **1.00/1.00** *theirs*. Review mode 4-fn diff 2 req / $0.00013. Fail-open no-verdict. **≠** huntedman/JevLint. Always qualify the owner. +- **shubhangi013/prune-review** — **delta**. 22-run numbers unchanged: winning 27.9% post hoc; all 22 incl. 305% outlier **1.18%**; excl. outlier 15.9%. Target ~20%. Cost not quality. Safety escarpment always keeps concurrency/auth/a11y/startup. +- **yottayoshida/jev-intent-review** — **delta**. Dual MIT/Apache-2.0. Whole-repo intent VERIFIED/VIOLATION/UNKNOWN/NOT_APPLICABLE. CLI works; GH Action not written. Empty search ≠ proof. +- **Eran-BA/Jev_from_GLiNER2** — spec-only; license null; size 0. GLiNER2-base-v1 → Choice/Score/Noul `/v1/systemone`. **No service, no training, no measurements.** Interface ≠ replica. Distinct from jeff GLiFormer. +- **bokuweb/grande** — Rust; license null. Rust/WebGPU System One; Archer/kev-shaped shared-state branches. JGLUE *theirs*: E2B zshot JNLI **0.614** ECE 0.252→**0.088** T=2.81; JCQA **0.853**. 270M **0.710/0.710**. Packed Δmax 7e-5. Isolation sibling 0.098 / state 0.996. Softmax ≠ Noul until calibrated. +- **jlt-commons/laya-jolt** — **delta** (empty skip §61). Clojure Apache-2.0. Byte-for-byte vs Python `system_one` on README quickstart. ~1e-7 last-digit drift. ~1.7 GB f32. +- **leesk212/JEV-CPU** — Python MIT. SemIf CPU semantic-if + web UI. **Meanblock/JEV-CPU 404** — only leesk212 exists. Cross-ref semif-serve §69. Demo GIF is PoC not a bench. +- **kunchenguid/local-jev** — **delta**. ONNX ModernBERT-large-zeroshot-v2.0. Measured vs jev-1.13.0 *theirs* (136 checkpoints): done **30%** / shape **57%** / r −0.06; gold done 26% vs Jev 87%; 112 min vs 21 s. Confidence omitted. Not equivalence. +- Toolbelt notes: **omkarghugarkar007/actiongate-jev** slogan already §64 (single-use ALLOW/REVIEW/BLOCK). **Nyarlathoteppppp/pi-heed** — TypeScript MIT; 3★. Persist user constraints across compaction; check side-effecting calls. Jev never writes policy. Fail-open. Shadow default. Bench *theirs* v0.8.0+Jev: recall **98.5%** / false block **0.0%** / $0.000058. Live: rule changed mid-session 8/13 off vs 0/13 on. +- MED: **david-j-lustig/system-one-responsible-ai** — MIT; size 0; README+LICENSE only. Framing stub. Pair with BBQ, not a substitute. + +Census **not re-derived** this hour (last §69). Archer still NOT landed. + +### Hourly ~06:43 Boise 2026-09-19 / 12:50 UTC (SGR-judge Harbor contract, control-plane productization, never-generates, recipes+life feed, NAR claim-audit) + +Architecture notes, not an npm / uvx / plugin catalog. `notes.md` §71. TypeSafe Jev is the exemplar, not the monopoly. Archer still Watch. No invented metrics. Do not re-fold §50–§70 HIGH except sibling contrast. Hunches labeled. Harbor-shaped **contract** before a quality headline. Control plane productization. Generation as a tree of Choices. Competing NAR claims are an **audit object**, not an endorsement. + +- **slavadubrov/jev-judge-bench** — Python; README MIT / GitHub SPDX NOASSERTION. Frozen SLA-150: Jev vs cheap schema-guided LLM judges (Luna / DeepSeek-flash / glm-5.3-flash). Human labels; invalid = FN; cost/latency first-class. **No quality headline yet.** 21 offline tests. Canaries *theirs* not quality: Jev OpenRouter 5/5; Luna 10/10; DeepSeek GA 10/10; DeepSeek beta 8/10; GLM 5.3 10/10; GLM 4.7 4/10 overload. $10 Berlin live in progress. Direct TypeSafe untested. Five-field/H5 untested. **≠** chenmingtang830/jevarena **≠** fstandhartinger/jevbench. Always qualify the owner. +- **IPECTER/jev-context-pruner** — **EMPTY SKIP.** Description-only Codex compression-proxy slogan; contents 409 empty. Sibling of fast-jev-compaction / jev-compactor / dizk/jev-lens. Do not invent files. +- **shitianfang/jev-use** — TypeScript MIT v0.4.1. Claude/Codex/pi plugin: hand no-text steps to Jev; writing stays with the LLM. Same author as jev-handoff. Vercel `typesafe-ai/jev` 95 calls *theirs*: p50 **220 ms** / p95 423; 12q **186 vs 2,672 ms**; 20-step 4.3 s / 0 escalated; gate **12/12** / p50 199 ms. Vercel drops confidence → margin default **0.4**; first loop 17/20 then 0/20. Fail-open gate. **≠** jev-ultrafast. Do not copy `npx`. +- **goodruizhan/pi-jev-control** — TypeScript; license null; v0.3.0 private. Pi System-One control plane (router/gate/retry/sieve/review/GUI). Compaction never modifies on-disk session. GUI < threshold → unknown, never force-click. No live quality numbers. Distinct from omp-jev-extensions / jevons / pi-heed / pi-om. Do not copy `pi install`. +- **florian-hoenicke/jev-gpt** — Python; license null. Extreme decider≠executor: never free-generates; one typed question per WordNet/jina tree choice, then rank texts. ~**400** calls / **75 s** / **2 cents** *theirs*. Architecture demo, not a product. Distinct from jeffrey. +- **nexibeo/jev-cookbook** — JavaScript MIT; 1★. 15 OpenRouter recipes. Samples 16–36, **not benchmarks**. Recipes 01–13: **425** calls / **$0.015**; median 0.34–0.45 s; browser **5/6**. Pattern: code prepares, Jev answers. Do not copy OpenRouter tilde-id. +- **fengyiqicoder/jevfeed** — JavaScript MIT. Personal browser-history feed. No likes/follows/accounts. Last 200 pages local; one Jev request per batch of ten (distribution *is* ranking). 17 tests, no network. Distinct from ThinkyMiner/Winnow and kevinpita/winnow. +- **Heman10x-NGU/openJev-verdict-2.0** — **claim-verification, not endorsement.** Python; README Apache-2.0 / GitHub SPDX NOASSERTION. Hub heman10x/openJev-verdict-2.0 HTTP 200. README *theirs* N=2000: acc **77.10%** / Brier **0.0636** / ECE corr **0.0144**. Jev table row is a Laya-catalogued vendor baseline, not independent. **Open PR #1** audits: throughput 24.7/s misread as 25 ms (actual 40.5 ms; 3.5× not 28×); Laya 76.60% inside 95% CI (parity); like-for-like dist ECE 15.13% vs 21.40%, Jev 14.40% slightly lower. **≠** IamBusy/OpenJev `/v1/decide`. + +Census **not re-derived** this hour (last §69). Archer still NOT landed. + +### Hourly ~07:49 Boise 2026-09-19 / 13:52 UTC (1-token logprob ≠ Noul, replica engine, Elixir SDK, commit/jevex VOI, inbox, laya-multilingual) + +Architecture notes, not an npm / uvx / plugin catalog. `notes.md` §72. TypeSafe Jev is the exemplar, not the monopoly. Archer still Watch. No invented metrics. Do not re-fold §50–§71 HIGH except sibling contrast. Hunches labeled. Mental models: decision theory, calibration, VOI, signal detection, anti-soundness-theater. + +- **taku-me/chakuho** — Python MIT. 1-token logprob local `/v1/systemone`. Softmax ≠ Noul. Coverage ≠ correctness. GUI 336 *theirs*: 27B **95%/92%** vs Jev **89%/82%**; `__none__` 97% vs 8B 10%. Cousin jevify / TypeAR / pcdServer / jevmlx. +- **zerodegress/jevinf** — Python MIT; ≥3.14. Open replica engine (NanoJev / decider-2b / Laya). **2.57×/2.27×** 100% argmax *theirs*. MPS only. Wire-compat ≠ replica. +- **phiat/typesafe-elixir-sdk** — Elixir MIT; 1★. Unofficial HTTP client. **≠** dannote/jev OTP peer. +- **jimmyhealer/jevex** — **rename** of jev-semantic-explorer. n=16 SWE *theirs*: 160s→**69s**, $8.74→**$3.13**, 16/16. Keep n=8 1/8→6/8. +- **yodablocks/commitjev** — Python MIT. Pre-review; middle band never rounded; regex first; Nouls decide. 0 false on 5 clean *theirs* (small control). Same owner as jev-orderby-bench. +- **Mrmimee/hermes-plugin-jev** — README MIT / GitHub SPDX null. **Agnes 3.0 Flash** chat-completions branded as Jev. **≠** hermes-jev-router. +- **dev-willbird1936/pi-jev-compact** — TypeScript MIT. Verbatim Pi summarizer replacement. **≠** vava-nessa/pi-jev-compaction. Fail-open if <25% saved. +- **IPECTER/jev-runway** — **EMPTY SKIP.** LICENSE only; created≈pushed 1s. Second IPECTER slogan. +- **Milo318/mailordinal** — TypeScript MIT. Decision-native inbox: nine signals → 100-point policy. Humans own ambiguity. Cousin jav-email-cascade. +- **aarora79/jev-samples** — license null. One runnable sample. Toolbelt, not a bench. +- **shaharia-lab/jev-cli** — Rust Apache-2.0 OR MIT; 0.0.0; 17 issues. **Not ready.** ≠ jevql. +- **convaiinnovations/laya-multilingual** — HF 200; Apache-2.0; 322M; 9 likes. MASSIVE **0.366/0.387** vs English laya 0.227/0.733. Khmer 0.000@0.952. Ships uncalibrated. Route by script. +- **mobarmg/jev-schema-scorer-deberta-v3-large** — Hub MIT; GitHub **404**. v2 Choice **0.841**. Peaked ranking ≠ calibration. +- Skip: jev-routing (host-surface delta: Cursor/Devin CLI); laya-typed-decisions; open-jev-laya-bench / jev-tree-choice-cap / INSTRUCT_JEV **HF 401**; jevlogs GitHub 404 + HF 401. + +Census **not re-derived** this hour (last §69). Archer still NOT landed. + +### User-provided ~08:37 Boise 2026-09-19 / 14:37 UTC (classifier.dev — productized System One HTTP) + +Architecture notes, not an npm / wrangler catalog. `notes.md` §73. Skip Archer. No invented metrics. Hunches labeled. Mental models: selective classification, calibration, VOI, signal detection, anti-soundness-theater. Not SWE-only. + +- **mrmps/classifier-dev** — MIT; **185★**; https://classifier.dev. Public zero-shot HTTP; no key. Jev primary; LLM fallback only. Distinct from ask-jev-ai wall. 400 headlines **650 ms** *theirs*. Smart re-asks single-label <0.7; multi-label **ignores** (re-judge worse, 23 s). Emotion ≥0.9 → **82%** / <0.5 → **29%**; gemini-3.8-flash **87.5→90.0** / **61.8→63.7**. Multi-label F1 **0.887** / **230 ms** vs cascade **0.799** / 1.5 s (eval 232 ms; AG News **87.7%** vs 82.0%). `/benchmark` = tracked JSON; read eval/README (n=7 train-on-test). Silent **FALLBACK**: granite F1 **0.546** vs advertised ~**0.800**. rh-guard owns the gate. Life/business (spam/inbox/feedback). Do not copy wrangler. + +### User-provided ~08:48 Boise 2026-09-19 / 14:48 UTC (choxos/jev-reviewer — systematic-review pointer) + +Architecture notes, not an npm / relay catalog. `notes.md` §74. **Delta of §48.** Skip Archer. Always write **choxos/jev-reviewer** or “systematic-review Jev Reviewer.” **≠** egma-ai/jev-reviewer. + +- **choxos/jev-reviewer** — MIT; **12★**; https://jevreviewer.xera.ac. Local-first Cochrane/PRISMA extraction. Jev picks line ids; code copies verbatim. Two-pass Choice + Noul (quotes ≥ 0.5 *theirs*). *Not found* is an answer. Human tick never overwritten. 18-q **4.6 s / $0.0101** *theirs* (spot check, not a validation study). Do not copy `npm start` / `.env`. + +### User-provided ~08:56 Boise 2026-09-19 / 14:56 UTC (githubnext/localjev — wire-compat ≠ logit-equiv) + +Architecture notes, not a bun / `.env` catalog. `notes.md` §75. Skip Archer. Always write **githubnext/localjev**. **≠** kunchenguid/local-jev. + +- **githubnext/localjev** — MIT; **261★**; GitHub Next. Local Bun `POST /v1/systemone` on DiffusionGemma via Chat Completions; TypeSafe SDK drop-in. Prompted JSON probs + entropy confidence — **not** razorback16 structured-read logits. 1,200-req bake-off *theirs* (M5 Max): Qwen3.6 **76.7%** / Gemma 4 26B-A4B **75.0%** / DiffusionGemma **74.2%** short macro; no definitive winner; do not treat as calibrated. LM Studio cannot load DiffusionGemma. Do not copy bun / `.env`. + +### User-provided ~09:07 Boise 2026-09-19 / 15:07 UTC (NandhaKishorM/laya — packaging, not a new species) + +Architecture notes, not a pip / Colab catalog. `notes.md` §76. Skip Archer. Always write **NandhaKishorM/laya**. Weights stay under convaiinnovations. **≠** TypeSafe `/v1/systemone`. **≠** githubnext/localjev. + +- **NandhaKishorM/laya** — Apache-2.0; **710★**; PyPI `laya`. Router over [`laya`](https://huggingface.co/convaiinnovations/laya) / [`laya-multilingual`](https://huggingface.co/convaiinnovations/laya-multilingual) / [`laya-typed-decisions`](https://huggingface.co/convaiinnovations/laya-typed-decisions). T4 *theirs*: 1q **32.8 ms** / 10q **72.3 ms**. Post-T ECE **0.081** vs Jev **0.246** (raw 0.213 vs 0.144). Banking77 **0.425** vs Jev **0.870**. typed-decisions **0.766** fine-tune (base below majority). Khmer **0.000@0.952**. 0.85 still soft. Jev rows third-party unpublished-here. Do not copy pip / preload. + +### User-provided ~09:14 Boise 2026-09-19 / 15:14 UTC (@airesearch12 Benchmark Heaven census — tweet, not scores) + +Architecture notes, not a leaderboard dump. `notes.md` §77. Skip Archer. Quote the tweet; mark likes **ephemeral**. **≠** [fstandhartinger/jevbench](https://github.com/fstandhartinger/jevbench) v1.1. Watch [benchmarkheaven.com/jev-models](https://benchmarkheaven.com/jev-models); do **not** paste live ranks here. + +- **@airesearch12** — [status/2101259522933186879](https://x.com/airesearch12/status/2101259522933186879) (Florian S). Named ~18 openjevs (system-one-open, openjev-sglang, DeBERTa open-jev, Needle 3, open-alternative-jev, Nimble 9B, SemIf, open-jev Dasein / JoshuaSP, OpenJev razorback16, mini-jev, system-one, system-one-gemma, jevlike, AlexWortega/openjev, **GLiNER2**, Succinct Router 14M, jev-model-router/Director/Loki). GLiNER2 + routers = **class-boundary**. Incomplete vs Laya / githubnext/localjev / kev / TypeAR / openvons. Engagement ephemeral (SIGNAL ~417/9/3; this pass 564/15/5). Do not copy Stripe. Scored sibling §78. + +### User-provided ~09:24 Boise 2026-09-19 / 15:24 UTC (JevBench v1.2 scored board) + +Architecture notes, not a hit list. `notes.md` §78. Skip Archer. Quote the board; mark numbers *theirs*. **≠** tweet census §77 **≠** jevbench v1.1 87.6 **≠** jev-judge-bench **≠** jevarena. + +- **JevBench v1.2** — [benchmarkheaven.com/jev-models](https://benchmarkheaven.com/jev-models); harness [fstandhartinger/jevbench](https://github.com/fstandhartinger/jevbench) (MIT; 0★; HEAD `27ed3d6c`; README SHA `bf1e79ba`; RESULTS SHA `fdfab1a2`). Protocol `jevbench::v1.2`; scored 19 Sept 2026; 15 × 534 (hard 220). Official Score = geometric mean I/C/S/K 25% each. Jev 1.13.0 **75.3**; SemIf **74.6** (−0.7); OpenJev razorback16 67.6. Luna I **96.8** rank **#7**. Cal **ON** rank. Option-order 72%→21%. Self-host latency ×2 assumption; many costs est. Laya absent (gap, not named-excluded). GLiNER2 mapping-excluded; apps out. Qwen3.8 27B Chutes TEE **≠** Archer. Do not copy Stripe / CLI. + +### Hourly ~08:42 Boise 2026-09-19 / 15:37 UTC (already folded — apply, don’t dump) + +Architecture notes, not a hit list. `notes.md` §79. Skip Archer. Named HIGHs already §73–§78. Skip thin noise. Do not copy action.yml / bun / pip. + +- **Apply-the-five** — wire-compat ≠ logit-equiv; productize label+p and mark `FALLBACK`; packaging ≠ new species / script-before-p; pointer-not-generator (two-pass; *Not found*; human tick); census ≠ scored bake-off / geo-mean weights are a design. +- **Skip** — uehaj/jev-semgrep (already §61; dedicated now §86); Akeel-Majeed/JEValuate (≠ ElshinQ/jevaluate); MM-sheng/jevspeak (jev-gpt cousin); fable-jev; jev-model-router already §77. +- **Soundness-theater skip** — [totally-tim/jev-gate](https://github.com/totally-tim/jev-gate) (0★; MIT) / [connectedGraph/claude-jev-warden](https://github.com/connectedGraph/claude-jev-warden) (1★; MIT). A Noul attends or escalates; do not hard-gate as merge/quality. **≠** jev-gateway / MongLong0214/jev-gate / jev-gate-student-b. + +### User-provided ~10:20 Boise 2026-09-19 / 16:20 UTC (AgentGhost wrap-as-execution + JP genre atlas) + +Architecture notes, not an npm / `.env` / OpenRouter catalog. `notes.md` §83–§84. Skip Archer. Quote README/tweet. Do not re-fold named HIGH apps except sibling contrast. rh-guard owns the gate cousin. + +- **reddpy/AgentGhost** — MIT; **2★**; TypeScript; HEAD `ac04e4fb`; README SHA `44145fa9`. Wrap-as-execution ALLOW/ASK/DENY. The wrap *is* the tool function; rules first; ASK throws; `failMode: closed`. Judge is a slot. Hosted provider tools out of reach. **≠** jwen5419807/agentghost **≠** vventirozos **≠** actiongate **≠** toolgate **≠** jev-use. Do not copy `AUTO_APPROVE`. +- **@studio_yebisu** — [status/2101065176069886152](https://x.com/studio_yebisu/status/2101065176069886152). JP genre atlas of high-star Jev apps + open replicas. Stars research-time (typesafe-computer-use 203→**427**; jev-voice-browser 40→**103**). Not verified evals. Engagement ephemeral (this pass 131,234 / 1,934 / 192). SAM 3.1 already §39. OpenRouter Jev no-waitlist is WATCH. **≠** @airesearch12 class census **≠** v1.2 board. Do not dump the 30 repos. + +### User-provided ~10:25 Boise 2026-09-19 / 16:25 UTC (Akshay Pachaar “Jev Clearly Explained”) + +Architecture notes, not an SDK / Python-sample catalog. `notes.md` §85. Skip Archer. Quote the article. Independent pedagogy, not a TypeSafe how-to. + +- **@akshay_pachaar** — [status/2101037514945597645](https://x.com/akshay_pachaar/status/2101037514945597645) / [article](https://x.com/i/article/2100940576741093376). LLM hammer; code owns branches; parallel questions; thresholds in code; **schema-safe ≠ correct**; placements = routing / tool-risk / verify with LLM; shadow-mode; questions-as-code. **200× / 400×** TypeSafe ceiling *theirs*. Text-only. Engagement ephemeral (this pass 233,495 / 2,280 / 235). **≠** official docs **≠** Flavio Copes **≠** LangChain harness **≠** AgentGhost. Do not copy the Python samples. + +### User-provided ~10:30 Boise 2026-09-19 / 16:30 UTC (uehaj/jev-semgrep dedicated) + +Architecture notes, not an npm / `npx` / `.env` / marketplace catalog. `notes.md` §86. Skip Archer. Quote README. Dedicated over §61. rh-guard skip (not a gate). + +- **uehaj/jev-semgrep** — JavaScript; MIT LICENSE / GitHub SPDX NOASSERTION; **51★** ephemeral (SIGNAL ★42; §61 0★); 2 forks / 0 issues. HEAD `21120e9`; README SHA `923e6a5`. Grep by meaning; proposition ≠ embedding; contrast-set refund; AND/OR/NOT after threshold (do not multiply p). Cross-lingual; no index. **≠** [semgrep.dev](https://semgrep.dev) **≠** jevgrep **≠** jev-combinators. 0.94/0.98 *theirs* 10×51, not Harbor. Do not copy npm / marketplace. + +Census **not re-derived**. Archer still NOT landed. + See `references/mixed-architecture.md` in the skill. Class-level family choice: `references/judgment-class.md`. Proof vs judgment (Alloy vs Apalache; DST trio Antithesis / Resonate HQ / PufferLib): diff --git a/research/archive/findings.md b/research/archive/findings.md index 39efe52..71d7fb6 100644 --- a/research/archive/findings.md +++ b/research/archive/findings.md @@ -648,5 +648,1690 @@ are producers not perceive; two call shapes removed from the optimizer card; $0.042/MTok tagged vendor-stated; GodsBoy 94.4% tagged exploratory. `notes.md` §43. - - +## Batch #27 (2026-09-18, ~12:58 Boise hourly) + +Note: `research/notes.md` §44. Docs-only. Archer still Watch (no +architecture rewrite). Do not rehash §42 HIGH. + +- **sqlite-jev (Contract as README):** in-engine SQLite extension; + batched `jev_rows`; sibling *pattern* to jevql, different serving + (DB sees `jev()`). Inspired by pg-jev. Semantic full scan, not an + index. License file absent. 0★. +- **bitrate-advisor (Empirical as a shape):** Jev proposes ABR; policy + is the envelope; never bolder. Missing model → policy answer. + Three-state receipt author-reported. +- **jev-routing (Contract as README):** Go host adapter, not MCP, for + Claude/Codex/Grok. Compact then one Choice + done. +- **jev-claw (Empirical as author's 10/10 + 11 offline tests):** Jev + classifies; `decide()` maps; path regex floors risk; confidence is + min. +- **jev-harness practice:** already §33; this hour assert-on-action + as Harbor-adjacent substrate; recipes across business/life. +- **openjev-lm / jev-gate / DeBERTa / mini-jev-runs / tree-cap / + jev-pref:** frames only (receipts economics; memory gate; encoder vs + decoder replica). No rewrite. +- **jev-voice-control / JevML:** README-only stubs. Hypothesis. +- **X:** mmalisper JOB hybrid +12% geomean, join-order 2× slower, + fail-open to Postgres (author-reported). Higgsfield GenAI auto-route + is a claim. + +Cross-repo addition: (as) structured-store semantic index has an +in-engine vs CLI fork; (at) soft judgment inside a hard envelope +(ABR, planner); (au) distill-to-device as a context sieve, not only +an action gate. + +## Batch #28 (2026-09-18) — kev runnable Archer reconstruction + +Note: `research/notes.md` §45. Docs-only. Folded into PR #2, not a +second PR. Archer 27B drop still Watch. + +- **kev (Empirical as named ID receipt; Contract as README/API).** + `jaredpalmer/kev`, Apache-2.0, 24★ this pass. Qwen2.5-0.5B LoRA + + pointer; `POST /v1/systemone`; typesafe-sdk `base_url`. Public gold, + not a Jev teacher. Isolation exact (Δ 3.7e-6; sibling p=0.03 vs + state 0.99). ECE 0.065 / 0.031 after T; acc 0.799 / 1,350 ID. + Permute 7.4%; IIA mean 0.13; boundary forgery held. 0.5B knowledge; + ID calibration only; not multimodal. +- **Place:** trained decision-only open path next to Laya / Nimble / + Watch. Cleanest *runnable* productization of Archer's reconstruction. +- **Contrast:** TypeAR (constrained AR ≠ Noul) vs encoder DeBERTa + (OOD measured) vs proprietary Jev vs openjev-lm (teacher-copy). +- **Eval:** jevals/Harbor bake-off candidate; mechanism tests mirror + Archer probes. +- **When-to-use:** laptop-local System One for development/eval; not a + knowledge/frontier substitute. No serve how-to. + +Cross-repo addition: (av) the trained decision-only path now has a +shipped API-compatible reconstruction (kev); Watch remains the 27B +announcement. + +## Batch #29 (2026-09-18, ~14:03 Boise hourly) + +Note: `research/notes.md` §46. Docs-only. Folded into PR #2. Archer +27B drop still Watch (Hub empty). No invented metrics. + +- **blackwood-rlcd (Empirical as named vendor receipt; Hypothesis on + your labels).** Open multimodal RLCD, CC BY-NC, Jev-compatible shim. + Web 0.907 vs Jev 1.13 text-only 0.480; letter-shuffle 0.133 vs 0.587; + ECE 0.037; ~200 ms H100. Jev still leads general text 0.850 vs 0.786. + Screenshot vs Jev-text is not the same input. Omni decide can ship + without waiting for Archer. Soft judgment over marked pixel + candidates; code clicks. +- **open-jev-laya-bench (Empirical as that named receipt; not a Jev + ranking).** 26+9 tasks, 11959 test / 3269 cal. Macro acc Δ +0.023 + [+0.013,+0.032] neutral, +0.229 [+0.198,+0.262] home. ECE/NLL/Brier. + LLM-as-judge is not the score. Harbor/jevals practice in the wild. +- **Foodoo1 decision-token QLoRA (Empirical as 200-case receipt).** + Train the single decision token under parallel constrained decode. + fraud_risk 64→95%, overall 85.2→98.8% at ~234 ms/4-field. Synthetic; + not a financial product. Softmax ≠ Noul. +- **jevgate frame (already §25):** allowlist *proves*; Jev judges only + unlisted; fail-open (cannot block). +- **wellposed (Empirical as request-lint recipe):** missing `other` → + confidence 1.00 wrong; gating cannot catch it. Broken state paths. + `tenbin` owns the lint skill. +- **jev-reflex-autonomy-lab:** S1 keeps control; optional S2 one-use + advice. No metrics. License null this pass. +- **MED:** jev-decision-layer (gate is part of the result); jev-e2e + (Playwright checks; confident model cannot substitute); jevpandas + (dataframe semantic index; LICENSE 404). + +Cross-repo addition: (aw) omni decide is a shipped open head, not a +Watch-only hole; (ax) bake-off substrate with ECE/NLL/Brier in the +wild; (ay) decision-token LoRA is how you train constrained-AR, not a +new species; (az) confidence gating cannot catch a forced Choice. + +## Batch #30 (2026-09-18) — Abide productized preference lint + +Note: `research/notes.md` §47. Docs-only. Folded into PR #2. Not a +second PR. Archer still Watch. No invented metrics. Text/diff only. + +- **coldteadotai/abide (Empirical as README + dated replay; Hypothesis + on your AGENTS.md).** MIT, created 2026-09-18, TypeScript, npm + `@coldtea/abide`. Productized Jev hooks: one Score per soft project + rule on the diff, never the conversation. Linter owns hard rules + (jevgate-family sandwich, different remainder). Edit vs turn is an + observation window. Bands ≥0.8 / 0.5–0.8 / <0.5 are their operating + point, not a universal 0.8. Fail-open hooks. Rubric quotes source + lines; calibrate/tune rewrite dead rules. Replay 93 sessions, 1,256 + edits / 147 turns, $0.22: independent-reviewer precision edit 26% / + turn 73% before tune. Turn-phase soft rules held up better. No + turn-number drift. Replay does not measure in-session repair. +- **Siblings:** jev-pref (contract Abide productizes); rh-guard + (eval-integrity, not project soft rules); wellposed (request lint + upstream); jevgate (hard envelope); JevLint (file-level conventions). + Do not merge products. Do not copy hooks. + +Cross-repo addition: (ba) preference lint has a productized compile / +calibrate / tune / replay path; (bb) observation window (edit vs turn) +is question design; (bc) banded fail-open means soft judgment is never +the sole hard veto; (bd) false positives in the rubric, not the model. + +## Batch #31 (2026-09-18) — kev delta (Hub + NOTA) + +Note: `research/notes.md` §45 delta. Docs-only. Folded into PR #2. +Not a rewrite of §45. No species change. No invented metrics. + +- **Hub weights.** [`jaredpalmer/kev-0.5b`](https://huggingface.co/jaredpalmer/kev-0.5b) + HTTP 200. `kev.publish`; `--run` accepts Hub ids; base still + downloads on first load. GitHub release tarball remains. 61★ this + pass (signal had 53★). +- **PEFT.** `task_type=FEATURE_EXTRACTION`; publish patches legacy + adapters with null task_type. Docs mention only kev-0.5b. +- **NOTA training (HIGH question-design).** First run learned "this + wording ⇒ pick it". Fix: add none-of-the-above as a wrong + alternative too; vary wording; dedicated `none_of_the_above` eval + (present vs removed). **No published rates.** wellposed still + owns request-shape lint; training must confront the residual + option. + +Cross-repo addition: (be) bake-off fetch path is a Hub id; (bf) +Choice `"other"` is a training confrontation, not only a request hatch. + +## Batch #32 (2026-09-18 ~14:52 Boise) — extractive / local surface / speed layer + +Note: `research/notes.md` §48. Docs-only. Folded into PR #2. Not a +second PR. Archer still Watch. No invented metrics. Do not re-fold +blackwood-rlcd, open-jev-laya-bench, decision-token LoRA, jevgate, +wellposed, jev-reflex-autonomy-lab, Abide, kev, jevpandas, +bitrate-advisor. + +- **AppitStudio/testimonial-miner (Empirical as README fixture).** MIT. + Gmail → numbered sentences in code → one broadcast (Choice/Noul/Score + + per-sentence Nouls). Model never writes. `redecide` retunes + thresholds on the log. 8-request fixture: 5 candidates / 3 rejected + / 3 header skips. Offline tests use fakes. +- **choxos/jev-reviewer (Empirical as sample study).** MIT. Pointer at + line ids; code copies verbatim with place. *Not found* is an answer. + 712-line article: 1q 10 req / 1.2–2 s; 9q 17 / 2.3 s. Spot checks, + not a validation study. +- **us/jev-local (Contract as surface; Hypothesis on your labels).** + `POST /v1/systemone` drop-in. Default scorer is a **deterministic + stub** until `JEVLOCAL_SCORER=hf`. LICENSE absent this pass. Not a + Jev reproduction. +- **hitakshiA/solari-reflex (Empirical as named table).** MIT. + Observe → decide → verified act; no screenshots. Vs Codex on Solari: + 60.2 s vs 194.9 s; 66 s vs 460 s; 24.2 s vs 98.4 s (~3–7×). +- **ktaletsk/jevframe (Empirical as shape).** MIT, PyPI. pandas and + Polars `.jev`; full `p__`; no silent renormalize. Sibling of + jevpandas, not a re-fold. +- **MED:** jev-hermes (route ≠ memory); agent-workflow-typesafe-ai + (advisory sidecar, Apache-2.0); dag-jev (structure induction; + experiment; no metrics); jev-agentworld-web-simulator (decision + control / generator content); jev-testbench (collab arms; LICENSE + absent); jevscan (AST ∩ semantic; `tenbin` owns lint); pi-jev-approver + (fail-closed without key; light note); Mattepiu/laya-onnx (~15 ms + CPU; do not copy vs-Jev table). +- **Spotcheck:** SemIf 1551★; jevlike 866★; Awesomejev 488/21644 not + re-derived; tracker lastModified 2026-09-18T20:12:57Z; Laya listed; + Blackwood not. + +Cross-repo addition: (bg) extractive keep/drop + offline re-threshold +is judge-once/re-policy; (bh) pointer-not-generator is citation +integrity; (bi) `/v1/systemone` drop-in is a surface — stub ≠ scorer; +(bj) observe→decide→verified-act needs no screenshots; (bk) route ≠ +memory; (bl) advisory sidecar never changes host routing; (bm) collab +arms belong in the measurement curriculum. + +## Batch #33 (2026-09-18 ~15:52 Boise) — boundary map / Harbor bake-off / dual-process + +Note: `research/notes.md` §49. Docs-only. Folded into PR #2. Not a +second PR. Archer still Watch. X discourse blocked this hour. No +invented metrics. Do not re-fold §48 items. + +- **Zaious/jev-capability-atlas (Empirical as an axis, not a + knowledge-breadth estimate).** README MIT / GitHub SPDX + NOASSERTION. Unofficial. Extractable-from-state vs needs-outside + knowledge. History suite table: A wrong@0.90 (Yongzheng; Kangxi by + popular convention; GT contested); B near-flat 0.07 (luck); + C right@0.97 with passage. N=3, single annotator. Internals ≠ FSM; + placement is a component node. Confidence is a distribution + statistic (RLCD). Dangerous-high ECE: DAIR Emotion 48% / 0.819 / + 16% p(correct)=0. Browser-use = DOM-as-text + speculative fan-out, + not vision. +- **nibzard/decision-model-benchmark (Empirical as Harbor/jevals + practice).** LICENSE absent this pass. Frozen protocol; v2 report + of record; $28.34. jev banking 76.3%, spam 93.0%, S3 100%* at + 72.7% valid coverage (256+ cap), S4 flip 13%, S5 admits 49.7% / + ECE 0.246; p50 264–276 ms; S1 $0.07/1k. No class wins on quality. + Do not merge Banking77 87% / 76.3% / 79.67%. +- **Jevals/jevals-data (Contract as feedstock).** CC-BY-4.0. Boards + + JSONL + suites. 2026-09-18 board, suite 0.1.0, 8 systems. Jev + banking77 acc 0.7967 / ECE 0.0981 / p50 467 ms / $0.043/1k on + *this* board. Recompute-from-logs; not a ranking. +- **taro1985/dual-process-ai (Empirical as a productized metaphor).** + MIT. Kahneman S1 decide / S2 generate. Routing fails open; safety + fails closed. Routing accuracy **unmeasured**. Keyword fallback ≠ + S1. +- **simonmesmith/jev-arc-agi-v1-experiment (Empirical as a + negative).** LICENSE absent. Direct Jev 4/400 (1%), 1.125%, ~$2.32, + 10 min. Cell-wise Choice. Combinatorial ≠ extractive. +- **ikermoel/open-alternative-jev (Empirical as packed-logprob + economics).** Apache-2.0. Not a Jev reproduction. RACE-H 92.9% @ + 4.55 q/s; interference 6–9%. +- **wfzyx/von (Empirical as extreme speed/econ surface).** + Apache-2.0. 14 MB SAN; authored144 52.6%. Not jev-local stub, not + kev, not a Jev replica. Do not copy vs-Jev table. +- **jaredpalmer/kev (light delta).** 100★ this pass. No species + rewrite. + +Cross-repo addition: (bn) extractable-from-state is the placement +axis; (bo) retrieve-then-state is VOI with a named receipt; (bp) +Harbor-style class bake-off includes constrained LLMs and +baselines; (bq) public JSONL boards are feedstock, not rankings; +(br) dual-process is S1 decide / S2 generate with unmeasured +routing accuracy; (bs) combinatorial assembly is not extractive +keep/drop; (bt) local `/v1/systemone` has three surfaces (stub / +pointer / tiny SAN). + +## Batch #34 (2026-09-18 ~16:22 Boise) — GLiNER2.5 extractive compaction + +Note: `research/notes.md` §50. Docs-only. Folded into PR #2. Not a +second PR. Archer still Watch. **Not Jev. Not multimodal.** No +invented metrics. Do not re-fold §48 extractive recipes, §49 +bake-off, Abide, kev, GLiGuard species. + +- **m-newhauser/gliner25-compaction (Empirical as README + behavior).** Apache-2.0. Created 2026-09-18T17:22:34Z; 1★ at + capture. Local GLiNER2.5 (`fastino/gliner2.5-base-v1`) chooses + `keep_full` | `keep_evidence` | `keep_call_only` | `drop` for + completed tool pairs and copies **exact character-offset** spans. + Not a prose summarizer. Mutating tools / unknown shell / control + operators → `keep_full`. Low-confidence / invalid evidence fail + closed to `keep_full`. `shadowMode` default true (log, do not + replace history). Reduction in characters, not tokens. No + published retention-quality rates. +- **Mental models:** (1) pointer/extractive vs generator summarizers + (family with testimonial-miner / jev-reviewer); (2) soft retention + Choice under a hard mutation envelope — fail-closed contrast vs + many fail-open Jev gates; (3) same compaction *job* as + fast-jev-compaction / pi-jev-compaction, GLiNER encoder backend, + Fastino/GLiGuard sibling class; (4) shadow mode as safe rollout. +- **rh-guard:** sibling note only (fail-closed retention, hard shell + mutation, shadow). Not reward-hack detection. + +Cross-repo addition: (bu) compaction that writes prose is a +different species from pointer keep/drop; (bv) fail-closed +`keep_full` names the *reduction* as the irreversible act; (bw) +compaction job is backend-agnostic (Jev Score/Noul vs GLiNER +encoder); (bx) shadow-mode default is the rollout for memory +mutation. + +## Batch #35 (2026-09-18 ~16:48 Boise) — CI merge-gate, fail-open wake, S1 indexer, claim-evidence + +Note: `research/notes.md` §51. Docs-only. Folded into PR #2. Not a +second PR. Archer still Watch. No invented metrics. Do not re-fold +§50 compaction, §49 bake-off, Abide, kev, pi-jev-approver, jevgate, +rh-guard species. + +- **CaseReed/latch (Empirical as README / offline demo).** MIT. + Created 2026-09-18T22:09:12Z; 0★. Cluster in code; Jev labels + cause; policy owns Gate PASS vs BLOCK. Judge never says ignore + alone. Reporter never fails Playwright; `--gate` is a separate + step. Demo: 8 connection errors → PASS; 5 assertions → BLOCK. + Message-based grouping fragments (`pallets/click` 13→10). Pair + Harbor + rh-guard. +- **shitianfang/wakegate (Empirical as safety table; 21/21 smoke).** + MIT. Skip only if Jev answers and p(wake)<0.2. Same author wrote + scenarios+question. Savings unmeasured. Contrast fail-closed + pi-jev-approver / cannot-block jevgate / fail-closed compaction. +- **GreyssonEnterprises/s1-graphify-indexer (+ s1-indexer).** GLiNER2 + default; escalate LLM only if backend loaded. 10–50× **unfilled**. + Degraded file-node graph if GLiNER cannot load; query does not + invent edges. License not on GitHub this pass. +- **VladyslavHontar/clear-head (Empirical as README).** MIT. 1★. + Stop hook: claims vs session lines. Keyword retriever; JEV_FIRM + 0.6 never blocks below. +- **reification-labs/foreman.** Description-only Phoenix scaffold. + No Jev dep. Distinguish from super-jev "foreman" loop. +- **vinilana/jev-gateway-bench (Empirical as *shape* + one-run + signal).** MIT. Chess perft hidden verifier; on vs off. Both + 36/36; 4 vs 6 LLM req. Author: not a measurement. Sibling + jev-gateway fail-open if Jev down. +- **LightningK0ala/jev-marshal.** Empty repo. Watch / description. +- **LilDojd/jevons (Empirical as README policy).** MIT. Bounded Pi + supervisor; shadow recovery; never generates commands. +- **MED:** Victor-Casado/if-ai (plain-English PR checks, fail-closed + on error); alexsatch/omp-auto-mode (one-line README); jolehuit/ + jev-downloads-sorter (device-loop Choice); LakshyaChaudhry/ + jev-label-desk (empty README); flaviomartil/herdr-jev (~260 ms + triage + triad; no-key heuristic). + +Cross-repo addition: (by) cluster in code / judge labels / policy +decides (CI flaky-vs-real); (bz) name the irreversible act before +picking fail polarity (wake skip vs merge PASS vs compaction drop); +(ca) Harbor on/off routing with a hidden verifier; (cb) S1 extract ++ escalate-S2 is not a measured 10–50× until the table is filled; +(cc) claim/evidence Stop is anti-hallucinated-done, not a test +runner; (cd) description-only greenfield is not a product receipt. + +## Batch #36 (2026-09-18 ~16:56 Boise) — GLiNER2 Ultrafast observe→score→act + +Note: `research/notes.md` §52. Docs-only. Folded into PR #2. Not a +second PR. Archer still Watch. **Not Jev. Not GLiNER2.5. Not +multimodal.** No invented metrics. Do not re-fold §48 solari, §50 +compaction, §51 CI merge-gate / wake, GLiGuard species, or jev-ultrafast 7.1 s as this demo. + +- **sahibzada-allahyar/gliner2-ultrafast (Empirical as README / + architecture behavior).** MIT. Created 2026-09-18T19:35:23Z; 12★ + at capture. Adaptation of browser-use/jev-ultrafast. Local + GLiNER2 (`fastino/gliner2-multi-v1`) extracts requirements and + scores observed a11y/DOM controls. No screenshots. No generated + selectors or JS. Code owns order, dates, freshness, clicks. + Hybrid: local decide; Mercury 2.5 via OpenRouter for TYPE. + `DONE` is loop termination; apps must verify outcomes + independently. Inspector scores are not calibrated P(success). + Demo (theirs, not re-run): NYC→SFO Flights 12.20 s visible / + 13.785 s loop / ~$0.0001 API. Demonstration, not a bake-off. +- **Mental models:** (1) observe→score-among-candidates→code-acts + is backend-agnostic (Jev Ultrafast ↔ GLiNER Ultrafast; same + lesson as compaction); (2) observed DOM/a11y candidates vs + screenshot multimodal (solari / laya-mind2web DOM-index Laya vs + blackwood-rlcd); (3) hybrid local decide + remote fill; (4) + composition + independent outcome check. +- **rh-guard:** light note only (do not trust `DONE`). Not + reward-hack detection. + +Cross-repo addition: (ce) computer-use selection is +backend-agnostic on the same observed-candidate hole; (cf) +screenshot multimodal is a different input from DOM-as-text; +(cg) local decide + remote TYPE is mixed-architecture economics, +not dual-process-ai; (ch) `DONE` ≠ Harbor-verified success. + +## Batch #37 (2026-09-18 ~17:15 Boise) — jev-pruner evidence-preserving Bash stdout prune + +Note: `research/notes.md` §53. Docs-only. Folded into PR #2. Not a +second PR. Archer still Watch. **Not a summarizer. Not session +compaction. Not GLiNER.** No invented metrics. Do not re-fold §50 +gliner25-compaction as this product, §52 observe→score-act, or +fast-jev-compaction as a duplicate. + +- **tamaratran/jev-pruner (Empirical as README / evals README + behavior).** MIT. Created 2026-09-18T03:00:58Z; 5★ attached + capture, 7★ live this pass. After Bash, Jev Noul-prunes stdout + chunks before the main LLM sees them. No summary. Hard envelope + (≤10k estimated tokens; JSON/XML/YAML/diff/binary; whole-document + commands) then soft Noul. Fail-safe keep original; full archive. + Marketplace id still `fast-jev-output`. Codex is opt-in wrapper. + Manual sweep (theirs, 2026-09-18): needles 24/24; mean reduction + 83% (71–92%) on trim scenarios; wrongly trimmed 0/12; 240 ms. + Harbor plugin-eval cannot reach Jev (fail-safe). Terminal-Bench + paired pilot is integration, not a full bench. +- **Mental models:** (1) evidence-preserving prune ≠ summarizer + (family with gliner25-compaction / jev-reviewer); (2) hard + size/format envelope then soft Noul; (3) fail-safe keep original + (reduction is the irreversible act); (4) stdout prune vs session + compaction are different jobs; host capability shapes the product. +- **rh-guard:** sibling note only (fail-safe / envelope). Not + reward-hack detection. + +Cross-repo addition: (ci) command-output sieve and session compaction +share extractive honesty, not a product; (cj) structure-first +pass-through (≤10k / JSON-diff) is the sandwich, Jev is the remainder; +(ck) Harbor plugin-eval refusing Jev is a fail-safe receipt, not a +missing metric; (cl) marketplace id may lag the repo name. + +## Batch #38 (2026-09-18 ~17:21 Boise) — Cua-S1 specialist form System One (source-only) + +Note: `research/notes.md` §54. Docs-only. Folded into PR #2. Not a +second PR. Archer still Watch. **Not TypeSafe Jev. Not GLiNER. Not a +general CUA. Not multimodal pixels-in.** No invented metrics. Do not +re-fold §52 gliner2-ultrafast as this product, §48 solari-reflex, +§53 jev-pruner, blackwood-rlcd as a screenshot cousin, or +laya-mind2web as a Laya DOM-index cousin. + +- **trycua/cua `libs/cua-s1` (Empirical as README / MODEL_CARD + behavior; weights Watch).** Parent MIT; ~23.3k★ this pass + (updated 2026-09-18T23:21:08Z). Research family of small specialist + computer-use models. Profile `cua-s1-form-v0`. Source-only: no + weights, datasets, or checkpoint scores. `tinyx` byte encoder + + option-attention: per observed element fill / check / click / + skip. Fill values selected from extracted `Label: value` pairs, + not generated. Code owns execution order. Plan ≠ execute; dry-run + default; `execute`/`submit` opt-in; fail-closed unknown checkbox / + fill without advertised token `set_value`. Tests exercise + implementation, not checkpoint quality. Offline metric *names* + only. +- **Mental models:** (1) specialist S1 vs general agent (narrow task + contract; membership ≠ general CUA); (2) Choice among observed + elements / fixed actions (same hole as jev-ultrafast / + gliner2-ultrafast / solari / laya-mind2web; contrast blackwood + screenshot); (3) plan ≠ execute, dry-run default, fail-closed + envelope; (4) parallel "System One" naming, not TypeSafe Jev. +- **rh-guard:** light note only (dry-run / submit opt-in / + fail-closed state). Not reward-hack detection. + +Cross-repo addition: (cm) observe→score-among-candidates→code-acts +is backend-agnostic including a specialist CUA head; (cn) "System +One" in CUA research is a parallel name, not a TypeSafe contract; +(co) plan ≠ execute / dry-run is the sandwich around a soft +specialist; (cp) source-only drops publish metric *names*, not +checkpoint scores — Watch for `cua-s1-form-v0`. + +## Batch #39 (2026-09-18 ~17:48 Boise) — CUDA replica, decision-native RAG, verbatim recall, Ruby primitive, FHIR Harbor, AMBIGUOUS baselines + +Note: `research/notes.md` §55. Docs-only. Folded into PR #2. Not a +second PR. Archer still Watch. X MCP flap; `since_id` not advanced. +Archive `234740` not present locally. No invented metrics. Do not +re-fold §50–§54. Student-b species already §33 — light delta only. + +- **Mintzs/jevify (Empirical as README behavior).** Python. Created + 2026-09-18T23:41:21Z; 0★. CUDA/PyTorch parallel Choice/Score/Noul + *shape* on Qwen2.5-1.5B (`ora_decision_engine` / `ora-decision`). + CUDA graphs, branch kernels, literal-label scoring. Independent of + Distillation. **Uncalibrated model likelihoods, not measured + correctness.** Default refund workflow is not a validated policy. + Default `--answer-encoding letters`. **No LICENSE file this pass.** +- **emergency-lee/decision-native-rag-skills (Empirical as + architecture; Hypothesis as a measured win).** MIT. Created + 2026-09-18T23:29:20Z; 0★. Retrieve wide → decide → evidence set → + conflict resolve → reason only over kept evidence. Provider- + agnostic. No bundled Python harness. No universal benchmark. + Offline replay → shadow → canary → A/B. +- **Dharundp6/jev-carryforward (Empirical as README behavior).** MIT, + 1★. npm `carryforward`. Verbatim JSONL ledger; Jev scores recall; + constraints/corrections never judged; fail-open dump. 9×3 hint, not + proof. +- **carldaws/hunch (Empirical as README / example suite).** MIT. Ruby + `chance`/`pick`/`rate`; English-as-config; validations `rescue nil` + fail-open at save. Stub backend. Cousin of probably-lang (library, + not a new language). +- **si618/explore-typesafe-ai (Empirical as named report; not + clinical validation).** Created 2026-09-18T23:48:44Z; 0★; license + not in API. 100 synthetic Synthea; labels first; 60 requests / 403 + judgments to jev-1.13.0. Report: NEWS2 10/20 → 1/20 under-triage; + 98% of 143 med statuses; 7/20 inbox auto-dispatch all correct; + p50 329 ms; $0.0038. Claude wrote labels. 20 cases/scenario. +- **ickma2311/jev-baselines-eval (Empirical as Harbor/jevals + practice, including the honest negative).** MIT. Created + 2026-09-18T22:57:35Z. Pre-registered vs nano/frontier/encoder. + **Both AMBIGUOUS.** CLINC150 Jev 0.870 vs nano 0.795 vs Terra + 0.915. Banking77 encoder **0.933 / 9 ms**. Cascade Δ +0.265 at + 1pp; **sign flips at exact parity** (R_jev=1.000) because + confidence=1.0 on 102/200 incl. 6 wrong. AUROC neither direction; + no ECE. Latency ~2.2× serving-path, not 40–200×. Errata ×3. +- **SargeDev/jev-gate-student-b (light delta).** HF card unchanged: + MAE 0.187 / Pearson 0.791 / 90% n=60; ~59 ms; fail-open; teacher- + copy. +- **MED:** fdemir/toolgate (MIT; allow/block/review; Jev not + authorization; 72-case synthetic); masa-med-ai/typesafe-screening-mcp + (MIT; 326 hits ~17s ~$0.014; screening aid); laurentfabre/ + databricks-jev-pdf-lab (honest negative; no OSS license); + yannip1234/codex-jev (Apache-2.0; 185→44 estimated tokens + integration demo; equal accuracy/lower cost not established); + kazuhideoki/jev-search (file+fzf; **not** superagents-lab web + search; no LICENSE). + +Cross-repo addition: (cq) uncalibrated local likelihoods ≠ Noul; +(cr) retrieve-wide → decide → evidence set is the RAG sandwich; +(cs) verbatim ledger + scored recall, rules never judged; (ct) +judgment as a language primitive (English-as-config); (cu) +confidence=1.0 theater flips cascade sign at exact parity; (cv) +encoder-with-labels still wins; (cw) serving-path ≠ model-speed; +(cx) kazuhideoki/jev-search ≠ superagents-lab/jev-search; (cy) +toolgate product ≠ ndolinschi vocab; (cz) Precision PDF honest +negative is a result. + +## Batch #40 (2026-09-19 ~00:38 UTC / ~18:38 Boise 2026-09-18) — classify-first MCP + living applied-mappings atlas + +Note: `research/notes.md` §56. Docs-only. Folded into PR #2. Not a +second PR. Archer still Watch. No invented metrics. Do not re-fold +§50–§55. + +- **kbhuw/jev-sift (Empirical as README / schema; Hypothesis as a + measured win).** JavaScript. Created 2026-09-18T00:13:31Z; 10★ + this pass; **no LICENSE file this pass.** Plugin + `0.2.0+codex.20260918200547`; package 0.2.0; author Kush + Bhuwalka. Classify first, read selectively: batch path / public + URL / inline text → Jev relevance or 1–8 typed questions. Direct + `POST /v1/systemone` `jev-latest`. Envelope (theirs): 50 items, + 60k char, 2 MB / 20 s, public-IP only, 3 redirects, no + JS/cookies/login, PDFs unsupported. Uncertain/errors/truncation ≠ + irrelevant. Transport tests (mocks) ≠ accuracy. Same + retrieve-wide → decide → evidence-set family as + decision-native-rag-skills. Topology A MCP; not jev-routing + (host adapter). Cousins: typesafe-screening-mcp, + kazuhideoki/jev-search, jev-pruner, carryforward. +- **jevable.com (Empirical as public showcase; 342 is *their* + count).** Independent curated atlas (Nikunj / `@nikunj` in + JSON-LD). HTTP 200 Railway. Claimed **342**; JSON-LD first page + **36**; `pageSize` 36. Categories: Agents, Browser extensions, + Creative tools, Data & research, Developer tools, Experiments, + Finance, Games, Marketing, Productivity, Robotics. No public API + this pass. Class patterns, not a 342-title dump: (1) intent + columns → jevpandas/jevframe / dabit3 formulas; (2) + score-among-observed → jev-ultrafast / gliner2-ultrafast / + solari / cua-s1 + Your Signal/Near Here (do not merge 7s/$0.0039 + with 12.20s; 100× is a claim); (3) VOI gates → tamara + compaction / jev-pruner / gliner25 / routeKit / Gmail embeddings- + first / jev-sift; (4) generative UI decide → json-render + + jev-agentworld-web-simulator; (5) robotics text-state MuJoCo + geometry-as-text / MOSS / jev-drone / Doom JSON (drawing-pixel + claim ≠ Archer); (6) draft-gate silence-as-safer needs fail-open + / heartbeat vs Abide `<0.5` (edit proceeds). Confirm-don't- + invent: Higgsfield claim, jev-trader, Cambium, SEO 584/139 + `other`, snacks 3000/28s/$0.11 claim, ai-cli/hunch, Manhattan/ + Sudoku ≠ replace A*. + +Cross-repo addition: (da) classify-first MCP is retrieve-wide → +decide → evidence-set on agent I/O; (db) errors/truncation ≠ +irrelevant; (dc) topology A MCP ≠ host-adapter routing; (dd) +living atlas extracts class patterns, not a 342-row dump; (de) +draft-gate silence ≠ safer (heartbeat); (df) robotics text-state ≠ +pixels; (dg) 342 is their count / JSON-LD 36 is page 1. + +## Batch #41 (2026-09-19 ~00:48 UTC / ~18:48 Boise 2026-09-18) — Stagehand experimental Jev pick-and-copy + +Note: `research/notes.md` §57. Docs-only. Folded into PR #2. Not a +second PR. Archer still Watch. No invented metrics. Do not re-fold +§50–§56. Draft stack — Watch merge. + +- **browserbase/stagehand #2951–#2955 (Empirical as PR-body + architecture + their local eval).** Parent MIT. All OPEN draft. + Author miguelg719. Created 2026-09-17T06:40Z. User link **#2955 + (5/5)**: extract completion **judge** + **pick-and-copy**. Jev + picks a11y elements; code copies text. `extract` `"off"` | + `"judge"` | `"pick"`. Both modes send page/extracted content to + TypeSafe. Schema leftovers / screenshot extract / failed gate → + LLM. Their card (gemini-3.8-flash, Browserbase, local, 25×3): + 69/75 vs 23/25 (92% both); **37/75** no-LLM ~0.5 s vs baseline + **4.37 s** / two LLM calls; LLM-off **36/75** — pick is a fast + path, not a replacement. Stack: #2951 editable ids (outline + unchanged); #2952 client + pick (`best`+`strict`; ambiguity + stops); #2953 act tree (LLM fallback); #2954 observe + cache-check + (errors never block replay). Same observe→score-among-candidates→ + code-acts *job* as jev-ultrafast / gliner2-ultrafast / cua-s1 / + solari, inside a major harness. Do not merge clocks. + +Cross-repo addition: (dh) harness pick-and-copy is pointer-not- +generator at product scale; (di) pick is a fast path, not a +replacement (36/75 LLM-off honesty); (dj) cache-check errors fail +open (never block replay); (dk) best+strict is NOTA at pick time; +(dl) draft stack #2951–#2955 Watch merge. + +## Batch #42 (2026-09-19 ~00:46 UTC / ~18:46 Boise 2026-09-18) — public wall, meaning-search, attention≠correctness, skills→oxlint, session-sticky route, measured RAG rerank + +Note: `research/notes.md` §58. Docs-only. Folded into PR #2. Not a +second PR. Archer still Watch. No invented metrics. Do not re-fold +§50–§57 (Stagehand is already §57). + +- **waynesutton/ask-jev-ai (Empirical as README / live wall).** + JavaScript. Created 2026-09-19T00:23:46Z; 0★; **license null.** + Public realtime judgment wall; 6 parallel questions/ask; policy + in `convex/questions.ts`; safety p≥0.6 blocked; no-key allowlist + (UI says Jev offline). Cost $0.000032–$0.000041/ask from TypeSafe + token counts → $32–$41/1M. Live askjev.ai. Productized System + One primitive surface. +- **Bentlybro/jevgrep (Empirical as stripped-repo card).** Python + MIT. Created 2026-09-19T00:09:47Z; 0★. Meaning-search CLI+MCP + without embeddings. 228 questions on docstring-stripped + Flask/httpx/Django/AutoGPT: **79% top-5** vs BM25 40% / grep 20%. + Keyword still wins exact (BM25 top-10 96% vs 85%). Packed+parallel + 0.9 s vs serial ~23 min AutoGPT 4,329 files. Distinct from + kazuhideoki / superagents-lab / jev-sift. +- **egma-ai/jev-reviewer (Empirical as README architecture).** + JavaScript MIT. Created 2026-09-19T00:38:48Z; 0★. Jev assigns + attention P0/P1/P2; OpenAI writes behavior deltas. Attention ≠ + correctness. **Not** choxos/jev-reviewer. Incomplete never P2. + Demo: real Jev + labeled prepared explanation copy; live OpenAI + pending funded API. +- **cephalization/jev-oxlint (Empirical as Phoenix experiment).** + TypeScript. Created 2026-09-19T00:34:49Z; 0★; license null. + Experiment; nothing published. AST/precheck prove; guidance + whole-file in state; remainder judged; not a hard gate. Phoenix: + answer-key agree on every fixture; flush-only-on-success noul + 0.07; routing 0.80–0.94 vs <0.50; coarse hint not. `tenbin` owns + the lint skill. +- **jxu-dev-c/jev-adaptive-thinking (Empirical as README session + machine).** Go. Created 2026-09-19T00:35:37Z; 0★; license null. + Session-sticky first-prompt classification; later never + reclassify; fail-closed lock to `gpt-5.6-sol`. Same family as + routeKit. Live testing left to the deployer. +- **Max-sm-yc/Jev-RAG (Empirical as one-run).** Python. Created + 2026-09-19T00:33:02Z; 0★; license null. ≥70% cost / 72% latency + vs Muse Spark *rerank*; full-context Spark still 10.60 s. Costs + include embeddings. +- MED: EpicEric/safe-sh (AGPL-3.0; static shell analysis); + ravikadam/jev-loan-triage (17 questions; policy in code); + TurboGuo/jev-fedspeech + jev-dating (Jev vs chat arenas; prior + empty search was a query miss); g-h-miles/jevbox (MIT; drums). + hermes/mcp packs: **no new pack this pass.** + +Cross-repo addition: (dm) productized public primitive surface +(six parallel questions; policy-in-code; cost-to-1M from tokens); +(dn) meaning-search without embeddings (packed parallel; keyword +still wins exact strings); (do) attention ≠ correctness on a PR +(anti-soundness-theater; not choxos pointer); (dp) skills→oxlint +AST prove ∩ remainder without hard-gating; (dq) session-sticky +first-prompt route fail-closed to a declared fallback; (dr) +measured RAG rerank vs generative rerank must keep the no-RAG +latency arm visible. + +## Batch #43 (2026-09-19 ~01:48 UTC / ~19:48 Boise 2026-09-18) — capability kernel, typed DSPy control plane, calibration arena + fan-out suite, engine-owns-truth, human-confirmed port cleanup + +Note: `research/notes.md` §59. Docs-only. Folded into PR #2. Not a +second PR. Archer still Watch. No invented metrics. Do not re-fold +§50–§58. Skip SPFreedom/jef-mcp (parody) and rayelzz/jevregist +(account farming). + +- **somoore/interlock (Empirical as README architecture).** Python + MIT. Created 2026-09-19T01:41:58Z; 0★. Capability kernel: LLM + ring 3 / Interlock ring 0. Secrets never in the agent. Closed + action space. Jev SENSOR; `policy.py` BLOCK/ASK/ALLOW. Type-safe + ≠ correct. Distinct from toolgate. 38-case local-judge set, not + a blind paper. +- **manikanda-kumar/jev-dspy-control-plane (Empirical as README + architecture + metric list).** Python MIT. Created + 2026-09-19T01:36:58Z; 0★. DSPy drafts AFTER route+action. OpenJEV + / DSPy / JSON Schema share ontology. Offline heuristic ≠ quality. +- **meetr1912/jev-arena (Empirical as their live card).** Python + MIT. Created 2026-09-19T01:28:11Z; 0★. 145 noul, Brier 0.0059, + ECE 0.0620, 2 requests / 710 ms; overconfident in low bins. + Siblings: jev-sonar (heatmap-as-policy), jev-vickrey (Jev never + bids), jev-bracket (live Brier 0.2853 vs Elo 0.2322 — trailed + Elo; honest). +- **JoelLewis/game-coach (Empirical as PRD; Hypothesis as shipped + product).** TypeScript GPL-3.0. Created 2026-09-19T01:08:53Z; + 0★. Wave 0. Stockfish owns truth; Jev owns judgment. +- **epiphany-dynamics/port-cleanup (Empirical as README safety + model).** Swift MIT. Created 2026-09-19T00:58:50Z; 0★. Human is + the only kill trigger; identity re-check; shields; mapped + explanations. +- MED: jev-pr-labeler, jevcumber, typedecide (not on npm), jevon, + dsh-jev (not on npm), fast-jev-compaction-pi, jev-tetris-benchmark + (not a rigorous eval), modelsystem (1★), opencode-system-one + (license null; 1★), browser-ai (design done), semantic-bookmark. + Star spike: SemIf 1607★ this pass; jevlike 897★. + +Cross-repo addition: (ds) capability kernel vs post-decision +firewall (secrets never in agent; Jev SENSOR; type-safe ≠ +correct); (dt) typed control plane around DSPy, not more LM knobs; +(du) native-probability calibration + fan-out as measurement +economics; (dv) engine owns truth / Jev owns judgment (anti- +soundness-theater with attention≠correctness); (dw) human- +confirmed kill + mapped explanations + identity re-check. + +## Batch #44 (2026-09-19 ~02:43 UTC / ~20:43 Boise 2026-09-18) — domain specialist vs few-shot hosted, decide→policy leftover cascade, ORDER BY calibration≠sortable, GLiFormer wire-compat backend + +Note: `research/notes.md` §60. Docs-only. Folded into PR #2. Not a +second PR. Archer still Watch. No invented metrics. Do not re-fold +§50–§59. + +- **help-er/Domain-jev-maker (Empirical as their RESULTS.md).** + Python MIT. Created 2026-09-19T02:17:54Z; 0★. Independent + CLINC gold, not a Jev teacher-copy. KL 0.168 vs 0.580 + banking; few-shot determinate McNemar n.s. Train specialist + when policy reads p. +- **skiingfalcon/jav-email-cascade (Empirical as README + architecture).** Python; license null. Created + 2026-09-19T02:28:24Z; 0★. Decide→policy→LLM leftover. Noul + 0.5 never rounded. Mock gen-json flat-confidence is *their + mock*. Distinct from dual-process-ai. +- **yodablocks/jev-orderby-bench (Empirical as independent + measurement).** Python MIT. Created 2026-09-19T01:31:53Z; + 0★. Six gates pass. Score ordinal 0.143 weak link; 53-way + 0.99 tie; recodelabs batch-40 fails ranking. Calibration ≠ + sortable. +- **logan-markewich/jeff (Empirical as their RESULTS.md).** + Python; license null. Created 2026-09-19T02:17:25Z; 0★. + GLiFormer-400M `/v1/systemone`. ~6× L4 HTTP / ~24× A10G + direct; AG News 75.5% vs 90.5%. Not a Jev replica. +- MED: hraness/sysone (MIT, TypeScript; loopback gateway; no + weights). + +Cross-repo addition: (dx) specialist vs few-shot as a function +of whether downstream reads p; (dy) Harbor-shaped +decide/policy/LLM leftover with native vs verbalized vs +logprob arms; (dz) ranking family vs calibration family +(ORDER BY; request shape); (ea) wire-compat encoder backend +as a product economics decision, not a quality clone. + +## Batch #45 (2026-09-19 ~03:39 UTC / ~21:39 Boise 2026-09-18) — active-learning triage / don't distill Jev as teacher, evidence-packet explorer, meaning-grep, closed-vote CU, Jev vs MLX PCD Harbor, host-owned waymode, OMP/pi fail-open gates + +Note: `research/notes.md` §61. Docs-only. Folded into PR #2. Not a +second PR. Archer still Watch. No invented metrics. Do not re-fold +§50–§60. Skip empty jev-compactor / laya-jolt. + +- **ThyFriendlyFox/jev-triage (Empirical as README architecture).** + Python MIT. Created 2026-09-19T03:37:23Z; 0★. Accept / teacher / + human. Log full distributions. Do not distill Jev as teacher + (~68% ceiling). Real outcomes stay the targets. +- **jimmyhealer/jev-semantic-explorer (Empirical as their + performance.md, author-run).** Python MIT. Created + 2026-09-19T03:21:01Z; 0★. jevex. 1/8→6/8 SWE-bench Verified + finish n=8 (empty = miss). Packet HitFile 0.233 vs BM25 0.159 + diagnostic, not product KPI. +- **uehaj/jev-semgrep (Empirical as README + judge test).** + JavaScript; LICENSE MIT (GitHub NOASSERTION). Created + 2026-09-19T03:18:28Z; 0★ then. AND/OR/NOT line Nouls; JP↔EN. + Precision 0.94 / recall 0.98 *theirs*. Dedicated fold §86. +- **buluoray/JevOnly (Empirical as README architecture).** Python + Apache-2.0. Created 2026-09-19T03:24:36Z; 0★. Closed-vote; no + planner LLM. 11/43/~$0.014/17s *theirs*. +- **mallahyari/system-one-benchmark (Empirical as their n=50 + table).** Python; license null. Created 2026-09-19T03:35:00Z; + 0★. Jev 84.0% / Brier 0.1096 vs PCD 52% / 0.3884. O(1) ≠ Noul. +- **mossburgh/waymode (Empirical as README + eval suite).** + TypeScript MIT. Created 2026-09-19T02:43:12Z; 0★. Host-owned + handlers. 24/26 + 34/36 *theirs*. Not a self-driving proof. +- **luw2007/omp-jev-extensions (Empirical as README fail + polarity).** TypeScript MIT. Created 2026-09-19T02:55:18Z; 0★. + Fail-open acceptance + route (`confidence: 0`). Contrast + pi-jev-approver fail-closed. + +Cross-repo addition: (eb) training-data VOI / don't distill Jev +as teacher of record; (ec) index-once evidence packets vs grep; +(ed) line meaning-grep AND/OR/NOT; (ee) closed-vote CU with no +planner LLM vs host-owned product surface; (ef) Harbor Jev vs +PCD: O(1) speed ≠ calibrated Noul; (eg) OMP/pi fail-open vs +fail-closed remainder gates. + +## Batch #46 (2026-09-19 ~04:38 UTC / ~22:38 Boise 2026-09-18) — permission vs probability (greenlight), judgment ≠ permission (skill-broker outline), eval integrity / instrument-not-score (dinostomp) + +Note: `research/notes.md` §62. Docs-only. Folded into PR #2. Not a +second PR. Archer still Watch. No invented metrics. Do not re-fold +§50–§61. Watch archive `/workspace/jev-archive/2026-09-18/223856` +absent this VM; live GitHub receipts. Do not treat skill-broker as +a production recipe. + +- **SemetricLabs/omp-greenlight (Empirical as measured traffic + + labelled corpus).** Python MIT. Created 2026-09-19T04:06:43Z; 0★. + 1,013 calls / 10 sessions. Default 40.9% prompts removed; 0 of 94 + unsafe auto-approvals on 140-row corpus. Operator owns thresholds; + plugin never self-tunes. Not a sandbox. +- **adamjralph/skill-broker (Hypothesis / outline only).** Language + null; license null. Created 2026-09-19T04:35:59Z; 0★. + PROJECT-OUTLINE.md authoritative. Code owns grants; Jev never + grants access. Not a production recipe. +- **collapseindex/dinostomp (Empirical as FINDINGS.md + demo card).** + Python; README Apache-2.0 / GitHub NOASSERTION. Created + 2026-08-09T07:59:32Z; 5★. Checks the instrument, not just the + score. 189 findings / 99 against itself. `dinostomp jev` 24-example + demo ECE 0.062 *theirs*. Beside jevals, not a Harbor taskset. +- MED: nrdz-labs/fast-jev-opencode (MIT TS; fail-open OpenCode V2 + port); yikangy873-gif/jev-desktop; MrDiamondBallz/jev-agent-integration; + bohutang/sift; CorieW/JevExplore. + +Census: Awesomejev 488/21644; SemIf 1641 (+13); jevlike 905 (+4); +tracker likes 41 (+1); lastModified unchanged; Laya yes; Blackwood +ABSENT; X MCP flapping (`pages_archived` 0). + +Cross-repo addition: (eh) permission vs probability / operator-owned +safety bar; (ei) judgment ≠ permission (Jev never grants access); +(ej) eval integrity / instrument-not-score (`dinostomp jev` as +if-statement hygiene). + +## Batch #47 (2026-09-19 ~05:40 UTC / ~23:40 Boise 2026-09-18) — constrained optimizer + S1 features (slo-router), privilege ≠ verdict (construct-auto-classifier), attention/VOI never-block (jev-lens) + +Note: `research/notes.md` §63. Docs-only. Folded into PR #2. Not a +second PR. Archer still Watch. No invented metrics. Do not re-fold +§50–§62. Watch archive `/workspace/jev-archive/2026-09-18/234027` +absent this VM; live GitHub + HF receipts. Hunches labeled. Do not +treat the eight-row slo-router demo as a benchmark. Do not treat +swift-jev LICENSE-only as a CLI product. Do not merge +sysone-help/sysone with hraness/sysone. + +- **zeeshan8281/slo-router (Empirical as live analysis *shape* / + negative for sync Jev).** Python; license null. Created + 2026-09-19T05:32:45Z; 0★. Jev features, constrained controller. + Same routes/accuracy; p95 77.93→490.38 ms *theirs*. Fail-open + local features. Exactness raises floor, never overrides + capability. +- **godspede/construct-auto-classifier (Empirical as + certification).** TypeScript Apache-2.0. Created + 2026-09-18T22:27:03Z; 0★. Effect-based shell gate. Jev 0 + dangerous / 975; every chat model leaked 16–104. Privilege ≠ + verdict. Fail-closed. Operator-owned dials. +- **rashedInt32/jev-lens + jev-lens.nvim (Empirical as README + architecture).** JS MIT created 2026-09-19T04:46:01Z; Lua MIT + created 2026-09-19T04:47:11Z; 0★. Never blocks; never edits; + never green unless sure. Attention filter / VOI, not a + permission gate. +- MED: sysone-help/sysone (MIT TS; evaluation-model-first SDK; + not hraness/sysone); ctaxnagomi/INSTRUCT_JEV (HF MIT; 119 rows + 47/51/21; jevals seed); ckaik/swift-jev (MIT; LICENSE-only). + +Census: Awesomejev 488/21644; SemIf 1652 (+11); tracker likes +42 (+1); lastModified unchanged; Laya yes; Blackwood ABSENT; +X MCP flapping. + +Cross-repo addition: (ek) constrained optimizer + S1 features / +never sole hot-path gate (Harbor-style latency cost); (el) +privilege ≠ verdict / effect contracts not tokens; (em) +attention filter / VOI for human review / never blocks / never +green unless sure. + +## Batch #48 (2026-09-19 ~06:39 UTC / ~00:39 Boise) — measurement owns endorsement (jev-packs), Jev supplies evidence / code owns authority (actiongate-jev), ranking ≠ calibration (does-jev-confidence) + +Note: `research/notes.md` §64. Docs-only. Folded into PR #2. Not a +second PR. Archer still Watch. No invented metrics. Do not re-fold +§50–§63. Watch archive `/workspace/jev-archive/2026-09-19/0039` +absent this VM; live GitHub receipts. Hunches labeled. Do not treat +jevassert as a shipped product (404). Do not re-fold +sysone-help/sysone. Do not cite `threat` (n=1). Do not merge +gqgs/laya-onnx with Mattepiu/laya-onnx or local-jev with +us/jev-local. + +- **dtduc-git/jev-packs (Empirical as registry tables / + Harbor-jevals practice).** Python CC0-1.0. Created + 2026-09-19T06:35:36Z; 0★; GitHub size 0. Evidence-gated + `pack.yaml` + `cases.jsonl` + `evidence.md`. Nine packs + `verified` *theirs* on pinned `jev-1.13.0` (citation-support + 800 / 0.919 / 0.022 through banking-intent 150 / 0.840 / + 0.090). `unknown` mandatory. Named runner jevassert **not + released** (404). Measurement owns endorsement. +- **omkarghugarkar007/actiongate-jev (Empirical as README + slogan).** TypeScript Apache-2.0. Created + 2026-09-19T06:20:33Z; 0★; size 0. Deterministic policy owns + ALLOW|REVIEW|BLOCK; Jev is semantic evidence only. Slogan: + "Jev supplies evidence. Code owns authority." Positive + score never overrides a hard fail. Fail-closed + financial/destructive/credential if Jev is down. 500-case + is label-baseline, not accuracy. +- **Adilmp/does-jev-confidence-mean-anything + jevcal + (Empirical as 8,000-judgment audit).** Python; license + null / MIT. Created 2026-09-19T04:51:58Z / 05:44:21Z; 0★ / + 1★. AUC ~0.91; stated ~75% vs human ~10%; ~96% ECE + removed without rank change. Vendor "calibrated" = + rank-correlation, not frequency units. jevcal ~100 labelled + rows. One domain (`civil_comments`). +- MED: gqgs/laya-onnx (complete Laya→browser int8; conversion + smoke; distinct from Mattepiu); kunchenguid/local-jev + (ModernBERT approximation — not equivalence; distinct from + jev-local stub and jeff). + +Census: Awesomejev 488/21644; SemIf **1660** (+8); jevlike +**910** (+5); TypeAR 9; tracker likes 42; lastModified +unchanged; Laya yes; Blackwood ABSENT; X MCP flapping. + +Cross-repo addition: (en) measurement owns endorsement / +evidence-gated question packs; (eo) Jev supplies evidence / +code owns authority (positive p never overrides a +deterministic security failure); (ep) ranking ≠ calibration +/ never hard-threshold raw p as a frequency. + +## Batch #49 (2026-09-19 ~06:50 UTC / ~00:39 Boise remainder) — hot-click CU (ego-jev), Jev judges relevance / code decides structure (jev-compactor), local-rules-then-remainder + anti-self-train (x-reply-filter) + +Note: `research/notes.md` §65. Same hour as §64. Docs-only. +Folded into PR #2. Archer still Watch. No invented metrics. Do +not re-fold actiongate / jev-packs / sysone-help. jev-compactor +was empty skip in §61; content landed. Hunches labeled. + +- **jiangkoumo/ego-jev (Empirical as n=3 medians / README + architecture).** JavaScript MIT. Created + 2026-09-19T06:33:03Z; 0★. Indexed viewport table → + operation+target. Code owns observe/execute/`--until`. + Text model only for type. HN 4.9 s vs 9.7 s; wiki 5.4 s vs + 10.1 s *theirs*. High variance. Not a benchmark. +- **edwardyen724-g/jev-compactor (Empirical as one-session + bench).** TypeScript MIT. Created 2026-09-19T03:31:54Z; + 1★. Was empty skip. Jev judges relevance; code decides + structure; never rewrite. 64.5% / 366 ms / $0.0004 / 0 + hallucinated / 4 of 4 vs Sonnet 96.2% / 1 invented path. + Dual fail polarity. Regex floor independent of Jev. +- **zhuyansen/x-reply-filter (Empirical as README + small + e2e).** JavaScript MIT. Created 2026-09-19T06:07:18Z; 0★. + Local rules first; remainder Nouls; never auto-train on + the model's own hides. 3-sample e2e 0.90/0.93 vs 0.08/0.10. + +Census: same as batch #48 (Awesomejev 488/21644; SemIf 1660; +jevlike 910; tracker likes 42). + +Cross-repo addition: (eq) hot-click CU / indexed element +table / S1 on click path; (er) Jev judges relevance, code +decides structure / never rewrite; (es) local rules first +then remainder / never auto-train on the model's own hides. + +## Batch #50 (2026-09-19 ~07:53 UTC / ~01:47 Boise) — control-plane combinators, receipts-not-leaderboard, skill VOI, OOD/AUC≠ECE, frontier-100, turnstile, jevmlx + +Note: `research/notes.md` §66. Docs-only. Folded into PR #2. +Archer still Watch. No invented metrics. Atlas axis already +§49; skillranker existence already §7 (hook polarity +**corrected** to fail-open); jevmlx reconstruction already +§1. Hunches labeled. + +- **voidning/decision-combinators (Empirical as README + architecture).** TypeScript. README MIT / GitHub license + null. Created 2026-09-19T07:25:16Z; 0★. Then / Gate / + Vote / Cascade / Weighted. Not literal AND/OR. No + measurements. +- **Zaious/jev-capability-atlas (delta: eval-integrity + cluster).** 10★ this pass (was 0). Receipts not + leaderboard; type-safe ≠ correct. Do not rehash history + suite. +- **Dicklesworthstone/skillranker (Empirical as README; + correction).** Rust; 52★. Two-pass + none-of-these. + Claude hook fail-open (quiet exit 0). VOI over skill + library. +- **scienthoon/jev-ood-calibration (Empirical as 900-ticket + + 3 public benches).** MIT. Created 2026-09-19T07:33:22Z. + Synthetic ECE 0.107 = 4.4× floor; priority 44.7% / mean p + 0.74 / T 3.40; boolean T 0.66. AUC ≠ ECE. +- **softpudding/jev-frontier-100 (Empirical as exploratory + 100×3).** MIT. Jev 77.0%; Qwen3.5 4B/2048 96.7%; 4B off + 56.0%. Not preregistered. Not a ceiling. +- **zyphr-labs/turnstile (Empirical as README + architecture).** Apache-2.0. Experimental alpha. Policy + first; Jev remainder; replay. Missing Jev → Review. + Actiongate-class. +- **bnsd55/jevmlx (Empirical as README library).** MIT; + 28★. One-pass schema→JSON+probs. Softmax ≠ Noul. No + local leaderboard yet. +- MED: chopratejas/invalidate (5★; 0 of 157 false + invalidations); yottayoshida/jev-intent-review (under + construction); shubhangi013/prune-review (22-run cost + 1.18% with 305% outlier). + +Census: not re-derived this hour (last §64/§65: Awesomejev +488/21644; SemIf 1660; jevlike 910; tracker likes 42). + +Cross-repo addition: (et) combinators / System One as +control plane; (eu) receipts not leaderboard / type-safe ≠ +correct jaggedness; (ev) skill-library VOI / abstention; +(ew) OOD / AUC ≠ ECE / sign by type; (ex) thinking-budget +bake-off; (ey) turnstile evidence≠authority + replay; +(ez) MLX one-pass replica economics. + + +## Batch #51 (2026-09-19 ~08:45 UTC / ~02:38 Boise) — TLA+ compose, SEAL coverage ledger, skill-broker sibling, sureness, JevBench v1.1, CI typed gate, Codex MCP + +Note: `research/notes.md` §67. Docs-only. Folded into PR #2. +Archer still Watch. No invented metrics. Do not re-fold §66 +HIGH except sibling contrast. skill-broker outline already +§62. Hunches labeled. + +- **copyleftdev/jev-labs (Empirical as TLA+ + 1,080 golden + chaos table).** Python MIT. Created 2026-09-19T08:07:12Z; + 0★. Never confidently wrong. 0 wrong golden; severe 46 + escalate. TLC 1,049,750 / 0 errors. Synthetic, not + clinical. +- **Reasonofmoon/seal (Empirical as README + BEYOND-JEV.md).** + Python MIT. Created 2026-09-19T08:10:05Z; 0★. No seal, no + advance. coverage.path auto|code|human|escalate. Mint ≠ + product brain. +- **adamjralph/skill-broker (delta: sibling table).** Still + outline; README restates grants-in-code. Contrast + turnstile / skillranker. +- **adarc8/how-sure-is-jev (Empirical as 60-q library).** + Python MIT. Choice confidence = max_prob. Most generous + metric. +- **fstandhartinger/jevbench (Empirical as v1.1 artifact).** + Python MIT. Unofficial. Main Score 0.6/0.2/0.2. Jev + 1.13.0 87.6. Calibration not scored. +- **NemanjaManic/ci-gatekeeper-bot-jev (Empirical as + own-repo latencies).** package.json MIT / GitHub SPDX + null. 504–629 ms. Conservative default. +- **teempai/jev-in-codex (Empirical as README).** TypeScript + MIT. Codex MCP. Ranking unbenchmarked. Lexical fallback. + +Census this hour (user-provided): Awesomejev 488/21644; +SemIf 1683 (+11); jevlike 923 (+5); tracker likes 43 (+1). +Archer still NOT landed. X MCP flap continues. + +Cross-repo addition: (fa) never confidently wrong / TLA+ +compose; (fb) no seal no advance / coverage ledger; +(fc) sureness / max_prob is generous; (fd) JevBench +calibration not in Main Score; (fe) CI typed gate before +expensive review; (ff) Codex MCP host adapter. + +## Batch #52 (2026-09-19 ~09:55 UTC / ~03:38 Boise) — attention redirect, pre-send views, tools≠use, observational memory, open-Jev class, physical-world S1 + +Note: `research/notes.md` §68. Docs-only. Folded into PR #2. +Archer still Watch. No invented metrics. Do not re-fold §67 +HIGH except sibling contrast. Hunches labeled. + +- **muse0509/jev-preflight (Empirical as owner-run hook + smoke).** Go MIT. Created 2026-09-18T13:54:37Z; 1★; + v0.1.0 Public Beta. Eight risk axes; assist=one + reinspect; fail-open; 0.85 uncalibrated. Claude Code + 2.1.267: no-key fail-open PASS; key-enabled exactly one + continuation. Not a merge blocker. +- **godspede/construct-auto-classifier (delta).** Cert + unchanged: Jev 0 dangerous / 975; $0.047/1k. Landed- + script trust; headless ≠ auto-approve. +- **edwardyen724-g/jev-compactor (Empirical as product-arm + table).** Later bench 73% / 350 ms / 4 of 4 *theirs*. + 30–250× cheaper. §65 64.5% is vs-Sonnet. +- **dizk/jev-lens (Empirical as 500-trajectory).** + TypeScript MIT. Created 2026-09-18T08:16:12Z; 0★. 79% + fewer tokens (11.6M→2.4M). Post-send prune +17% cost. + Distinct from rashedInt32/jev-lens. +- **Dharundp6/jev-carryforward (delta).** recall 0/4. + tools≠use. SessionStart > hoping. +- **willfish/pi-observational-memory-jev (Empirical as + README).** TypeScript MIT. Created 2026-09-18T06:42:35Z; + 0★. Keep/kind verbatim; model-free compact. +- **genai-craft/openvons (Empirical as their docs).** + Python Apache-2.0 LICENSE / GitHub SPDX NOASSERTION; 7★. + Independent open-Jev class. JevPick 3.2–4.8×. Wire-compat + ≠ replica. +- **AboveColin/HA-Jev (Empirical as README + measurements).** + Python MIT; 17★. First real card. Not for locks/heaters. +- **kylemclaren/jevql (architecture note).** Judgment + outside the store. CLI; DB never sees jev(). +- **bohutang/sift (short bullet).** ~$0.00003/post. + +Census this hour (user-provided): Awesomejev 488→561 +(+73, agent tooling 87→107); SemIf 1683→1704. Archer +still NOT landed. + +Cross-repo addition: (fg) attention redirect not merge +blocker; (fh) compress-before-first-send; (fi) tools≠use +/ SessionStart; (fj) observational keep/kind; (fk) +open-Jev class / wire-compat ≠ replica; (fl) physical- +world S1 / not for locks; (fm) judgment outside the store; +(fn) landed-script / headless≠auto-approve. + +## Batch #53 (2026-09-19 ~10:55 UTC / ~04:39 Boise) — digital-design combinators, VOI cache, skill-routing Harbor, zeroshot displacement, typed handoff + +Note: `research/notes.md` §69. Docs-only. Folded into PR #2. +Archer still Watch. No invented metrics. Do not re-fold §68 +HIGH except sibling contrast / combinators rename. Hunches +labeled. + +- **voidning/jev-combinators (delta: rename + extended five).** + Same repo as decision-combinators (§66). Package MIT / + GitHub SPDX null. Router / Loop / Retry / Fallback / + Memory. Digital-design metaphor ≠ literal AND/OR. No + measurements. +- **kushals256/jevcache (Empirical as n=100 live eval).** + TypeScript MIT. Created 2026-09-19T09:40:18Z; 0★. Jev + 0 FP / precision 1 / recall 0.38 / fpr 0 vs Jaccard@0.35 + fpr 0.48; $0.00174 *theirs*. Fail-open. +- **iamdin/pi-jev-skill-bench + pi-jev-skill-suggestion + (Empirical as README harness; no live Jev numbers).** + TypeScript MIT. 43 gold; roster 50–500. No-key no-op. +- **zhuyansen/jev-zeroshot-vs-bert (Empirical as 7-set + table).** Python MIT. +0.05–+0.13 vs DeBERTa-c; + contamination 0.901 vs `-c` 0.763; ≈230 / >2048 labels; + DiD 0.035 vs 0.112 *theirs*. +- **shitianfang/jev-handoff (Empirical as README + 40 tests).** + TypeScript MIT. Alpha v0.1. Gate never grants. Fail-open. + Vercel drops confidence. +- **ThinkyMiner/Winnow (Empirical as unreviewed goldens).** + TypeScript MIT. 80%/90% *theirs*. Distinct from + kevinpita/winnow. +- **rsdkrasen/hermes-jev-router (Empirical as README + + offline pytest).** Python; license null. WHETHER/HOW/WHAT. + Community plugin, not vendor. +- **mleyvaz/jev-typed-evaluation-collapse (Empirical as + NCML field note v0.3).** License null. Noul collapse vs + named Choice p=1.0 vs binary red 0.67–0.85 *theirs*. +- **DowLucas/browser-jev (Empirical as README).** TypeScript; + license null. Sample-from-distribution. Playwright + executes, Jev chooses. +- **IamBusy/OpenJev (Empirical as RESULTS.md).** Apache-2.0. + 45/60 vs v0.2 39/60. `/v1/decide` ≠ TypeSafe drop-in. + Distinct from hraness/sysone runners. +- **dddanielliu/semif-serve (Empirical as latency table).** + pyproject MIT / GitHub SPDX null. 1164 vs 178 ms *theirs*. + Wire-compat ≠ replica. +- Toolbelt notes: win4r/jev-security-scan (MIT; not a cert); + bojansandhaus/jev-decisions (MIT; 1★; advice never stops + commands); TeoMastro (license null; summary.md 404). + rh-guard owns reward-hack. +- **ctaxnagomi/DGUI_HYPERMEM-JEV (feedstock schema).** HF MIT; + 6 rows. Sibling INSTRUCT_JEV. + +Census this hour (user-provided): Awesomejev flat 561/27007; +tracker likes 43→45; SemIf 1714 (+10); jevlike 926 (+3). +Archer still NOT landed. + +Cross-repo addition: (fo) digital-design combinators / +extended five; (fp) VOI cache admission / 0 FP; (fq) +Harbor skill-routing roster-size; (fr) zeroshot vs BERT +displacement / contamination DiD; (fs) typed baton never +grants / inverted loop; (ft) worth-your-attention VOI / +qualify Winnow owner; (fu) WHETHER/HOW/WHAT; (fv) conflict +≠ ignorance / named Choice escape; (fw) Playwright +executes Jev chooses; (fx) OpenJev `/v1/decide` ≠ drop-in; +(fy) SemIf runoff wire; (fz) decision-as-memory flywheel. + +## Batch #54 (2026-09-19 ~11:55 UTC / ~05:46 Boise) — record/replay CI, BBQ, decider≠executor, sentence-as-rule lint, open replica substrates + +Note: `research/notes.md` §70. Docs-only. Folded into PR #2. +Archer still Watch. No invented metrics. Do not re-fold +§50–§69 HIGH except sibling contrast / jevassert landing / +prune-review, intent-review, laya-jolt, local-jev deltas. +Hunches labeled. Meanblock/JEV-CPU **404**. + +- **dtduc-git/jevassert (Empirical as README + Action @v0).** + Python Apache-2.0. Created 2026-09-19T06:48:15Z; 0★; size + 64. **LANDED** (was 404 §64). Record/replay CI: + accuracy/ECE/Brier/cost/latency offline from recordings. + Exit 0/1/2. McNemar. Backend-neutral. +- **dtduc-git/jev-packs (delta: runner pairing + 2,990 + matrix).** CC0-1.0; size 0→458. Jev/Sonnet 5 accuracy + tie Δ≤0.018; Jev better calibrated 7/9; ~250× cheaper + *theirs*. sms-spam this-pass 0.953 / ECE 0.040. +- **chenmingtang830/jevarena (Empirical as runnable + harness; not findings).** TypeScript Apache-2.0; size + 724. Failure-finding not a leaderboard. **≠** + meetr1912/jev-arena. Always qualify owner. +- **simonmesmith/jev-bbq-experiment (Empirical as full + 58,492).** R; license null; size 5365. 97.28%; bias + 0.04/0.34; $0.3429 / 7.75 min *theirs*. Not a general + bias cert. +- **thomasbrueggemann/jeffrey (Empirical as README).** + TypeScript MIT; size 71. Decider ≠ executor. Pick ≠ + fill. Risk≥0.5 pause. Stuck ladder 2 Jev / 0 steps. +- **mizchi/jevlint (Empirical as 13/15 corpus).** + TypeScript MIT; size 338. ast-grep × `ask:`. 13/15 + 1.00/1.00 *theirs*. **≠** huntedman/JevLint. Always + qualify owner. +- **shubhangi013/prune-review (delta: size/star).** + 22-run numbers unchanged 1.18% with 305% outlier + *theirs*. ~20% cost target. Cost not quality. +- **yottayoshida/jev-intent-review (delta: CLI).** Dual + MIT/Apache-2.0; size 81. VERIFIED/VIOLATION/UNKNOWN. + Empty search ≠ proof. Action not written. +- **Eran-BA/Jev_from_GLiNER2 (spec-only).** Size 0. + Interface ≠ replica. Distinct from jeff GLiFormer. +- **bokuweb/grande (Empirical as JGLUE + isolation).** + Rust; license null; 1★; size 583. JNLI 0.614 ECE + 0.088 T=2.81; JCQA 0.853; 270M 0.710/0.710 *theirs*. + Softmax ≠ Noul until T. +- **jlt-commons/laya-jolt (delta: content landed).** + Clojure Apache-2.0; size 5112. Byte parity vs Python + system_one. ~1e-7 last-digit drift. +- **leesk212/JEV-CPU (Empirical as PoC).** Python MIT; + 1★. SemIf CPU + UI. Meanblock/JEV-CPU 404. +- **kunchenguid/local-jev (delta: measured).** 136 + checkpoints *theirs*: done 30% / shape 57% / r −0.06. + Not equivalence. +- Toolbelt notes: actiongate-jev slogan already §64. + **Nyarlathoteppppp/pi-heed (Empirical as 79-session + bench).** TypeScript MIT; 3★; size 621. Recall 98.5% + / false block 0.0% / $0.000058 *theirs*. Jev never + writes policy. +- MED: david-j-lustig/system-one-responsible-ai size 0 + framing stub. + +Census not re-derived this hour (last §69). Archer still +NOT landed. + +Cross-repo addition: (ga) record/replay CI / calibration+ +cost first-class; (gb) evidence-gated packs have a +runner / 2,990 matrix; (gc) failure-finding arena ≠ +leaderboard / qualify jevarena owner; (gd) BBQ +stereotype/uncertainty/cost not a bias cert; (ge) +decider≠executor / pick ≠ fill; (gf) sentence-as-rule +lint / qualify jevlint owner; (gg) VOI hunk prune / +safety escarpment; (gh) whole-repo intent UNKNOWN; +(gi) GLiNER2 spec ≠ replica; (gj) open replica +substrates (grande / laya-jolt / JEV-CPU / local-jev); +(gk) persist constraints across compaction. + +## Batch #55 (2026-09-19 ~12:50 UTC / ~06:43 Boise) — SGR-judge Harbor contract, control-plane productization, never-generates, recipes+life feed, NAR claim-audit + +Note: `research/notes.md` §71. Docs-only. Folded into PR #2. +Archer still Watch. No invented metrics. Do not re-fold +§50–§70 HIGH except sibling contrast. Hunches labeled. +IPECTER/jev-context-pruner **empty**. + +- **slavadubrov/jev-judge-bench (Empirical as frozen + contract; not a quality score).** Python; README MIT / + GitHub SPDX NOASSERTION. Created 2026-09-19T12:20:43Z; + 0★; size 0 lag. SLA-150. Invalid = FN. 21 offline tests. + Canaries ≠ quality. **No headline yet.** ≠ jevarena / + jevbench. Always qualify owner. +- **IPECTER/jev-context-pruner (empty skip).** Description + only; contents 409 empty. Sibling contrast vs + fast-jev-compaction / jev-compactor / dizk/jev-lens. +- **shitianfang/jev-use (Empirical as 95-call card).** + TypeScript MIT v0.4.1; size 196. p50 220 ms; 186 vs + 2,672 ms; gate 12/12; Vercel margin 0.4; 17/20 then + 0/20 *theirs*. Fail-open. ≠ jev-ultrafast. Same author + as jev-handoff. +- **goodruizhan/pi-jev-control (Empirical as README).** + TypeScript; license null; v0.3.0 private. Control plane. + GUI never force-click; compaction never writes session. + No live quality numbers. +- **florian-hoenicke/jev-gpt (Empirical as README demo).** + Python; license null. Never free-generates. ~400 calls / + 75 s / 2¢ *theirs*. Distinct from jeffrey. +- **nexibeo/jev-cookbook (Empirical as recipe samples, not + benches).** JavaScript MIT; 1★; size 165. 425 calls / + $0.015; browser 5/6 *theirs*. 16–36 handmade. +- **fengyiqicoder/jevfeed (Empirical as README + 17 + tests).** JavaScript MIT. No social graph. One request + per batch of ten. Distinct from both Winnow products. +- **Heman10x-NGU/openJev-verdict-2.0 (claim-audit, not + endorsement).** Python; README Apache-2.0 / GitHub SPDX + NOASSERTION. 77.10%/0.0636/0.0144 *theirs* unverified. + Open PR #1: throughput≠latency; Laya parity; like-for-like + ECE. ≠ IamBusy/OpenJev. + +Census not re-derived this hour (last §69). Archer still +NOT landed. + +Cross-repo addition: (gl) Harbor SGR-judge contract / +canaries ≠ quality / no headline yet; (gm) hand no-text +steps / gateway drops confidence; (gn) Pi control plane / +GUI never force-click; (go) generation as tree of Choices; +(gp) recipe atlas / samples not benches; (gq) +personal-history ranking without a social graph; (gr) +competing NAR dual-channel ECE as claim-verification; +(gs) empty compaction-proxy skip. + +## Batch #56 (2026-09-19 ~13:52 UTC / ~07:49 Boise) — 1-token logprob ≠ Noul, replica engine, Elixir SDK, commit/jevex VOI, inbox, laya-multilingual + +Note: `research/notes.md` §72. Docs-only. Folded into PR #2. +Archer still Watch. No invented metrics. Do not re-fold +§50–§71 HIGH except sibling contrast. Hunches labeled. +Local archive `/workspace/jev-archive/2026-09-19/0747/` +missing; live GitHub + HF ~13:52 UTC. IPECTER/jev-runway +**LICENSE-only empty**. GitHub `size: 0` lags on repos +that have files. + +- **taku-me/chakuho (Empirical as 336-case GUI).** Python + MIT. 1-token logprob `/v1/systemone`. Softmax ≠ Noul. + Coverage ≠ correctness. 27B 95%/92% vs Jev 89%/82% + *theirs*. `__none__` 97% vs 8B 10%. Cousin jevify / + TypeAR / pcdServer / jevmlx. +- **zerodegress/jevinf (Empirical as README speedup).** + Python MIT; ≥3.14. 2.57×/2.27× 100% argmax *theirs*. + MPS only. Wire-compat ≠ replica. +- **phiat/typesafe-elixir-sdk (Contract).** Elixir MIT; + 1★. Unofficial HTTP client. ≠ dannote/jev OTP peer. +- **jimmyhealer/jevex (delta of §61).** Rename of + jev-semantic-explorer. n=16 160s→69s / $8.74→$3.13 / + 16/16 *theirs*. Keep n=8 1/8→6/8. +- **yodablocks/commitjev (Empirical as 13 labelled).** + Python MIT. Middle band never rounded. 0 false on 5 + clean *theirs* (small control). Same owner as + jev-orderby-bench. +- **Mrmimee/hermes-plugin-jev (identity lock).** README + MIT / GitHub SPDX null. Agnes 3.0 Flash branded as + Jev. ≠ hermes-jev-router. +- **dev-willbird1936/pi-jev-compact (Empirical as README + + latency).** TypeScript MIT. ≠ vava-nessa/pi-jev- + compaction. Fail-open if <25% saved. +- **IPECTER/jev-runway (empty skip).** LICENSE only. +- **Milo318/mailordinal (Empirical as README).** + TypeScript MIT. Decision-native inbox. Humans own + ambiguity. +- **aarora79/jev-samples (toolbelt).** License null. One + sample. Not a bench. +- **shaharia-lab/jev-cli (not ready).** Rust Apache-2.0 + OR MIT; 0.0.0; 17 issues. ≠ jevql. +- **convaiinnovations/laya-multilingual (Empirical as + Hub card).** Apache-2.0; 322M; 9 likes. MASSIVE + 0.366/0.387 vs 0.227/0.733. Khmer 0.000@0.952. Ships + uncalibrated. +- **mobarmg/jev-schema-scorer-deberta-v3-large (Hub; + GitHub 404).** MIT. v2 Choice 0.841. Peaked ranking ≠ + calibration. +- Skip: jev-routing host-surface delta (Cursor/Devin); + laya-typed-decisions; open-jev-laya-bench / + jev-tree-choice-cap / INSTRUCT_JEV HF 401; jevlogs + GitHub 404 + HF 401. + +Census not re-derived this hour (last §69). Archer still +NOT landed. + +Cross-repo addition: (gt) 1-token logprob ≠ Noul / +coverage ≠ correctness; (gu) open replica engine +argmax-parity; (gv) unofficial Elixir SDK ≠ OTP peer; +(gw) jevex n=16 files-to-read VOI; (gx) commit +attention≠verdict / middle band; (gy) branding ≠ +backend (Agnes as Jev); (gz) Pi compact ≠ compaction; +(ha) decision-native inbox; (hb) unofficial CLI not +ready; (hc) multilingual Laya confident-wrong OOD; +(hd) schema-conditioned peaked ranking; (he) second +IPECTER empty skip. + +## Batch #57 (2026-09-19 ~14:37 UTC / ~08:37 Boise) — productized System One HTTP, escalate-under-threshold, silent FALLBACK + +Note: `research/notes.md` §73. Docs-only. Folded into PR #2. +Skip Archer. No invented metrics. Hunches labeled. +User-provided SIGNAL + live README `6ae8bda` / +eval/README `2009d1f`. + +- **mrmps/classifier-dev (Empirical as README + + eval/README).** MIT; **185★**; https://classifier.dev. + Public zero-shot HTTP; no key. Jev primary + (`src/jev.ts`); LLM fallback only. 400 headlines + **650 ms** *theirs*; packing 100 = one-at-a-time. + Smart: single-label <0.7; multi-label ignores. + Emotion ≥0.9 → 82% / <0.5 → 29%. gemini-3.8-flash + 87.5→90.0 / 61.8→63.7. Multi-label F1 **0.887** / + **230 ms** vs cascade **0.799** / 1.5 s (eval 232 ms; + AG News **87.7%** vs 82.0%; emotion **60.5%** vs + 57.0%). n=7 train-on-test; ~0.03 coin flip. + granite-4.0-h-micro F1 **0.546** vs advertised + ~**0.800**; digest marks `FALLBACK`. Distinct from + ask-jev-ai. rh-guard owns the gate card. Life/business + (spam/inbox/feedback). + +Census not re-derived. Archer still NOT landed. + +Cross-repo addition: (hf) productized System One HTTP / +label+p public contract; (hg) escalate-under-threshold / +multi-label ignores tier; (hh) vs_jev tracked JSON not +transcription; (hi) silent FALLBACK / granite 0.546 vs +advertised 0.800; (hj) classification API for +spam/inbox/feedback. + +## Batch #58 (2026-09-19 ~14:48 UTC / ~08:48 Boise) — choxos/jev-reviewer systematic-review pointer (delta of §48) + +Note: `research/notes.md` §74. Docs-only. Folded into PR #2. +Skip Archer. **≠** egma-ai/jev-reviewer. No invented metrics. +Hunches labeled. User-provided SIGNAL + live README `d0220a1`. + +- **choxos/jev-reviewer (Empirical as README; delta of + §48).** MIT; **12★**; https://jevreviewer.xera.ac. + Pointer-not-generator at Cochrane/PRISMA scale. Jev + picks line ids; code copies verbatim. Two-pass Choice + + Noul (quotes ≥ 0.5 *theirs*). *Not found* is an + answer. Human tick never overwritten. 18-q template + **4.6 s / $0.0101** *theirs* (spot check, not a + validation study). + +Census not re-derived. Archer still NOT landed. + +Cross-repo addition: (hk) two-pass relative Choice + +absolute Noul; (hl) *Not found* first-class; (hm) human +check as productized judgment; (hn) evidence-synthesis +as class application beyond SWE; (ho) name lock vs +egma-ai. + +## Batch #59 (2026-09-19 ~14:56 UTC / ~08:56 Boise) — githubnext/localjev wire-compat ≠ logit-equiv + +Note: `research/notes.md` §75. Docs-only. Folded into PR #2. +Skip Archer. **≠** kunchenguid/local-jev. No invented metrics. +Hunches labeled. User-provided SIGNAL + live README `39939e6` +/ eval `418cae7`. + +- **githubnext/localjev (Empirical as README + + evaluation-results-2026-09-18).** MIT; **261★**; + GitHub Next. Bun `POST /v1/systemone` on DiffusionGemma + via Chat Completions; TypeSafe SDK drop-in. Prompted + JSON → validate/retry → normalize + entropy confidence. + **Not** razorback16 structured-read logits. **Not** + IamBusy/OpenJev `/v1/decide`. Bake-off *theirs*: 1,200 + req / ~23.5 min / M5 Max. Short macro Qwen3.6 **76.7%** + / Gemma 4 26B-A4B **75.0%** (SST-5 MAE **0.533**) / + DiffusionGemma **74.2%**. Qwen vs Gemma 26B = 2/120. + Do not treat as calibrated. LM Studio cannot load + DiffusionGemma (18 Sep 2026). + +Census not re-derived. Archer still NOT landed. + +Cross-repo addition: (hp) wire-compat ≠ logit-equiv / +prompted JSON ≠ structured read; (hq) GitHub Next +institutional local `/v1/systemone`; (hr) Harbor-shaped +1,200-req bake-off with caveats; (hs) LM Studio runner +gap / structured-read primitives for OpenJev parity; +(ht) name lock vs kunchenguid/local-jev. + +## Batch #60 (2026-09-19 ~15:07 UTC / ~09:07 Boise) — NandhaKishorM/laya packaging, not a new species + +Note: `research/notes.md` §76. Docs-only. Folded into PR #2. +Skip Archer. No invented metrics. Hunches labeled. +User-provided SIGNAL + live README `f12882b`. +**≠** TypeSafe `/v1/systemone`. **≠** githubnext/localjev. + +- **NandhaKishorM/laya (Empirical as README; delta of + Hub Laya §18 / §42 / §46 / §72).** Apache-2.0; + **710★**; 62 forks; Python. PyPI + `Router` over + convaiinnovations/{laya, laya-multilingual, + laya-typed-decisions}. T4 *theirs*: 1q **32.8 ms** / + 10q **72.3 ms** (~7.8× vs Jev p50 236–276 ms cited + third-party). Post-T ECE **0.081** vs Jev **0.246**; + raw 0.213 vs 0.144. Banking77 **0.425** vs **0.870** + (77 vs 72; ~3–4 tok/label). typed-decisions **0.766** + fine-tune (base 0.362/0.342 vs majority 0.461). + Soft-acc 0.471 vs 0.580. Khmer **0.000@0.952** — + Router because gating cannot catch. 0.85 still soft. + Jev rows unpublished-here. + +Census not re-derived. Archer still NOT landed. + +Cross-repo addition: (hu) packaging ≠ new species / +Router script-before-p; (hv) where Jev leads +(Banking77 / soft-acc / raw ECE); (hw) post-T ECE ≠ +raw ECE / latency; (hx) 0.85 still soft / Khmer OOD +productized; (hy) vs-Jev third-party unpublished-here. + +## Batch #61 (2026-09-19 ~15:14 UTC / ~09:14 Boise) — @airesearch12 Benchmark Heaven openjev census (tweet, not scores) + +Note: `research/notes.md` §77. Docs-only. Folded into PR #2. +Skip Archer. No invented metrics. Hunches labeled. +User-provided SIGNAL_4b0c + live X MCP this pass. +**≠** fstandhartinger/jevbench v1.1. Do **not** paste +live board ranks. + +- **@airesearch12 status/2101259522933186879 + (Empirical as tweet).** Florian S / Benchmark Heaven. + Named ~18 openjevs including GLiNER2 and routers + (class-boundary). Engagement ephemeral (SIGNAL + ~417/9/3; this pass 564/15/5). Incomplete vs Laya / + githubnext/localjev / kev / TypeAR / openvons / + chakuho / jevinf / grande / laya-jolt / blackwood / + classifier-dev. Watch + https://benchmarkheaven.com/jev-models. Do not copy + Stripe. + +Census not re-derived. Archer still NOT landed. + +Cross-repo addition: (hz) external census ≠ scored +bake-off; (ia) GLiNER2+routers class-boundary; (ib) +incomplete census vs watch; (ic) Harbor honesty watch +(cal / cost / latency / silent fallback). + +## Batch #62 (2026-09-19 ~15:24 UTC / ~09:24 Boise) — JevBench v1.2 scored board (measurement, not a hit list) + +Note: `research/notes.md` §78. Docs-only. Folded into PR #2. +Skip Archer. No invented metrics. Hunches labeled. +User-provided SIGNAL_988e + live board + GitHub README +`bf1e79ba` / RESULTS-v1.2 `fdfab1a2` / HEAD `27ed3d6c`. +**≠** tweet census §77 **≠** v1.1 87.6. + +- **JevBench v1.2 (Empirical as board).** Protocol + `jevbench::v1.2`; 534 decisions; geo-mean I/C/S/K + 25% each. Jev 1.13.0 **75.3**; SemIf **74.6** (−0.7); + OpenJev razorback16 67.6 *theirs*. Luna I **96.8** + rank **#7**. Cal **ON** rank. Option-order 72%→21%. + Self-host ×2 assumption; many costs est. Laya + absent (gap). GLiNER2 mapping-excluded. Qwen3.8 27B + Chutes TEE **≠** Archer. Do not copy Stripe / CLI. + +Census not re-derived. Archer still NOT landed. + +Cross-repo addition: (id) geometric-mean product / +weak axis dominates; (ie) cal now ON rank (delta +from v1.1); (if) weight sensitivity; (ig) +option-order 72→21; (ih) instruction models +class-boundary; (ii) Harbor honesty ×2/est.; (ij) +Laya absent gap / Qwen3.8 27B ≠ Archer. + +## Batch #63 (2026-09-19 ~15:37 UTC / ~09:37 Boise) — hourly 0842 already-folded watch (apply, don’t dump) + +Note: `research/notes.md` §79. Docs-only. Folded into PR #2. +Skip Archer. No invented metrics. Hunches labeled. +**Do not re-card** §73–§78. Leftover thin is skip. +**≠** a hit list. + +- **Apply-the-five (recipe).** Wire-compat ≠ logit-equiv + (githubnext/localjev §75); productize label+p + mark + FALLBACK (classifier-dev §73; granite 0.546 vs 0.800 + *theirs*); packaging ≠ new species / script-before-p + (NandhaKishorM/laya §76); pointer-not-generator + two-pass (choxos/jev-reviewer §74 ≠ egma-ai); external + census ≠ scored bake-off / geo-mean I/C/S/K (§77+§78). +- **Skip thin.** uehaj/jev-semgrep already §61 (dedicated now §86); JEValuate + 0★ ≠ ElshinQ/jevaluate; jevspeak 1★ (jev-gpt cousin); + fable-jev 1★; jev-model-router already §77. +- **Soundness-theater skip.** totally-tim/jev-gate 0★ + MIT; connectedGraph/claude-jev-warden 1★ MIT. Soft + Noul ≠ merge seal. **≠** jev-gateway / MongLong0214/jev-gate + / jev-gate-student-b. + +Census not re-derived. Archer still NOT landed. + +Cross-repo addition: (ik) already-folded hourly as a +recipe not a dump; (il) hard-gate Noul as PR/quality +is soundness theater. + +## Batch #64 (2026-09-19 ~15:50 UTC / ~09:50 Boise) — khordoo/jev-reflex-autonomy-lab delta of §46 + +Note: `research/notes.md` §80. Docs-only. Folded into PR #2. +Skip Archer. No invented metrics. Hunches labeled. +Quote README. Not a hit list. **Do not re-card** the +§46 one-liner. User SIGNAL_2f18 (~08:48 Boise; ★6) ++ live this pass **7★** / 1 fork; license null; +README SHA `130987c9`; ARCHITECTURE SHA `48da0769`; +HEAD `e3297ebe`. Demo watch-demo.html. + +- **S1 never stalls / S2 one-use advisory.** Jev keeps + steering; S2 does not fly and does not grant. +- **Consumption telemetry.** Purple confidence = + that decision used returned S2; purple S2 bar = + arrival; red = fail. Arrival ≠ used. +- **Local controller ≠ localjev.** Rule-based built-in + vs hosted `jev-latest`. 20% starting gate *theirs* + still soft; does not start a mission. Seed = + geometry ≠ async replay. +- **No pixels.** Application contracts ≠ TypeSafe SDK; + confidence ≠ selected probability; physics owns + collisions. README GLM 5.3 vs ARCHITECTURE + muse-spark-1.3-contributor — quote both *theirs*. + +Census not re-derived. Archer still NOT landed. + +Cross-repo addition: (im) escalate-without-stall; +(in) mixed-initiative consumption mark; +(io) Local-vs-Live reflex A/B / Local ≠ localjev; +(ip) seed≠replay / 20% still soft / no-pixels +class discipline. + +## Batch #65 (2026-09-19 ~16:05 UTC / ~10:05 Boise) — awlevin/typesafe-computer-use productized OCR+AX CU + +Note: `research/notes.md` §81. Docs-only. Folded into PR #2. +Skip Archer. No invented metrics. Hunches labeled. +Quote README. Not a hit list. **Do not re-card** the +census one-liner ($0.0002/step; rebuild pixel-free +reasoning). User SIGNAL_0e9d (~08:49 Boise; MIT; +★419) + live this pass **427★** / 24 forks / 7 issues; +README SHA `369f4a6a`; HEAD `cc7b5066`. + +- **Productized observe→score→act.** OCR+AX → numbered + items → hosted TypeSafe Choices → code clicks/types. + Same hole as jev-ultrafast / gliner2-ultrafast / + cua-s1 / Stagehand / ego-jev. Backend here is hosted + Jev, **not** GLiNER2 and **not** Cua-S1. +- **Mutually exclusive action set.** Overlapping + options read as doubt (confidence is concentration). + Split `kind`/`item`/`site`/`offscreen` as VOI / noise + control. +- **Decision ≠ answer-reader.** Never ships a screenshot + to frontier for the *decision*. The one-shot **answer** + writer may receive the capture. Not omni. Skip Archer. +- **Writer/decider + soft Noul.** Writer only for free + text. Post-type 0.5 and `--min-confidence` 0.4 still + soft. AX bonus never sole (Spotify 0 *theirs*). + `done` ≠ verified success. +- **Harbor-shaped one-screenshot table.** $0.0002 vs + Opus $0.032 (155×) *theirs*; dates.py caveat. Not a + taskset. **≠** jev-ultrafast **≠** cua-s1 **≠** + jev-macos-loop **≠** camoufox. + +Census not re-derived. Archer still NOT landed. + +Cross-repo addition: (iq) OCR+AX desktop CU family; +(ir) exclusive CU options / split questions; +(is) screenshot-to-frontier fail-closed on the +decision; (it) 155× one-screenshot ≠ Harbor score. + +## Batch #66 (2026-09-19 ~16:10 UTC / ~10:10 Boise) — moritzkremb/jev-voice-browser productized ASR CU + +Note: `research/notes.md` §82. Docs-only. Folded into PR #2. +Skip Archer. No invented metrics. Hunches labeled. +Quote README. Not a hit list. **Do not re-card** the +§39 tweet (~300 ms / $0.0002 *theirs*). User +SIGNAL_89ca (~08:50 Boise; MIT; ★100) + live this +pass **103★** / 12 forks / 0 issues; README SHA +`fa033303`; HEAD `054db0f3`. + +- **Productized ASR observe→score→act.** Web Speech + partials → Playwright snapshot ≤100 → one 9–11- + question hosted Jev request → policy. Same CU hole + as typesafe-computer-use OCR, different producer + (waveform → transcript, not screenshot → OCR). + Jev never hears audio. Skip Archer. +- **Partial-speech VOI.** Closed-set may fire on a + partial; free-text waits for final or 600 ms + silence. `complete` Noul + 900 ms silence *theirs*. +- **Pointer-not-generator.** Regex spans; Jev picks; + code copies. Numbered overlay, no second model. +- **Spoken confirm ≠ auth.** `destructive ≥ 0.5` → + say "confirm". README *theirs*: convenience, not a + guarantee. Control-port reach is the grant. + rh-guard owns the gate cousin. 0.5 / 0.55 / 0.6 + still soft. +- **27/27 fixtures ≠ Harbor.** ~$0.0002/call; p50 + ≈ 300 ms *theirs*. **≠** jev-voice-control **≠** + nikolas-j **≠** Aj1905 **≠** typesafe-computer-use. + +Census not re-derived. Archer still NOT landed. + +Cross-repo addition: (iu) ASR perception front-end; +(iv) partial-speech wait policy / free-text VOI; +(iw) spoken confirm ≠ auth; (ix) overlay +disambiguate without another model. + +## Batch #67 (2026-09-19 ~16:20 UTC / ~10:20 Boise) — AgentGhost wrap-as-execution + studio_yebisu JP genre atlas + +Note: `research/notes.md` §83–§84. Docs-only. Folded into PR #2. +Skip Archer. No invented metrics. Hunches labeled. +Quote README and tweet. Not a hit list. Dual-signal +turn. User SIGNAL_7c0f + SIGNAL_a92b (~08:50 Boise) ++ live this pass. + +- **Wrap-as-execution ALLOW/ASK/DENY.** + [reddpy/AgentGhost](https://github.com/reddpy/AgentGhost) + TypeScript; MIT; **2★** / 0 forks / 0 issues; + created 2026-09-18T21:16:55Z; HEAD `ac04e4fb`; + README SHA `44145fa9`. The wrap *is* the tool's + execution function. Rules first; ASK/DENY throw; + `failMode: closed`. Judge is a slot. Hosted + provider tools out of reach. `AUTO_APPROVE` is a + demo hatch, not a grant. **≠** jwen5419807/agentghost + **≠** vventirozos **≠** actiongate **≠** toolgate + **≠** jev-use. rh-guard owns the gate cousin. +- **JP genre atlas, not a bake-off.** + [@studio_yebisu](https://x.com/studio_yebisu/status/2101065176069886152) + 2026-09-18T21:45:48Z. Apps by hole. Stars + research-time (typesafe-computer-use 203→**427**; + jev-voice-browser 40→**103**). Not verified evals. + Engagement ephemeral (this pass 131,234 / 1,934 / + 192). SAM 3.1 already §39. OpenRouter Jev + no-waitlist is WATCH. **≠** @airesearch12 **≠** + v1.2. Do not dump the 30 repos. + +Census not re-derived. Archer still NOT landed. + +Cross-repo addition: (iy) wrap-as-execution / ASK +throws / fail-closed; (iz) application genre atlas +≠ class census ≠ scored board; (ja) star-count +drift as pedagogy. + +## Batch #68 (2026-09-19 ~16:25 UTC / ~10:25 Boise) — Akshay Pachaar “Jev Clearly Explained” + +Note: `research/notes.md` §85. Docs-only. Folded into PR #2. +Skip Archer. No invented metrics. Hunches labeled. +Quote ARTICLE.md. Not a hit list. User SIGNAL_55e5 +(~08:51 Boise; ~183k views) + live this pass +**233,495** / **2,280** / **235** / **3,652**. + +- **External pedagogy / how-to-apply.** Independent + of TypeSafe docs. LLM hammer for bounded + decisions; code owns branches; parallel + questions; thresholds in code; schema-safe ≠ + correct; three placements with an LLM (routing / + tool-risk / verify); shadow-mode; questions-as-code. +- **200× / 400× are TypeSafe ceiling claims.** + Article *theirs*: 70–500 ms, $0.042/MTok input, + output free; treat multiples as a ceiling, not a + promise. Not Harbor. +- **Text-only.** Convert environment to text/JSON + first. Not looking at the screen. Skip Archer. +- **Name locks.** ≠ official docs ≠ Flavio Copes + (cited further reading) ≠ LangChain harness + (cited; not a wrap card) ≠ AgentGhost §83. + +Census not re-derived. Archer still NOT landed. + +Cross-repo addition: (jb) public pedagogy receipt +for mixed architecture; (jc) schema-safe ≠ correct +as the safer hallucination sentence; (jd) marketing +multiples as ceiling, not Harbor. + +## Batch #69 (2026-09-19 ~16:30 UTC / ~10:30 Boise) — uehaj/jev-semgrep dedicated + +Note: `research/notes.md` §86. Docs-only. Folded into PR #2. +Skip Archer. No invented metrics. Hunches labeled. +Quote README. Not a hit list. User SIGNAL_9afa +(~08:54 Boise; ★42) + live GitHub this pass **51★** +(ephemeral). HEAD `21120e9`; README SHA `923e6a5`. + +- **Proposition ≠ embedding.** Cross-encoder + line+question → does the proposition hold, not + topical cosine. Contrast-set: all six about a + refund; only customer-*asking* pass. Angry agent + vs angry customer cosine ~1. +- **Boolean composition of soft Nouls.** AND/OR/NOT + after threshold in code; do not multiply p. + ≠ jev-combinators metaphor. +- **Semgrep.dev name collision.** Rename if both + installed. **≠** SAST. **Not a gate** (ranking + fail-open; rh-guard skip). +- **0.94/0.98 *theirs*.** LLM-as-judge 10 cases × + 51-line corpus. Not Harbor. Stars research-time. + +Census not re-derived. Archer still NOT landed. + +Cross-repo addition: (je) proposition ≠ embedding / +contrast-set; (jf) boolean composition of thresholded +Nouls; (jg) Semgrep.dev namesake + not-a-gate. diff --git a/research/notes.md b/research/notes.md index f14c821..cef4339 100644 --- a/research/notes.md +++ b/research/notes.md @@ -101,6 +101,24 @@ regex/LLM/code, Jev selects). - awlevin/typesafe-computer-use: $0.0002/step vs $0.032 Opus; key caveat: "every piece of reasoning the frontier model does for free has to be rebuilt here as deterministic state" (OCR + explicit date parsing). + **Productized HIGH:** `notes.md` §81. Do not re-card this one-liner. +- moritzkremb/jev-voice-browser: voice → Playwright; partial + transcripts → one 9–11-question Jev request (~300 ms); + pointer-not-generator for spans; `is_command` / `complete` / + `destructive` gates. **Productized HIGH:** `notes.md` §82. + Do not re-card this one-liner. +- reddpy/AgentGhost: wrap-as-execution ALLOW/ASK/DENY; + rules first, Jev remainder; ASK throws; fail-closed. + **Productized HIGH:** `notes.md` §83. rh-guard owns the + gate cousin. Do not re-card this one-liner. +- Akshay Pachaar “Jev Clearly Explained”: LLM hammer for + bounded decisions; schema-safe ≠ correct; 200×/400× + TypeSafe ceiling. **Pedagogy HIGH:** `notes.md` §85. + Do not re-card this one-liner. +- uehaj/jev-semgrep: grep by meaning; proposition ≠ + embedding; AND/OR/NOT after threshold; Semgrep.dev + name collision; not a gate. **Productized HIGH:** + `notes.md` §86. Do not re-card this one-liner. - RomanSlack/jev-drone: 500Hz control + 50Hz safety in code, classical CV to symbols at 15Hz, Jev advisory ~2.5Hz (Choice maneuver + Score risk + Noul lost-vs-occluded). "Cannot be the perception layer or run at control rate." @@ -1701,7 +1719,9 @@ System One decides on utterances. Independent public instance, not Basit: [Moritz Kremb](https://x.com/moritzkremb/status/2100577979021832365) (2026-09-17T13:29:51Z, `note_tweet` present). Talk → transcript → Jev probabilities → browser click. His ~300 ms and $0.0002 are his -receipt, not a class number. +receipt, not a class number. **Productized HIGH:** +[`moritzkremb/jev-voice-browser`](https://github.com/moritzkremb/jev-voice-browser) +`notes.md` §82. Do not re-card this tweet. **Not native omni System One.** Information dies at the interface: the decision call sees the schema you serialized, not the pixels or the @@ -2121,3 +2141,10411 @@ Kept, and patched in the cards: Not patched, standing risk: doctrine is copied across cards and the next fold is where copies drift; `notes.md` §1 is a one-day pin, not a live contract. + +## 44. Hourly fold ~12:58 America/Boise (2026-09-18) — store siblings, hard envelopes, distill-to-device, harness practice + +Window: America/Boise ~12:58 ≈ 18:58 UTC. Watch run 185349. Novel +versus §42 (~11:59). Docs-only. No Jev wrapper, no serving-stack +how-to, no copied SQL/`advise()`/proxy ports/`predict()`. Local +`/workspace/jev-archive` path for this run was not present; GitHub +READMEs, Hub cards, and X posts fetched live. HTTP 200 on cited URLs. +**Archer drop still WATCH** — no architecture rewrite this hour. User +watch says ~65% done, Qwen3.8-27B multimodal no-audio still expected +~19 Sep Boise. This pass did not retrieve a new Archer status tweet +(rate-limit / query constraint); Hub was not treated as a landing. + +Do **not** rehash pcdServer, jevql's already-folded CLI frame, +OpenSmoke, jev-mode, jot, typesafe-jev-tools, openevals, +hermes-north-star. Already-folded HIGH from §33 (jev-harness, +openjev-lm, jev-gate-student-b, open-jev-deberta, mini-jev-runs, +jev-tree-choice-cap, jev-pref) get a *frame*, not a second card. + +Six frames, then the artifacts. + +**(a) Structured-store semantic index.** Cheap exact predicates first; +typed questions on the remainder. Two *forks* of the same hole: +**in-engine extension** (`mgaitan/sqlite-jev`; inspired by +`realZachi/pg-jev`) vs **out-of-process CLI** (`kylemclaren/jevql` — +vanilla Postgres never sees `jev()`). zoxide (`joxide`, §42) is the +same hole over a path index. A semantic full scan is not an index. +Row contents leave the store: same residency warning as AU health +(§33). + +**(b) Distill-to-device memory/context gates.** System One as a +**sieve**, not only an action permit. Already §33 / `applied-mappings.md` +§1: `jev-gate-student-b` P(relevant) from yes/no logits. This hour the +meta is: vector recall → local yes/no gate → inject or stub. Fail-open +on errors. Teacher-copy ≠ gold. + +**(c) Encoder vs AR open replicas.** Already the three-open-paths cut +(§42). Encoder DeBERTa = public gold, OOD drop; decoder LoRA +(openjev-lm, jev-gate) = teacher-copy; constrained AR (TypeAR / +pcdServer) = softmax over allowed tokens, not a Noul. Do not pick a +path until meta-VOI says a model is needed at all. + +**(d) Soft judgment inside a hard safety envelope.** The model may +only **match the envelope or be more conservative**. Deterministic +policy is load-bearing; missing the model must still be safe. +bitrate-advisor is the named live-stream shape. mmalisper's JOB +planner is the named search/control shape (Postgres plans first; Jev +overrides only when confident). + +**(e) Perception → decision.** ASR / vision produce schema'd state; +System One decides; code acts. SAM and ASR are **producers**, not the +perceive species (§43). `jev-voice-control` is a README-only stub of +that pipeline. + +**(f) Harbor/jevals-style harness + shadow + confidence.** Assert on +the **action**, not on free text. LLM-as-judge is not the primary +System One score (`faq.md`). jev-harness already named this; this hour +the practice is first-class: recipes across alerts / RTB / sports-bet / +prediction-markets, offline fixtures, shadow until evals pass. + +### HIGH + +1. **[`mgaitan/sqlite-jev`](https://github.com/mgaitan/sqlite-jev)** + (created 18:25Z, C, license file absent, 0★). Loadable SQLite + extension: batched NL Noul/Choice/Score over row objects + (`jev_rows` virtual table; up to 40 rows per shared state). + Inspired by [`realZachi/pg-jev`](https://github.com/realZachi/pg-jev) + (in-engine Postgres, 152★ this pass — pointer, not a second card). + Sibling *pattern* to jevql, **not** the same serving choice: here + the database *does* see `jev()` as SQL. README limits: semantic + full scan, not an index; deterministic SQLite filters first; + `max_rows` is a spend guard; row contents go to TypeSafe. Thresholds + stay in SQL so they can rise with false-positive cost. Mock-server + tests never call TypeSafe. Do not copy `.load`, env, or SQL + signatures. Card: `mappings.md` §4. + +2. **[`AntonioCoppe/jev-harness`](https://github.com/AntonioCoppe/jev-harness)** + — already §33 (policy, confidence gate, shadow, 24-row 48.9s Claude + CLI → 1.3s Jev). This hour the *practice*: Harbor/jevals-adjacent + measurement substrate. Eval CLI asserts on the **action**, not on + prose. Recipes span alerts, NL row-filter, high-freq reflex (order / + RTB / fraud), prediction-market and sports-bet gates — business and + life, not SWE-only. Explicitly **not** a `/compact` replacement. + Do not copy the client. `validation.md`; gallery. + +3. **[`DECRUX9812/openjev-lm`](https://github.com/DECRUX9812/openjev-lm)** + — already §25. This hour the *receipts pattern*: overnight 6-vCPU, + $0/call, two independently written harnesses both 65/70 = 92.9% on + the same 70 hand-gold rows; 98.1% on 106 later postings is teacher + **agreement**, not gold; rare-class cells are one-row wide. Encoder + vs this decoder LoRA is the (c) fork, not a new architecture. Do + not copy train commands. + +4. **`SargeDev/jev-gate-student-b` + `jev-distill-corpus`** — already + §33. This hour the meta: **context sieve / memory gate**, not only + an action gate. Vector recall → P(relevant) yes/no logits → inject + or stub. Fail-open. Teacher-copy. `applied-mappings.md` §1. + +5. **`com-kotobalabs/open-jev-deberta-v3-large`** — already §33 + (434M, Banking77/SST5/BoolQ, in-domain ECE 0.022, OOD acc 0.690). + This hour: the **encoder** arm of open-replica, as opposed to + decoder LoRA (openjev-lm / jev-gate) and constrained AR. Public + gold, not a Jev teacher. Do not overwrite those numbers. + +### MED + +6. **[`affirmitv/bitrate-advisor`](https://github.com/affirmitv/bitrate-advisor)** + (created 15:27Z, MIT, TypeScript, 0★). Live-stream ABR: Jev + proposes initial rung / ceiling / resolution / next step from + telemetry + venue/carrier history; **deterministic policy fences + it in**. Jev may be as bold as the measured network allows and as + cautious as it likes, **never bolder**. Without an API key the same + call returns the policy's answer (`source: "policy"`). Power plan + (finish-the-game battery/thermal) is deterministic; Jev may only + make it more conservative. Author-measured 2026-09-18, three + states, OpenRouter billed ~$0.000041–0.000044, 0.25–0.39 s; ~$0.015 + per hour of stream at one decision / 10 s — **author-reported**, + not re-run. Domain: youth-sports phone streams (life / business), + not a codec tutorial. Empirical as a *shape* for §12 / §15 / §18. + Do not copy `advise()` or keys. + +7. **[`chris-wozniczek/jev-voice-control`](https://github.com/chris-wozniczek/jev-voice-control)** + (created 18:47Z, license absent, README-only this pass, 0★). + Speech → Jev typed decisions → macOS actions. Menu-bar Swift + **claim**. Hypothesis as a product; useful as the perception→decision + pipeline with ASR as producer (§39 / §43). Do not invent a Swift + API. + +8. **`doeixd/jev-pref`** — already the preference-lint contract + (YOU define the rule / Jev classifies / code maps outcome). No + rewrite. This hour it sits next to bitrate's envelope: policy-as- + prefs is the same ownership split on a diff instead of a live + stream. + +9. **[`nekowasabi/jev-routing`](https://github.com/nekowasabi/jev-routing)** + (created 08:13Z, MIT, Go, 0★). Host **adapter**, not an MCP + server: one Go binary in front of Claude Code / Codex / Grok + Build. Compacts tool results (verbatim drop/truncate, same + contract as fast-jev-compaction), then one Choice (next tool) + + Noul (done), then shrinks `tools[]` to **one schema**. README: + adding this via `mcp add` makes the catalog *worse*. No key → + on-device classifier. Do not copy ports, env, or install. + Host-adapter breadth, not a new species. `applied-mappings.md` §5. + +10. **[`trietphan/jev-claw`](https://github.com/trietphan/jev-claw)** + (created 18:44Z, MIT, JS, 0★). Typed routing for OpenClaw + agents. **Jev classifies** (task_type / complexity / risk / + second-opinion); **`decide()` in code** maps to a route. + Sensitive-path regex floors risk even if Jev underrates a + migration. Confidence is the **minimum** across classifications, + not the average. 11 offline policy tests (no network). Live eval + 10/10 on the author's 10 samples — **author-reported**, tiny. + Same hole as routeKit (§33): Jev does not get to skip the + escalation `if`. Do not copy the eight route names as doctrine. + +11. **[`gamesonrblx/JevML`](https://github.com/gamesonrblx/JevML)** + (created 18:47Z, license absent, README-only, 0★). "PCA / MCMC / + text-diffusion / NCA primitives + a harness that picks the right + tool." Meta-VOI adjacent: which primitive, if any. Hypothesis. + Do not invent those APIs. + +12. **`Mikhail/mini-jev-runs`** — already §33 (27.9k option-logit + runs, frozen Qwen3-4B, no token generated). Calibration / + constrained-decode corpus. No rewrite. + +13. **`reachjalil/jev-tree-choice-cap`** — already §33 / mappings §5. + Hierarchical Choice under the 255 cap. No rewrite. + +### X discourse this hour (verified) + +- [@mmalisper](https://x.com/mmalisper/status/2101001041903009987) + (17:30:57Z) and thread: Jev-assisted Postgres query planner on the + Join Order Benchmark. Join-order Choice **2× slower** (defaulted to + smallest table). Cardinality estimates helped when outside context + informed the plan; when Jev was wrong, one query was an **order of + magnitude slower**. Hybrid: Postgres plans first; Jev overrides + **only when confident** → **+12% geomean**, no dramatic slowdowns. + Downside: a Jev call is 100s of ms, not yet practical on every + plan. **Author-reported**, not re-run. Frame (d) + search/control + (§9): the planner is the envelope; confidence is the gate. + Fail-open to Postgres. + +- [@higgsfield_ai](https://x.com/higgsfield_ai/status/2101022473248727177) + (18:56:07Z) and [demo](https://x.com/higgsfield_ai/status/2101022133753430365) + (18:54:46Z): "perfect use case" — Jev auto-routes GenAI (video/image) + models for cost/speed/quality on the Higgsfield API. **Claim**, no + labeled catalog receipt. Same hole as routeKit / jev-claw: + classify requirements, policy picks the generator. Hypothesis + until *your* catalog. + +Cards: `mappings.md` §4 / §9 / §12 / §15 / §18; `mixed-architecture.md` +gallery + host-adapter note; `applied-mappings.md` §1 / §5; +`validation.md` harness practice; `mental-models.md` envelope + +planner; FAQ in-engine vs CLI; `judgment-class.md` encoder vs decoder +replica (no new species). No wrapper. + +## 45. kev — runnable Archer reconstruction, not a distill (2026-09-18) + +HTTP 200 this pass: +[jaredpalmer/kev](https://github.com/jaredpalmer/kev) README (raw +`main`), [MODEL_CARD.md](https://github.com/jaredpalmer/kev/blob/main/MODEL_CARD.md), +[LICENSE](https://github.com/jaredpalmer/kev/blob/main/LICENSE) +(Apache-2.0; GitHub API `license.key` apache-2.0), +[release v0.1.0](https://github.com/jaredpalmer/kev/releases/tag/v0.1.0), +[Archer Hume, *Jev's Architecture Unmasked*](https://archerhume.com/posts/jevs-architecture-unmasked) +(with and without trailing slash), TypeSafe +[System One API](https://docs.typesafe.ai/api). + +GitHub API this pass: created 2026-09-17T20:49:39Z; pushed +2026-09-18T19:50:58Z; 24★; language TypeScript (playground); default +branch `main`. Description: "tiny Jev-like model built on top of +Qwen2.5-0.5B you can train and run on your MacBook." Not +`Kevthetech143/super-jev`. + +**What it is (Contract as README + model card).** Jared Palmer, Apache-2.0 +adapter/head (Qwen2.5-0.5B under the Qwen license). LoRA (r=16) + a +pointer readout on that 0.5B causal backbone. Typed questions in, +calibrated probabilities out, one prefill pass, no decode. State and +questions packed into one sequence; a block-causal mask lets each +question see the document and never a sibling. The pointer head scores +each option against the question's `` token and softmaxes. +Trained with cross-entropy against labelled outcomes. Architecture +explicitly follows Hume's reconstruction (§31): shared state, isolated +questions, pointer head, CE vs labelled outcomes. API follows TypeSafe +`POST /v1/systemone`; official `typesafe-sdk` works with a `base_url` +change. Weights `kev-0.5b` (38 MB: LoRA adapter, readout head, +tokenizer) on GitHub release v0.1.0; base model from the Hub on first +load. Trains ~1h45m on an Apple M5; ~160 ms for a six-question +request. Model card: research prototype, not production, not Jev. +Do not copy serve flags, ports, or train commands into skill cards. + +**Not a distill.** Six public datasets converted to TypeSafe-shaped +requests (Banking77, AG News, MNLI, BoolQ, SST-5, Yelp): 9,000 records, +13,500 questions, two epochs. No LLM-generated data. Distinct from +openjev-lm / jev-gate on the *same* 0.5B backbone: those copy a hosted +Jev teacher. Distinct from Nimble: 9B contrastive synthetic labels, no +measured ECE. Distinct from encoder open-jev: bidirectional DeBERTa, +OOD drop measured. Distinct from TypeAR / pcdServer: those decode a +constrained next token. Distinct from proprietary Jev: closed weights, +~32k envelope. Distinct from Archer Watch: 27B announced, multimodal, +no Hub weights this pass. + +**Evidence (README / model card; not re-run; Empirical as their named +receipt).** Isolation exact: packed vs separate max Δ **3.7e-6**; +secret-in-sibling / absent / in-state **p = 0.03 / 0.03 / 0.99**. +Held-out ECE **0.065** (10 bins) on 1,350 ID questions; **0.031** after +one-parameter temperature scaling (T=1.47). Overall acc **0.799**. +Permute (4 orders, Choice K ≥ 3): argmax flips **7.4%**. IIA: log-odds +shift from one irrelevant option **mean 0.13**, p90 0.34. Boundary +forgery: option count unchanged, forged option p ≤ 0.09. Per-source +cells (acc / ECE) stay in the model card; do not promote them into a +ranking against Jev. + +**Honest limits (their words).** 0.5B knowledge (on the TypeSafe docs' +structured-criteria example kev picks `return_policy` where Jev picks +`return_status`). Calibration is in-distribution; ECE on the training +datasets says nothing about a new workflow. Not multimodal. Trained at +384 state / 1,024 branch tokens; serving caps at 8,192 vs Jev ~32k. +Score confidence is a stand-in; TypeSafe has not published theirs. +Choice `confidence` uses `(p_max − 1/K) / (1 − 1/K)` — the same +arithmetic Hume reconstructed in the official adapter (§31), as *their* +API derivation, not a TypeSafe contract. + +**Placement.** (a) trained decision-only open path next to Laya / Nimble +/ Archer Watch; (b) cleanest *runnable* productization of Archer's +reconstruction (API-compatible); (c) contrast vs TypeAR (constrained +AR decode) vs encoder open-jev vs proprietary Jev; (d) jevals/Harbor +bake-off candidate; mechanism tests mirror Archer probes — they +falsify the reconstruction, they do not prove kev = Jev; (e) when to +use: laptop-local System One API drop-in for development/eval; not a +knowledge/frontier substitute. **Empirical** as a public repo + named +ID receipt. **Hypothesis** that it substitutes for Jev on *your* +labels. Cards: `judgment-class.md`; FAQ; `validation.md`; +`mental-models.md`; `mixed-architecture.md`; `formal-methods.md` +(pointer-softmax is still a sensor); `optimizer-integration.md`. +No wrapper. + +### Delta (~14:35 Boise) — Hub weights, PEFT task_type, NOTA training + +Not a rewrite of §45. No architecture species change. HTTP 200 this +pass: Hub [`jaredpalmer/kev-0.5b`](https://huggingface.co/jaredpalmer/kev-0.5b) +(Apache-2.0 adapter; `peft` / `text-classification` card); GitHub +README now names that Hub id as the fetch path; release tarball +remains. GitHub this pass: **61★**, pushed ~20:29Z (signal had 53★). +Do not copy `--run` / publish flags into skill cards. + +1. **Hub weights.** `kev.publish` uploads the adapter; `--run` accepts + Hub ids (`jaredpalmer/kev-0.5b`); the Qwen2.5-0.5B base still + downloads on first load. Bake-off fetch path for jevals/Harbor + (`validation.md`). +2. **PEFT.** `LoraConfig(task_type="FEATURE_EXTRACTION")` (`kev/model.py`). + Publish patches legacy adapters that saved `task_type` null (Hub + warns; PEFT treats both the same on a bare backbone). Docs mention + only kev-0.5b. Pointer readout, not a new species. +3. **HIGH — `none_of_the_above` is a training question, not only a + request hatch.** First training run learned "this wording ⇒ pick + it" when NOTA appeared only as the correct answer (`kev/data.py` + comment). Fix: add NOTA as a **wrong alternative** too + (`p_none_distract`); vary wording (`NONE_OPTIONS`: "None of the + above" / "Something else" / "Not listed here" / …); dedicated + `test_none_of_the_above` (true option still present → little mass + on none; true option removed → pick none; a shortcut model picks + it in both cases). **No published rates this pass — do not + invent them.** Model card already augments with p=0.10 + true→`other: None of the above`; the *delta* is confronting the + hatch as a distractor as well. Cross-link wellposed / Choice + `"other"`: request-shape lint puts the residual option on the + offered set; **training must confront that option** or the hatch + becomes a wording shortcut. Cards: `question-design.md`; FAQ + forced Choice; `validation.md`. + +## 46. 14:03 Boise hourly — open multimodal RLCD, bake-off substrate, decision-token LoRA (2026-09-18) + +America/Boise 14:03 = 20:03 UTC. Docs-only fold into PR #2 +(`cursor/augustus-store-envelope-00b4`). Archer Hume 27B drop still +**WATCH** (Hub authors `archerhume` / `4rcherhume` empty this pass; +user watch still ~2026-09-19). No invented metrics. Frames first, +not a hit list. Identity lock vs `typesafe-ai` / `tenbin` / +`decision-first` holds. No wrapper, no serve how-to. + +HTTP 200 / Hub fetch this pass: +[`BlackwoodAI/blackwood-rlcd`](https://huggingface.co/BlackwoodAI/blackwood-rlcd) +README; [`pngwn/open-jev-laya-bench`](https://huggingface.co/datasets/pngwn/open-jev-laya-bench) +`RESULTS.md` (no dataset README); [`Foodoo1/Qwen3-14B-RLCD-Decision-LoRA`](https://huggingface.co/Foodoo1/Qwen3-14B-RLCD-Decision-LoRA) +README; GitHub READMEs for +[`thevibeworks/jevgate`](https://github.com/thevibeworks/jevgate), +[`suraj-phanindra/wellposed`](https://github.com/suraj-phanindra/wellposed), +[`khordoo/jev-reflex-autonomy-lab`](https://github.com/khordoo/jev-reflex-autonomy-lab), +[`Wany-i/jev-decision-layer`](https://github.com/Wany-i/jev-decision-layer), +[`perixtar/jev-e2e`](https://github.com/perixtar/jev-e2e), +[`yalindogusahin/jevpandas`](https://github.com/yalindogusahin/jevpandas). + +### HIGH + +1. **[`BlackwoodAI/blackwood-rlcd`](https://huggingface.co/BlackwoodAI/blackwood-rlcd)** + (Hub, 2026-09-18 v1; `pipeline_tag: image-text-to-text`; CC BY-NC 4.0; + likes=1 this pass). **Open multimodal RLCD decision model.** + Screenshot or text in; Choice / Score / Boolean out in one prefill; + Jev-compatible `/v1/systemone` shim. Options listed as letters; + answer from option-letter logits at the last position; temperature- + calibrated; no sampling; `output_tokens` always 0. Not Archer's 27B + drop. Not CLIP/SigLIP (those are vision *scorers*). Same *decide* + species as Jev / Laya / kev, with image-in. + + **Architecture (load-bearing).** Omni perception→decision can ship + **without waiting for Archer**. Soft judgment over **pixel + candidates code already marked** (letters drawn on the screenshot; + criteria keyed by those letters) inside a deterministic click/act. + Specialist composition (SAM / OCR / AX → text → Jev) still exists + (`§39`); this is the shared multimodal System One that card was + waiting for. Information still dies at the *act*: the model picks a + letter; code clicks. A Noul is still not a proof. CC BY-NC: research + / personal, not a commercial drop-in. Do not copy vLLM flags, the + shim, or curl bodies into skill cards. + + **Evidence (model card; paired per item; not re-run; Empirical as + their named receipt).** Web element choice, 300 held-out steps: + acc **0.907** vs Jev 1.13 text-only **0.480**; letter-shuffle flip + **0.133** vs **0.587** (lower better); ECE **0.037** vs **0.091**; + latency **~200 ms** (1×H100) vs 441 ms (Jev via OpenRouter). + 4-lettering (four orderings averaged, 4× compute) acc **0.953**. + Desktop held-out acc **0.76** vs **0.654**. NL predicates over + records **0.973** vs **0.965**. Tetris lines / 60 pieces **14.2** vs + **13.6**. General text, 85 public sets / 8,456 items: blackwood + **0.786**, Jev **0.850** — Jev still leads. Domain sets 27 / 12,746: + **0.710** vs **0.703**. Pixel rows use a randomized viewport crop. + Jev is text-only by design: screenshot rows compare screenshot + input with Jev's *text* input on the same steps. **Hypothesis** that + it substitutes for Jev on *your* labels. Boolean questions on skewed + sets can be over-confident (their limit). Cards: `judgment-class.md` + (family, holes, when-to-use, dedicated card, perception rewrite); + FAQ wait-for-Archer; `validation.md` letter-shuffle; `formal-methods.md` + one sentence; `mixed-architecture.md` gallery; `applied-mappings.md` §2. + +2. **[`pngwn/open-jev-laya-bench`](https://huggingface.co/datasets/pngwn/open-jev-laya-bench)** + (`RESULTS.md` this pass; no dataset README). Shared bake-off: + [`pngwn/system-one-qwen3.5-4b-scorer`](https://huggingface.co/pngwn/system-one-qwen3.5-4b-scorer) + `@ e6464dce` vs [`convaiinnovations/laya`](https://huggingface.co/convaiinnovations/laya) + `@ 7c76b622`. **26 neutral + 9 home** tasks, **11,959** test items, + **3,269** cal items. Hardware: NVIDIA A100-SXM4-80GB. Config chosen + on `cal` (scorer vs `laya_text`). **Not TypeSafe Jev vs Laya** — do + not launder this table into a proprietary-Jev ranking. Laya's + published in-task acc/ECE sit on unpublished sets and are context, + not like-for-like. + + **Harbor / jevals practice in the wild.** Measurement substrate + exemplar: accuracy + **ECE / NLL / Brier** (and acc@50% coverage), + held-out `test`, temperature chosen on `cal`, leave-one-task-out, + prompted-base / prompted-instruct arms on the same 4B, per-task T + vs shipped T. **LLM-as-judge is not the primary System One score.** + That is the eval path the skill already wanted (`validation.md` + Eval & hill-climb; `notes.md` §40). Do not copy the scoring scripts. + + **Evidence (`RESULTS.md`; not re-run; Empirical as that named + receipt).** Neutral, shipped T: System One 4B macro acc **0.760**, + Laya **0.737**, Δ **+0.023 [+0.013, +0.032] \*** (interval excludes + 0). Neutral ECE-10 **0.074** vs **0.122**; macro NLL **0.569** vs + **1.272**; macro Brier **0.324** vs **0.392**. Home (the 4B's own + held-out split): **0.715** vs **0.486**, Δ **+0.229 [+0.198, +0.262] + \***. Home ECE-10 **0.072** vs **0.272**; NLL **0.697** vs **3.836**; + Brier **0.363** vs **0.759**. Ahead on 13/26 neutral and 8/9 home. + Neutral prompted-instruct (same 4B, ≤26 options) macro acc **0.763** + vs fine-tune **0.760**, Δ **+0.003 [-0.004, +0.011]** — interval + includes 0; do not slogan "fine-tune always wins zero-shot." Home + prompted arms drop banking77 / ticket-queue (enum >26). Latency + one-question p50 on the same GPU: tweet_emotion Laya 27.9 ms vs + scorer 76.7 ms; home_banking77 (K=77) 30.0 vs 382.8 ms. **Hypothesis** + as a ranking of *your* head. Cards: `validation.md` bake-off; FAQ + LLM-as-judge; `judgment-class.md` Laya companion. + +3. **[`Foodoo1/Qwen3-14B-RLCD-Decision-LoRA`](https://huggingface.co/Foodoo1/Qwen3-14B-RLCD-Decision-LoRA)** + (Apache-2.0 adapter; base `Qwen/Qwen3-14B`; PEFT). **Decision-token + QLoRA** under **parallel constrained decoding**: one prefill, KV + broadcast across fields, logit slice over candidate first tokens — + the TypeAR / pcdServer / `stephanj/parallelConstraintDecoding` + inference pattern (`findings.md` batch #1). Loss computed **only on + the single decision token** of each field. The base already hits + easy fields (language/sentiment 100%); reasoning-heavy fields fail + because the decision happens at one token with no room to think. + Pattern, not a fraud product: train the decision token, do not + fine-tune generated prose. Evaluated with + [`harshatheg/Qwen-2.5-1B-RLCD`](https://huggingface.co/harshatheg/Qwen-2.5-1B-RLCD) + inference code (that Hub repo still has no weights — stock Qwen + + custom code; already archived). Synthetic fraud-triage schema; + **do not use for real financial decisions** (their limit). Do not + copy the prompt or LoRA flags into skill cards. + + **Evidence (card; held-out 200-case, 4 fields, 4-bit NF4, RTX 3090; + not re-run).** fraud_risk (4-way) **64.0% → 95.0%**; block_account + **77.0% → 100%**; language and sentiment stay **100%**; overall + (800 decisions) **85.2% → 98.8%**. Mean latency **~234 ms** per + case (all 4 fields in one broadcast). Remaining errors all + LOW→ELEVATED (conservative); zero high-risk judged low. Train: + 16,608 single-token examples from 4,152 synthetic cases (648 + templates × 72 messages EN/ZH/ES/JA), 1 epoch, ~97 min on one 3090. + **Empirical** as that named receipt. **Hypothesis** as a general + recipe on *your* schema. Cards: `judgment-class.md` constrained-AR; + `methods-catalog.md`; FAQ open-weights vs constrained decode. + +4. **[`thevibeworks/jevgate`](https://github.com/thevibeworks/jevgate)** + — already §25 / `mappings.md` §18. This hour the *mental model*, + matching the GitHub description: **an allowlist proves** every verb + is a listed read-only tool; **Jev judges only unlisted** leftovers; + **fail-open (cannot block)**. Hard envelope in code + soft judgment + only on the residual. Proven runs in 4 µs and never leaves the box; + Refused never asks the model (so a comment cannot talk a writer + through); Unknown is five Nouls, admit iff every p < 0.2. Jev on + `/bin/ls` alone leaks (0.04) — that is why it is the third tier. + Real-traffic counts already on the README (116,979 Bash calls: + 26.4% proven; 16.2% reach Jev; 47 admitted, all read-only by hand) + stay author-reported; do not re-promote 0/59 as if new. MIT. 0★ + this pass. No rewrite of the 0.2 threshold as a constant. Cards: + `mappings.md` §18 "proves"; FAQ allowlist-then-judge; methods-catalog. + +5. **[`suraj-phanindra/wellposed`](https://github.com/suraj-phanindra/wellposed)** + (MIT, JS, 0★, created 2026-09-17). **Lint the request before it + comes back confidently wrong.** Neighbor-skill identity: `tenbin` + still owns design-time lint/measure; this is an Empirical *recipe* + of that hole, not a second Augustus skill and not a wellposed + how-to. Standing Choice contract (`SKILL.md`): probabilities are + conditional on the offered set; absent candidates can never be + chosen; add `other` where coverage is open. wellposed is the + receipt that **violating that contract is silent**. + + **Evidence (README; not re-run).** 40 generated requests: **0/40** + syntax errors (API validation already covers that); **16/40 (40%)** + asked something that did not make sense; **0/11** list-questions + included a none-of-the-above. Live probe: unsubscribe email, four + department options, no `other` → `"support issue"` **confidence + 1.00**; with `other` → `"other"` confidence 0.93. Overlapping + `angry`/`furious` collapsed confidence to **0.19** — loud; ordinary + gates catch it. Forced wrong Choice is quiet. Structural lint (35 + rules, offline) then optional jev-on-jev semantic layer. Labeled + corpus: recall **22/26 = 85%**, precision **22/24 = 92%** (computed + live from the corpus). Honest limits: one-model labels, so 40% is a + floor; 2026-09-18 adversarial audit added 36 items because the + metric could not see ordinary-English false positives. Broken state + paths (`ticket.assigned_agent.name` with no such path) are + *provably* wrong. Confidence gating **cannot** catch a forced + Choice. Cards: `question-design.md` diagnosis; validation gate #2; + FAQ; tenbin identity lock. + +6. **[`khordoo/jev-reflex-autonomy-lab`](https://github.com/khordoo/jev-reflex-autonomy-lab)** + (TypeScript, 0★, created 2026-09-18T19:46Z; GitHub license null this + pass). Interactive multi-drone lab. **S1 Jev reflex keeps control**; + optional S2 planner (OpenRouter / GLM 5.3) is **one-use advice** on + low confidence. Jev does not pause while the planner responds. S2 + does not fly the drone. Kahneman row already taught the split + (`toolbox-mapping.md`: S2 proposes, S1 discriminates; never the + reverse) — this is that split as a control loop, not a flight + controller. Experimental visualization; mock mode without keys; + live fleet success varies. No metrics to promote. Do not copy the + adapter, `.dev.vars`, or ports. Cards: `mixed-architecture.md` dual + orchestration; `agent-self-assessment.md`; toolbox Kahneman row. + **Delta (consumption telemetry, Local-vs-Live A/B, seed≠replay, + no-pixels, 20% still soft):** `notes.md` §80. Do not re-card this + one-liner. + +### MED (pointers, not cards of their own) + +7. **[`Wany-i/jev-decision-layer`](https://github.com/Wany-i/jev-decision-layer)** + (MIT, Python stdlib, 0★). Wrap the decision model as a **business + decision tool**: caller names the *judgment*, not the model. + `decide(name, fields)` → outcome + confidence + **`gate` (part of + the result)**. Registry JSON; hard guards in code; `other` required + on Choice. OpenRouter `POST /api/alpha/decisions` (not + `chat/completions` — that 400s). Text-only: screenshots must be + textualized — contrast with blackwood. 28 offline tests, no key. + Unofficial. Do not copy the registry or MCP install. + +8. **[`perixtar/jev-e2e`](https://github.com/perixtar/jev-e2e)** + (MIT, TypeScript, alpha, 0★). Natural-language cases; Jev selects + observed controls; **Playwright executes and independently checks + expectations**. Verdicts PASS / FAIL / BLOCKED. A completed + navigation or a confident model response **cannot substitute for + checked expectations**. Harbor/jevals practice on a browser taskset + (score on the task; harness rolls out). Alpha: controlled demo does + not establish reliability across arbitrary sites. Do not copy CLI + flags or ports. + +9. **[`yalindogusahin/jevpandas`](https://github.com/yalindogusahin/jevpandas)** + (Python; GitHub LICENSE 404 this pass). pandas frame as the store: + `evaluate` / `filter` / `classify` / `score` / batched `ask`. + Classify example includes `other`. Failures never become negative + predictions. No generative chat, no joins, no training on review + labels. Store-as-semantic-index cousin of jevql / sqlite-jev + (`mappings.md` §4) over a dataframe instead of SQL. Samples in + `data/` are synthetic. Do not copy the client. + +Cards: `judgment-class.md`; `validation.md`; `question-design.md`; +`mappings.md` §18 / §4; `mixed-architecture.md`; `faq.md`; +`agent-self-assessment.md`; `mental-models.md`; `formal-methods.md`; +`applied-mappings.md` §2; `methods-catalog.md`; `toolbox-mapping.md`. +No wrapper. + +## 47. Abide — productized soft-rule preference lint (2026-09-18) + +America/Boise ~14:34 = 20:34 UTC. Docs-only fold into open PR #2 +(`cursor/augustus-store-envelope-00b4`). Not a competing PR. Archer +27B drop still **WATCH**. Identity lock vs `typesafe-ai` / `tenbin` / +`decision-first` holds. No wrapper, no hook how-to, no copied `npx` / +init / ports. Text/diff only — **not multimodal**. Do not invent +metrics. + +HTTP 200 this pass: GitHub README for +[`coldteadotai/abide`](https://github.com/coldteadotai/abide) +(MIT, TypeScript, created 2026-09-18, npm `@coldtea/abide`, 4★ this +pass); replay method in-repo at `benchmarks/replay/README.md`. +Attached SIGNAL + replay README + repo.json as primary placement +guidance. + +### HIGH + +1. **[`coldteadotai/abide`](https://github.com/coldteadotai/abide)** + — productized Jev hooks for Claude Code / Codex / OpenCode that + enforce **soft project rules** (AGENTS.md / CLAUDE.md / instruction + files) no linter can check. On every edit (or turn) it asks Jev + **one Score / probability per rule against the diff only**, never + the conversation. A break names the rule and asks the agent to + repair in-session. jev-pref already stated the contract (YOU define + the rule / Jev classifies / code maps outcome). Abide is the + fuller productized path: compile / calibrate / tune / replay / + audit + multi-host. Do not clone the CLI. + + **Architecture (load-bearing; not a hit list).** + + 1. **Soft rules → soft judgment; hard rules → linter.** Rules a + linter can check are handed to the linter. Same layering family + as jevgate's hard envelope (`mappings.md` §18): structure + **proves** what it can; typed judgment only on the residual. + jevgate's remainder is unlisted shell verbs; Abide's remainder + is project-instruction soft rules. Different holes; same + sandwich. Soft judgment is never the sole hard veto. + 2. **Edit-phase vs turn-phase is an observation-window question.** + Each compiled rule runs at `edit` (this hunk) or `turn` (the + whole turn diff). "Did this add more than was asked?" has no + answer after edit 1 of 12. Question design must name when the + evidence exists (`question-design.md`). + 3. **Banded confidence + fail-open.** Their product bands: ≥0.8 + repair in-session; 0.5–0.8 human note, agent silent; <0.5 + silence. That 0.8 is **their** operating point, not a universal + threshold (`mappings.md` §2 already forbids magic 0.8). Hooks + always exit 0, hard deadline; no key / no network → the edit + proceeds and the miss is logged. Soft judgment never sole hard + veto. + 4. **Rubric as editable artifact.** `.abide/rubric.json` quotes the + source instruction line. A wrong verdict is a rule rewrite. + `calibrate` scores rules against recent git history; `tune` + rewrites dead rules. False positives live in the question + (scope, criteria), not in the model. + 5. **Eval honesty / Harbor-adjacent.** Replay of 93 real Claude + Code sessions against each repo's own AGENTS.md (two private + repos + public `pr-lens`; hunks unpublished). Nothing is + re-run: diffs come from transcripts. Independent reviewer + (Claude, reading each rule's own text strictly; owner + spot-checked four first-pass comment flags and agreed). Replay + does **not** measure whether the agent repairs when told. + Harbor/jevals practice in the wild: frozen transcripts, phase + split, independent confirmation — not a Harbor taskset and not + a second jevals (`validation.md`). + + **Evidence (README + `benchmarks/replay/README.md`; not re-run; + Empirical as that named receipt).** 93 sessions with edits; 1,256 + edits judged; 147 turns judged; Jev cost **$0.22**; wall **~2 + min**. Edits flagged at 0.8+: **39 (3.1%)**. Turns flagged at + 0.8+: **15 (10%)**. After independent review: edits **10/39 = + 26%** precision; turns **11/15 = 73%**; all **21/54 = 39%**. + Author's gloss: 8 confirmed violations per 1,000 edits; 1 turn in + 13 ends with a confirmed violation of a rule no linter could + express; 11 of 93 sessions contained at least one. Recall probe + from 20 cleared hunks closest to the line (0.31–0.48): one real + miss (props-ordering). Most false positives from two edit-phase + rules (`plain-error-for-expected-failure`, `comment-volume`) — + fixable with `scope` and criteria, which is what calibrate / tune + exist for; the table is **before** either ran on the tightened + rubric. No turn-number drift: pr-lens-app flat ~2.5% through early + / mid / late turns; coldtea falls 5.8% → 1.3% because big + new-file writes happen early. Agents break these rules from the + first edit at a steady rate. Economics (README, author-measured + 2026-09-18, direct to TypeSafe): a check was ~2,500 tokens with an + ordinary LLM (cent+, seconds, prose to parse); Jev ~300 ms, + **$0.00004–0.00007** per check on this repo's 13 rules (1,000–1,600 + input tokens). Do not promote those $ / ms figures as class + constants. + + **Siblings — complementary, do not merge.** + + - **`doeixd/jev-pref`:** earlier watch; the contract Abide + productizes. Keep the contract quote; point here for compile / + calibrate / tune / replay / multi-host. + - **`24601/rh-guard`:** reward-hacking / eval-integrity on agent + tool use. Abide: project-instruction soft rules on diffs. Same + hook-host surface, different judgment class. Shared fail-open / + host-adapter lessons; no code dependency. + - **`suraj-phanindra/wellposed`:** request-shape lint still + upstream of any Score call (`tenbin` owns the skill). + - **`thevibeworks/jevgate`:** hard envelope owns safety; Abide is + soft residual judgment on residual soft rules. + - **`huntedman/JevLint`:** file-level convention Nouls; sibling, + not a substitute. + + **Placement.** Verifier over rules the project already wrote + (`mixed-architecture.md` preference lint; composition-algebra #9). + Pillar: selective classification / abstention + structural-prove ∩ + remainder. Hole: gate. Family: closed decision API (typed Score + per rule). Fail-open, banded. Eval path: replay + independent + review (named receipt above); not a substitute for jevals/Harbor + on *your* rubric. **Empirical** as README + dated replay. **Hypothesis** + that the same bands / precision transfer to *your* AGENTS.md. + Cards: `mixed-architecture.md`; `question-design.md`; + `validation.md`; `mappings.md` §2 / §18; `faq.md`; + `agent-self-assessment.md`; `toolbox-mapping.md`; + `methods-catalog.md`; `mental-models.md`. No wrapper. + +### Omni / Jev-omni + +Text/diff only today. Usage + measurement exemplar (replay harness, +precision by phase). Not a multimodal substrate and not a reason to +wait on Archer. + +## 48. Extractive selection, pointer-not-generator, local contract drop-in (2026-09-18 ~14:52 Boise) + +America/Boise ~14:52 = 20:52 UTC (run 205210). Docs-only fold into +open PR #2 (`cursor/augustus-store-envelope-00b4`). Not a competing +PR. Archer 27B drop still **WATCH**. Identity lock vs `typesafe-ai` / +`tenbin` / `decision-first` holds. No wrapper, no install.sh, no +copied ports / Docker / PyPI how-tos. No invented metrics. Do not +re-fold blackwood-rlcd, open-jev-laya-bench, decision-token LoRA, +jevgate, wellposed, jev-reflex-autonomy-lab, Abide, kev, jevpandas, +bitrate-advisor. + +Backend-agnostic: these are categorization / scoring *placements* +(extractive keep/drop, pointer evidence, contract-compatible local +scorer, structured observe→decide→act, dataframe columns, route≠memory, +advisory sidecar, structure induction, AST∩semantic). TypeSafe Jev is +the exemplar in the READMEs, not a monopoly. + +### HIGH + +1. **[`AppitStudio/testimonial-miner`](https://github.com/AppitStudio/testimonial-miner)** + (MIT, Python, created 2026-09-18T20:42Z, 0★ this pass). Gmail/IMAP + → numbered sentences in **code** → **one** typed request per email + (Choice kind + app, Nouls for user/praise/problem/English, Score + quotability, **one Noul per sentence**). **The model never writes + text.** The stored quote is the sender's sentences, selected by + those Nouls and joined in order. All thresholds live in + `Thresholds` and `redecide` reapplies them to logged answers + **without new model calls**. Header rules skip newsletters / + outbound / quoted replies before any call. Pattern: **extractive + selection + multi-question broadcast + offline re-thresholding.** + Cousin of exact-text keep/drop (`applied-mappings.md` §2), not a + testimonial product and not a Gmail how-to. + + **Evidence (README; not re-run; Empirical as that named receipt).** + Developed against `typesafe-sdk` 0.7.0 / `jev-1.13.0`. Their + *product* bars (not class constants): candidate if user ≥0.5, + praise ≥0.6, quality ≥1.5 on 0–3; borderline praise ≥0.35 / + quality ≥1.0; quote sentences ≥0.6. Live fixtures (8 requests, + ~22k input tokens): **5** candidates, **3** rejected, **3** header + skips. Offline tests use fakes and say nothing about model + accuracy. Author cost gloss: ~$0.042 / M input; typical email + 2–3k tokens; 10k judged emails ~$1. Do not copy IMAP setup. + Cards: `applied-mappings.md` §2; `mappings.md` §2; `methods-catalog.md`. + +2. **[`choxos/jev-reviewer`](https://github.com/choxos/jev-reviewer)** + (MIT, JavaScript, 1★, created 2026-09-18T14:47Z). Systematic-review + data extraction: PDF/Office/HTML/CSV stay in the **browser**; + segmenter assigns line ids (`A001`…); the decision model **points + at ids**; **code copies verbatim quotes with file and place**. + Nothing is paraphrased, so nothing can be invented. Two-pass: + screening Choice "which line answers q?" (+ none) over chunks, + then verifying Nouls "does this line itself answer q?". **Not + found is an answer.** Speculative fan-out: every question against + the same text. Pattern: **pointer-not-generator** for evidence + synthesis / citation integrity. Same species as keep/drop over + candidates code already holds; GLiNER locate is the cousin when + the answer *is* a span the encoder proposes. + + **Evidence (README; sample study, Sep 2026; not a validation + study).** 17-page article + 12-page analysis plan + CONSORT, + 712 lines: 1 question 10 req / 1.2–2 s / $0.0016; 9 questions + 17 / 2.3 s / $0.0052; 18-question template 27 / 4.6 s / $0.0101. + Chunks of 12k characters matched 7k with a third fewer requests. + Treat as spot checks. Scanned PDFs need OCR first (their limit). + Do not copy the relay. Cards: `applied-mappings.md` §2; + `methods-catalog.md` claim–evidence; `mental-models.md`. + +3. **[`us/jev-local`](https://github.com/us/jev-local)** + (LICENSE absent this pass; Python; 0★; created 2026-09-18T19:28Z). + Local `POST /v1/systemone` drop-in (Docker / `pip` / SDK + `base_url`). Open-weights *path*, no waitlist. **Default scorer is + a deterministic stub and carries no intelligence** until + `JEVLOCAL_SCORER=hf`. That flag is load-bearing: a green smoke + test on the stub is not a local Jev. Pattern: **contract-compatible + local scorer for offline/dev**, beside kev's Hub drop-in and + Laya's native head — not a fourth species. Their README: interface- + compatible baseline, **not a reproduction of Jev's undisclosed + model or training**. Softmax ignores level ordering; option order + can shift logits. Do not copy `install.sh`, ports, or model ids + into skill cards. + + **Evidence (README; Empirical as *their* named tables, Hypothesis + on *your* labels).** Official `typesafe-sdk==0.6.0` with only + `base_url` pointed at the server (4B backend) is the drop-in + proof. When `hf` is on they publish Wilson-CI / ECE / NLL tables + on templated sets (set1 / set2 / set3); do not re-promote those + rows as a ranking of TypeSafe Jev. Head-to-head vs published + jev-1.13.0 on 5 questions: 4/5 top-answer agree; payout + billing-vs-technical is an ambiguous prior, not a prompt bug. + Cards: `judgment-class.md`; `faq.md`; `validation.md`. + +4. **[`hitakshiA/solari-reflex`](https://github.com/hitakshiA/solari-reflex)** + (MIT, TypeScript, 0★, created 2026-09-18T15:24Z). Computer-use + **speed layer** on Solari (browser + Linux desktop): **one + structured observation → one typed decision → one verified + action**. **No screenshots in the loop.** Observation is + numbered controls / a11y tree / visible text (Calc: used visible + rows, never 2³¹ cells). Decision: which operation, which target + (speculative, same request), done/blocked. Write: a small model + only for TYPE as strict JSON. Act: guard check, refuse covered + controls, pipeline input; model output **never** becomes a + selector, coordinate, or script. Deny lists are **absent from the + question**, not merely disfavoured. Sibling of + `browser-use/jev-ultrafast` / jev-use / typesafe-computer-use. + Pattern: **perception → decision → act** with Harbor-style + measurement (score the task; independently checked by the app / + Stripe API / answer key). + + **Evidence (README + [solari-fast-showcase](https://github.com/hitakshiA/solari-fast-showcase); + vs Codex CLI GPT-6 Astra on the same Solari machines; not re-run).** + Stripe Checkout (qty 2, promo, card): **60.2 s**, $0.011 vs + **194.9 s**, 34 tool calls. Six different Stripe checkouts: + **66 s**, $0.064 vs **460 s**, 86 tool calls. 30 Calc expenses: + **24.2 s**, $0.0008 vs **98.4 s**, 77 tool calls. Ratios on that + table sit in ~3–7×. Their step table: decide ~400 ms / ~$0.0001; + TYPE write ~600 ms. **0.6** is *their* Advisor handoff, not a + universal threshold. Do not copy npm git-install. Cards: + `mixed-architecture.md`; `validation.md`; `applied-mappings.md` §2; + `agent-self-assessment.md`. + +5. **[`ktaletsk/jevframe`](https://github.com/ktaletsk/jevframe)** + (MIT, Python 3.10+, PyPI `jevframe`, pandas **and** Polars, + created 2026-09-18T17:06Z, 0★). `.jev` accessor: `noul` / `choice` + / `score` / `evaluate` with **full probability columns** (`p__…`), + preserved row order/indexes, one input row per request, questions + about a row share the request, default `max_concurrency` 16. + **No result is thresholded or silently renormalized.** `score` is + the expected zero-based level, not a probability. Sibling of + [`yalindogusahin/jevpandas`](https://github.com/yalindogusahin/jevpandas) + (`notes.md` §46): both are dataframe-native semantic columns; + jevpandas is the store-as-index cousin (`evaluate`/`filter`/ + `classify`/`score`); jevframe is the accessor + Polars + packed + struct layout. Neither is SQL `jev()`. Pattern: **dataframe-native + semantic columns** (class, not vendor). v0: no chat, no generated + records, no custom dtypes. Do not copy the client. Cards: + `mappings.md` §4; `mixed-architecture.md` gallery. + +### MED (pointers, not cards of their own) + +6. **[`de-niji/jev-hermes`](https://github.com/de-niji/jev-hermes)** + (MIT, Python, 0★). Intent **gate before a Hermes turn**: + `calendar` / `mail` / `status` → config + flat tools, **no memory + search spam that turn**; `complex` / people / prefs / "what did + we…" → memory + normal agent. Memory providers still **write** in + the background. Token savings come from skipping long tool/memory + *tours*, **not from turning memory off**. Pattern: **route ≠ + memory**; S1 gate preserves S2+memory for hard routes. OpenRouter + `POST /api/alpha/decisions` (same surface as jev-decision-layer). + Do not copy `docker cp`. Cards: `applied-mappings.md` §5; + `mixed-architecture.md` dual orchestration. + +7. **[`ngallodev-software/agent-workflow-typesafe-ai`](https://github.com/ngallodev-software/agent-workflow-typesafe-ai)** + (Apache-2.0, Python, 0★). Advisory-only host plugin. Projects + bounded redacted evidence into typed questions; normalizes answers + into versioned secret-free **semantic receipts**. **Never changes + host routing**, executor, model policy, lifecycle, evaluation, + review, or acceptance. Missing credentials / SDK / uncertain + answers / failures → distinct `no_action` outcomes. Pattern: + **soft sidecar receipts** (fail-open evidence). Complementary to + Abide (Abide is in-session Score on a diff; this is advisory + metadata the host may ignore). Do not copy the TOML. + +8. **[`Joymfl/dag-jev`](https://github.com/Joymfl/dag-jev)** + (Rust + petgraph; README empty this pass; 0★; created + 2026-09-18T20:42Z). GitHub description: DAG from unordered items + via Jev. Source: pairwise "does task i depend on j?" over a bag + of numbered steps (reads/writes in `input.txt`); answers intended + to build a `petgraph`. Experiment; graph wiring incomplete in + `main.rs` this pass; **no metrics**. Pattern: **structure + induction over bags** — code owns the DAG; the model only answers + pairwise (or Choice) dependency questions. Do not clone the + request builder. Cards: `mappings.md` §9 / §5. + +9. **[`knowlet/jev-agentworld-web-simulator`](https://github.com/knowlet/jev-agentworld-web-simulator)** + (MIT, TypeScript, 0★). Fictional web: Jev Choice for **search + intent** and (same request) **layout / palette**; an + OpenAI-compatible generator writes `SearchDocument` / + `PageDocument`; Zod validates; a **deterministic** compiler emits + json-render spec; SQLite is the world. Model cannot add + components or handlers. Mock mode does **not** silent-fallback + to live. Pattern: **decision for control, generator for content** + in a simulated world. Live smoke (when run) is an integration + test (budget **3 Jev + 3 generator**), **not** calibration or + world-consistency. Offline CI: 27 unit/contract + 2 Chromium + browse tests; [CI #2](https://github.com/knowlet/jev-agentworld-web-simulator/actions/runs/35391369860) + on `b1dd6d7` passed. Do not copy `.env`. Cards: + `mixed-architecture.md`. + +10. **[`ufx7/jev-testbench`](https://github.com/ufx7/jev-testbench)** + (LICENSE absent this pass; TypeScript; 0★). Two tools: black-box + **determinism / latency / context / concurrency** (`src/bench/`); + collaboration harness (`src/collab/`) with three arms — + `llm_autonomous` (unconstrained; illegal actions tracked), + `scripted_plus_jev` (code enumerates legal actions; Choice + picks; **no LLM**), `llm_plus_jev` (LLM proposes; Choice + arbitrates on the legal set; low p escalates then **stops** + rather than guessing). **Jev is not a peer arm.** Reports + Wilson intervals, exact McNemar per level (p < 0.05 **and** ≥5 + discordant pairs), ceiling-effect flags, cost/latency **per + model**. Pattern: bake this into the jevals/Harbor **measurement + curriculum**, not a second product. Grid-task demo uses a mock + LLM to prove the harness; swap before trusting a real model. + Cards: `validation.md`. + +11. **[`alexykn/jevscan`](https://github.com/alexykn/jevscan)** + (MIT, Python 3.12+, `0.2.0rc4`, 0★). Tree-sitter extracts + lexical units (Python/Rust/Perl/TS/JS); typed questions on + those targets; **does not execute or import the scanned code**. + Release candidate: **not a claim of calibrated semantic + accuracy.** Pattern: **structural AST + semantic judgment + compose** (sibling of riff / JevLint / Abide). `tenbin` still + owns the lint *skill*; this is a recipe of AST∩remainder, not + a second Augustus skill. Do not copy YAML. Cards: + `toolbox-mapping.md` spec/lint; `mixed-architecture.md`. + +12. **[`phin-tech/pi-jev-approver`](https://github.com/phin-tech/pi-jev-approver)** + (LICENSE present this pass; TypeScript; 0★). Pi **shell safety + gate**. Facts (git branch, path scope, registry-publish regex) + computed in code; typed Score `risk_level` + Nouls + Choice + `primary_concern` on the remainder. `commandRules` regex + **short-circuits** allow/deny — a `deny` rule is a hard block + no LLM or human prompt can overturn. **No key → fail closed.** + Optional LLM escalation can only reduce how often you are + asked, never replace the human as last resort. Author + side-by-side (their `jev-test` project): classification + **~500 ms** vs **~2.5–3.5 s** chat-model. Live-tested: `aws s3 + rm` 1.72/2, `ec2 terminate-instances` 1.80/2 in the ask band; + read verbs matching `action: allow` never called the model. + **Light note only** — rh-guard-adjacent (coding-agent tool + gate), different remainder from jevgate (unlisted verbs) and + Abide (soft project rules). Do not copy the Pi install. + +13. **[`Mattepiu/laya-onnx`](https://huggingface.co/Mattepiu/laya-onnx)** + (Apache-2.0 ONNX export of [`convaiinnovations/laya`](https://huggingface.co/convaiinnovations/laya); + Hub likes **1** this pass; updated 2026-09-18T10:45Z). Open + **replica deployment path**: non-autoregressive marker-token + head in onnxruntime. Card example: **~15 ms on CPU** for one + Noul. **Do not copy the inherited vs-Jev accuracy table** — + those rows remain vendor claims (`notes.md` §18, + `judgment-class.md`). Pattern: ONNX/runtime port of a trained + decision-only head, beside Laya native and kev Hub. Not Archer. + +### Spotcheck (not a fold) + +- **[`TheoLeeCJ/SemIf`](https://github.com/TheoLeeCJ/SemIf):** **1551★** + this pass (2026-09-18T20:51Z). Watch cited awesome claim **1491** + → **+60**. Independent; not affiliated with Jev/TypeSafe. +- **[`vinnylarouge/jevlike`](https://github.com/vinnylarouge/jevlike):** + **866★** this pass. +- **Awesomejev 488 / 21644:** watch cited "unchanged." This pass did + **not** independently re-derive 488/21644. + [awesomejev.com](https://awesomejev.com/) still showed **410 + entries / 10,093 stars** (refreshed 2026-09-17) in the public + snapshot fetched here. Do not invent a new count. +- **Tracker** [`multimodalart/jev-reproductions-tracker`](https://huggingface.co/spaces/multimodalart/jev-reproductions-tracker): + Hub `lastModified` **2026-09-18T20:12:57Z**. Space `models[]` + still lists **`convaiinnovations/laya`**. **`BlackwoodAI/blackwood-rlcd` + is not in that array.** Archer remains a *promised* trained-head + card. Omni decide that shipped (blackwood) is still off the + tracker this pass. + +Cards: `judgment-class.md`; `validation.md`; `applied-mappings.md` +§2 / §5; `mappings.md` §4 / §9; `mixed-architecture.md`; `faq.md`; +`mental-models.md`; `methods-catalog.md`; `toolbox-mapping.md`; +`agent-self-assessment.md`. No wrapper. + +## 49. Boundary map, Harbor bake-off vs constrained LLMs, dual-process (2026-09-18 ~15:52 Boise) + +America/Boise ~15:52 = 21:52 UTC (archive 214908). Docs-only fold into +open PR #2 (`cursor/augustus-store-envelope-00b4`). Not a competing +PR. Archer 27B drop still **WATCH** (still ~2026-09-19; not landed). +X MCP flap blocked discourse this hour — no tweets invented. Identity +lock vs `typesafe-ai` / `tenbin` / `decision-first` holds. No wrapper, +no install.sh, no copied `uv` / `von serve` / pip how-tos. No invented +metrics. Do not re-fold §48 items, blackwood-rlcd, open-jev-laya-bench, +decision-token LoRA, jevgate, wellposed, jev-reflex-autonomy-lab, Abide, +kev species, jevpandas, bitrate-advisor. + +Backend-agnostic: these are **placement / measurement** cards +(extractable-from-state axis; Harbor-style frozen-protocol bake-off; +recompute-from-logs feedstock; Kahneman S1 decide / S2 generate; +combinatorial assembly ≠ extractive keep/drop; packed one-forward open +LLM economics; tiny non-AR local surface). TypeSafe Jev is the +exemplar in the READMEs, not a monopoly. kev is a **star/activity +delta only** this hour (100★ this pass; user cited 97). + +### HIGH + +1. **[`Zaious/jev-capability-atlas`](https://github.com/Zaious/jev-capability-atlas)** + (README claims MIT; GitHub SPDX **NOASSERTION** this pass; Python; + created 2026-09-18T21:30:40Z; 0★; unofficial, not TypeSafe). + Independent **when-it-holds map with real API receipts**, not a + leaderboard. They cite [jev-benchmarks](https://github.com/AbdelStark/jev-benchmarks) + and thaiexam charts; they do not redo them. Primary mental model + this hour: **place the task on one axis** — + + > Is the correct answer fully recoverable from the `state` you hand + > the model, or does it require outside knowledge that is not in + > `state`? + + **Self-contained (strong):** classification they cite from + jev-benchmarks (AG News 91%, Banking77 **87%** — a *different* + protocol from DMB 76.3% and jevals.com 79.67%; do not merge); + citation support-checking with claim + quote both given + (`paraphrase_support` and `reversed_meaning_high_overlap` both + correctly judged); sarcasm when the trigger is in the given text. + **Not self-contained (fails, often confidently):** pure recall + without a supporting passage; scoring that needs a whole-field + comparison; genuinely overlapping categories. + + **History suite** (`suites/history-recall-context/`, receipt + `runs/2026-09-19.json`; N=3, single annotator, Chinese history + only — qualitative axis proof, **not** a knowledge-breadth + estimate). Cite the table, not the README's 30-second slogan: + + | Case | Question | Context? | Pick | Conf | Dist | Correct? | + |---|---|---|---|---|---|---| + | A | 2nd Qing emperor | No | Yongzheng | 0.90 | 0.04 / 0.93 / 0.03 | ✗ (Kangxi by popular convention; ground truth itself contested) | + | B | 7th Qing (obscure) | No | Xianfeng | 0.07 | 0.25 / 0.37 / 0.38 | ✓ by luck (near-flat) | + | C | Same as B | Passage in state | Xianfeng | 0.97 | 0.00 / 0.02 / 0.98 | ✓ | + + Informal unsaved B rerun picked Daoguang at 0.08 — both near-flat. + Teaching: **bare memory is unreliable; reading comprehension over + supplied text is reliable.** Conflating the two misjudges the risk. + Retrieve first; put the passage in `state`. + + **Not a state machine internally** (distributed LM understanding — + paraphrase vs reversed-meaning contrast). **Correctly placed as a + component node** in *your* code ("a node in your state machine" is + the architectural instinct; "its internals are a state machine" is + not). Confidence is a **statistic from the distribution** (RLCD + trains the distribution; `confidence` is arithmetic on how peaked + it is), not a second trained correctness score. Calibration is + **population-level** and can fail **dangerous-high**: DAIR Emotion + via jev-benchmarks — 48% acc, mean conf **0.819**, 16% of items + p(correct)=0. That is the "blind guessing" concern that actually + lands: overlapping blurred categories, overconfident. + + **Browser-use strength is placement, not vision.** Third-party + `jev-ultrafast` (Browser Use): Google Flights 9.5s → 7.1s; 12-task + vs Playwright MCP 1.5× faster / 1.6× cheaper, comparable accuracy; + standalone loop ~1.8s / $0.0005 / 97% (their figures, not + re-run). Jev is text-only. What held: **DOM snapshot as `state`** + (a visual task translated into extractive text) + **speculative + fan-out** over candidate elements. Not screenshots. Same + component-node placement as lizard-agent / solari-reflex. + + Pattern: **boundary map / placement judgment.** Cards: + `mental-models.md` (primary); `faq.md`; `question-design.md`; + `applied-mappings.md` §2 / §6; `mappings.md` §3 / §6; `mixed-architecture.md`. + +2. **[`nibzard/decision-model-benchmark`](https://github.com/nibzard/decision-model-benchmark) + (DMB)** (LICENSE **absent** this pass; Python; created + 2026-09-18T21:48:07Z; 0★; no vendor sponsorship). Independent + **frozen protocol**: typesafe:jev vs **8 constrained LLMs** vs + **3 deterministic baselines** (keyword / majority / random). Five + suites, 60 cells, **$28.34** measured spend; every raw log + published. **`results/v2/v2.md` is the report of record** (v1 + majority-prior defect corrected; mixed protocol versions merged + by explicit `--allow-protocol-mix`). Harbor-style measurement + exemplar: protocol frozen before the run; negative results ship; + recompute from logs; unknown usage is never a measured zero; later + runs replace cells whole. Do not copy `uv` how-to. + + **jev (v2, named receipt, not a ranking):** + + | Suite | Acc | Notes | + |---|---|---| + | S1 banking77 77-way | **76.3%** | ECE 0.083; cost/1k **$0.07**; p50 **274 ms** | + | S2 SMS spam | **93.0%** | ECE 0.249 (overconfident vs GLM 0.042) | + | S3 cardinality | **100%*** | valid coverage **72.7%** (225 failed = **256+ Choice cap**, `400 Too many choices`); 254–255 still 100% | + | S4 order permute | **76.7%** | flip **13%** (worst LLM 37%) | + | S5 honesty | **14.3%** | admits-ignorance **49.7%** vs most LLMs 97.3–100% (gpt-5.4-mini **64.7%**); ECE **0.246**; mean conf on no-good **0.543** | + + p50 across jev cells **264–276 ms**, flat 2→255 options. Fastest + *measured* vs thinking-mode LLMs is **10–16×**, **1.2×** vs + gpt-oss-120b on Cerebras — **not** the vendor 40–200× claim + against LLMs left in their slowest default. gpt-oss-120b banking + **81.3%**; glm-5.3 **80.4%** / spam **94.9%**. **No class wins on + quality.** Axes that *do* separate: latency, cost, schema-validity + (jev 0% malformed), Choice cap, honesty. Two OpenAI models sit + *below* the 87.7% majority baseline on spam. + + **Do not collapse Banking77:** atlas/jev-benchmarks **87%**, DMB + **76.3%**, jevals.com 2026-09-18 **79.67%** — n / split / protocol. + Cite named receipts. Pattern: **Harbor/jevals practice + + when-to-use-vs-constrained-LLM table.** Cards: `validation.md`; + `judgment-class.md`. + +3. **[`Jevals/jevals-data`](https://github.com/Jevals/jevals-data)** + (CC-BY-4.0; created 2026-09-18T21:36:00Z; 0★). Release boards + + per-decision JSONL run logs + suite files behind + [jevals.com](https://jevals.com). Cite "Jevals (jevals.com), + release \". **2026-09-18** board: suite **0.1.0**, **8 + systems**, tasks banking77 / helpsteer2 / pubmedqa. Formulas at + https://jevals.com/methodology/. **Recompute-from-logs pattern**, + not a third ranking to merge with DMB or atlas. + + Jev on *this* board (native probabilities; n=300 × 5 repeats; + **not** DMB n): banking77 acc **0.7967**, ECE **0.0981**, p50 + **467 ms**, cost/1k **$0.043**; helpsteer2 Score acc **0.4127** + (label-prior 0.4167 — barely above chance on that primitive); + pubmedqa Noul acc **0.9127**, ECE **0.0504**, p50 **438 ms**. Do + not dump the board as a ranking. Feedstock for Harbor/jevals: + frozen suite files, item ids pointing at public datasets (item + text not republished), run header + per-decision rows + (`item_id`, `epoch`, `target`, `order_seed`, `output`, + `usage`, `cost_usd`, `seconds`, `malformed`, `refusal`, + `retries`). Cards: `validation.md`. + +4. **[`taro1985/dual-process-ai`](https://github.com/taro1985/dual-process-ai)** + (MIT; Python; created 2026-09-18T20:55:08Z; 0★). Explicit + **Kahneman S1 (Jev) / S2 (Gemini)** design pattern. S1: typed + decision + confidence, ~70–500 ms. S2: free-form text. Mechanism: + `confidence ≥ τ → S1 decides; else escalate to S2`. **Routing + fails open** (low conf → S2; the router never refuses). **Safety + fails closed** (unparseable denied). **Routing accuracy is not + measured yet** (misroute rate / escalation rate / Brier on + held-out — the numbers this project needs and does not have). + Without a key it runs **degraded keyword mode** — **not an + equivalent S1** (no calibrated confidence; anything not on the + allowlist escalates). Do not copy hooks / Discord / pip. Crossover + metaphor for **business/life**, not only SWE: cheap classify / + route / gate on S1; write / reason / generate on S2. Same split as + jev-reflex-autonomy-lab (S1 keeps control) and jev-hermes (route ≠ + memory), productized as a cascade. Cards: + `mixed-architecture.md`; `toolbox-mapping.md`; `mappings.md` §2; + `agent-self-assessment.md`; `faq.md`. + +### MED (brief) + +5. **[`simonmesmith/jev-arc-agi-v1-experiment`](https://github.com/simonmesmith/jev-arc-agi-v1-experiment)** + (LICENSE **absent** this pass; Python; created 2026-09-18T21:32:19Z; + 0★). Direct Jev on ARC-AGI-1 public eval: **4/400 (1%)** fully + solved; **1.125%** task-weighted (4 + 0.5 / 400); **5/419** exact + grids; **~$2.32**; **10 minutes**; `jev-1.13.0`; two guesses per + grid. **Cell-wise Choice assembly:** height/width 1–30, then one + of ten colours per cell; cells do not see one another; transpose + for the second guess. Dimensions ~**90%** on first attempt; rarely + a complete grid. No partial credit for cells. Teaching: + **combinatorial grid tasks ≠ extractive keep/drop.** A program + library with no competing candidates **never called Jev** (6 + tasks / 1.5%). Combined policy 2.375% includes hand-written + programs — tells you less about Jev. Frozen protocol, traces, + `PROTOCOL.md`. Not a ceiling claim. Cards: `validation.md`; + `mappings.md` §9; `faq.md`. + +6. **[`ikermoel/open-alternative-jev`](https://github.com/ikermoel/open-alternative-jev)** + (Apache-2.0; Python; 3★; Space + [`IkerMoel/open-alternative-jev`](https://huggingface.co/spaces/IkerMoel/open-alternative-jev)). + Packed **one-forward** System One on **any open-weights LLM** + (HF + vLLM). **Not a Jev reproduction** — packages a capability + chat APIs hide; no claim about how Jev works. RACE-H (250 + passages × 4 questions, n=1000) on Qwen3.6-27B 8-bit: packed + **92.9% @ 4.55 q/s** vs one-at-a-time 92.6% @ 1.66; 2.5× fewer + tokens. Interference **6–9%** of answers move vs a 2.7% numeric + noise floor; order rotation 8% MMLU / 2.4% RACE-H. Temperature + scaling ECE 5.4%→2.1% MMLU, 2.8%→1.1% RACE-H. Small 4B packing + costs 2.8 points — use `separate` when accuracy at stake. On + vLLM, prefix-cache `separate` is fastest. Economics/architecture + of **open replicas on the constrained-AR / logprob path**, not + trained decision-only. Cards: `judgment-class.md`. + +7. **[`wfzyx/von`](https://github.com/wfzyx/von)** (Apache-2.0; + Python; 3★). **14 MB** Needle 3 SAN; non-AR local `POST + /v1/systemone` drop-in; sub-15 ms CPU *claim* / ~**38 ms** embed + in their table; ~28 MB RAM. Default needle **52.6%** balanced acc + on OpenJev `authored144` — **not a calibrated Jev replica**. Other + backends (berta-v3 / modern / laya) are optional heavier heads. + **Do not copy the vs-Jev ranking table** (includes a speculative + Jev weight estimate). Distinguish: jev-local **stub until hf**; + kev **trained pointer** on Qwen2.5-0.5B; von **tiny SAN** at the + extreme of the speed/econ class. Cards: `judgment-class.md`; + `faq.md`. + +8. **[`jaredpalmer/kev`](https://github.com/jaredpalmer/kev) delta** + — still active this hour; **100★** this pass (user cited 97; + prior fold 61★). Pushed through 2026-09-18T21:51Z. **Light note + only.** No species rewrite. Hub weights + NOTA training remain + §45. + +### Not this hour + +Archer drop **not landed**. X discourse **blocked** (MCP flap) — no +invented tweets. Do not re-fold §48. + +Cards: `mental-models.md` (boundary map); `validation.md` (DMB + +jevals-data + ARC); `judgment-class.md` (vs constrained LLM; von; +open-alternative-jev); `mixed-architecture.md` (dual-process; +component node; DOM-as-text); `faq.md`; `mappings.md` §2 / §3 / §6 / +§9; `applied-mappings.md`; `question-design.md`; `methods-catalog.md`; +`toolbox-mapping.md`; `agent-self-assessment.md`. No wrapper. + +## 50. GLiNER2.5 extractive compaction — encoder backend, same keep/drop job (2026-09-18 ~16:22 Boise) + +America/Boise ~16:22 = 22:22 UTC. Docs-only fold into open PR #2 +(`cursor/augustus-store-envelope-00b4`). Not a competing PR. Archer +27B drop still **WATCH**. Identity lock vs `typesafe-ai` / `tenbin` / +`decision-first` holds. No wrapper, no `--plugin-dir` / `uv` how-to, +no copied timeouts or Hub download scripts. No invented metrics. +**Not Jev. Not multimodal.** Do not re-fold §48 extractive recipes, +§49 bake-off, Abide, kev, GLiGuard as a species rewrite, or +pi-jev-compaction as a new product. + +Backend-agnostic: this is a **compaction / context-sieve placement** +(pointer keep-drop, soft Choice under a hard mutation envelope, +shadow-mode rollout). TypeSafe Jev is one backend for that *job* +(`tamaratran/fast-jev-compaction`, `vava-nessa/pi-jev-compaction`); +GLiNER2.5 is another. Augustus stays family-first. + +### HIGH + +1. **[`m-newhauser/gliner25-compaction`](https://github.com/m-newhauser/gliner25-compaction)** + (Apache-2.0; Python + Claude Code plugin; created + 2026-09-18T17:22:34Z; 1★ at capture). Local, **evidence-first** + context compaction. Default checkpoint + [`fastino/gliner2.5-base-v1`](https://huggingface.co/fastino/gliner2.5-base-v1) + (checkpoint model card Apache-2.0 per their README) chooses a + retention action for **completed** eligible tool interactions and + extracts **exact source spans** when the full result is unnecessary. + **Not a prose summarizer.** User and assistant text unchanged. + Retained evidence is copied from original **character offsets**. + Mutating tool interactions are preserved in full. Inference runs in + a local Python worker after Hub download; transcript analysis does + not require a remote inference API. + + Per completed pair, one retention action (closed set): + + - `keep_full` — complete call and result + - `keep_evidence` — exact excerpts from the result + - `keep_call_only` — call stays; result replaced with a rerun notice + - `drop` — paired call and result removed + + Inputs: current goal, nearby conversation, tool name and input, + original tool result. Deterministic safeguards override uncertain + predictions, protect validated diagnostic spans, preserve recent + messages and mutations, validate exact offsets, and reject orphaned + tool results. + + **Four load-bearing mental models (architecture, not a plugin + catalog):** + + 1. **Pointer / extractive, not generator.** Same family as + [testimonial-miner](https://github.com/AppitStudio/testimonial-miner) + and [jev-reviewer](https://github.com/choxos/jev-reviewer) + (`notes.md` §48): the model **selects**; code **copies + verbatim**. Compaction that invents a prose summary is a + different (worse) species for auditability. GLiNER locate is not + a footnote here — character offsets *are* the keep/drop + candidates. One local multi-head does **categorize** (which + retention action) and **locate** (which spans) in the same job + (`judgment-class.md` species map). + + 2. **Soft retention Choice under a hard envelope.** Code owns the + mutation monitor: mutating tools, unknown shell, and shell + control operators / pipelines / substitutions / redirections are + treated as mutating → `keep_full`. Missing or invalid evidence + **fails closed to `keep_full`**. Low-confidence retention + predictions likewise fail closed to `keep_full`. Contrast: many + Jev *preference / remainder* gates fail-open (Abide; jevgate + cannot block). Compaction *drop* (and lossy `keep_evidence`) is + the irreversible act, so the authorized reduction fails closed. + From the evidence-preservation view the outcome looks like + context-sieve "keep on error" (`applied-mappings.md` §1) — name + the *act*, not the slogan. Conservative shell over-retention is + their documented limit, not a bug to "fix" by failing open. + + 3. **Same compaction job, encoder backend.** + [`tamaratran/fast-jev-compaction`](https://github.com/tamaratran/fast-jev-compaction) + asks two Nouls (should the *call* stay? should the *result* stay + verbatim?). [`vava-nessa/pi-jev-compaction`](https://github.com/vava-nessa/pi-jev-compaction) + is the Pi cousin: verbatim drop, never summarize. Here GLiNER2.5 + chooses discrete retention actions + evidence spans. Fastino / + GLiGuard sibling *class* (schema-in-encoder, local) — not a + GLiGuard safety-schema clone, not a Jev Score, not a Noul. + Augustus does not pick a vendor for the hole. + + 4. **Shadow mode as safe rollout.** Public default `shadowMode: + true`: local analysis logs the proposed reduction **without + replacing session history** until explicitly set false. Same + rollout instinct as jev-harness / is-malicious (log would-do + first). Compaction mutates memory; shadow is the default because + a bad drop is not a reversible token cost. + + **Limits (theirs, README; experimental).** Reduction measured in + **characters, not tokens**. Conservative shell policy may retain + commands that are actually read-only. Only completed + tool-call/result pairs are candidates. Domain-specific tuning and + broad production evaluation remain future work. No published + token-reduction or retention-quality rates this pass — do not + invent them. Their config defaults (`minimumConfidence` 0.7, + `minimumEvidenceConfidence` 0.5, `minReductionRatio` 0.25, + `preserveRecentMessages` 6) are *their* knobs, not class constants. + + **Siblings — complementary, do not merge.** + + - **`24601/rh-guard`:** reward-hack / eval-integrity on tool use. + Shared notes only: fail-closed retention, hard shell mutation + policy, shadow-mode rollout. Different hole. This is not + reward-hack detection. + - **GLiGuard:** Fastino encoder sibling (safety-schema classify). + Compaction is locate+categorize on tool transcripts, not LLM I/O + moderation. + - **Abide:** same Claude Code hook-host surface; Abide is fail-open + Score on diffs; this is fail-closed `keep_full` on compaction. + + **Placement.** Context sieve + exact-text keep/drop + (`applied-mappings.md` §1–§2). Pillar: selective classification / + SDT criterion (false drop >> false keep) + runtime-assurance + sandwich (mutation monitor in code). Hole: sieve / keep-drop. + Family: GLi\* encoder (GLiNER2.5 local multi-head). Fail-closed on + the reduction. Eval path: none published this pass (experimental); + characters≠tokens is the honesty constraint. **Empirical** as + README behavior. **Hypothesis** that the same envelope transfers to + *your* transcript domain. Cards: `judgment-class.md` (primary); + `applied-mappings.md` §1–§2; `mappings.md` §12 / §18; + `mixed-architecture.md`; `faq.md`; `mental-models.md`; + `methods-catalog.md`; `toolbox-mapping.md`; + `agent-self-assessment.md`. No wrapper. + +### Omni / Jev-omni + +Not multimodal. Usage: extractive context management as a +perception/memory hygiene stage **before** decide. Archive + +landscape pointer. + +## 51. CI merge-gate, fail-open wake VOI, S1 indexer, claim-evidence Stop (2026-09-18 ~16:48 Boise) + +America/Boise ~16:48 = 22:48 UTC. Docs-only fold into open PR #2 +(`cursor/augustus-store-envelope-00b4`). Not a competing PR. Archer +27B drop still **WATCH**. Identity lock vs `typesafe-ai` / `tenbin` / +`decision-first` holds. No wrapper, no npm/wrangler/install.sh/devenv +how-to, no copied ports, thresholds as class constants, or invented +metrics. Do not re-fold §50 GLiNER2.5 compaction, §49 bake-off, Abide, +kev, pi-jev-approver, jevgate, or rh-guard as a species rewrite. + +Backend-agnostic: this hour is **decision-model as classifier / +gate / coordinator** (flaky-vs-real CI, VOI resume, S1 extract + +escalate-S2, claim/evidence integrity, Harbor on/off routing, +policy-as-judgment). TypeSafe Jev is the documented exemplar, not +the monopoly. Augustus stays family-first. + +### HIGH + +1. **[`CaseReed/latch`](https://github.com/CaseReed/latch)** + (MIT; TypeScript; created 2026-09-18T22:09:12Z; 0★ at capture). + Merge-gate triage for a **finished** red test run (Playwright / + Jest / pytest / JUnit XML): cluster failures by signature in + **code**, TypeSafe Jev labels each cause (one call per cluster, at + most 8; cached free), **code owns** Gate: PASS (infra noise) vs + Gate: BLOCK (real failure). The judge never gets to say "ignore" + alone. `ignore_as_infra` needs `env_cascade` + an infra + fingerprint. Missing key still prints clusters (`needs_human` / + `no_key`) and **never fails Playwright** — the reporter is + fail-open; `--gate` is a **separate** CI step. + + Demo (offline, README): 8 identical connection errors → 1 cause → + Gate PASS; 5 failing assertions → Gate BLOCK. Policy (theirs, not + class constants): no key / API error / `cause.confidence < 0.55` / + `same_root < 0.5` → `needs_human`; `env_cascade` + `same_root >= + 0.7` + fingerprint → `ignore_as_infra`; flake / locator_drift → + `fix_test`; `assertion_bug` and (`blocks_merge >= 0.55` or Jev + `action = fix_product`) → `fix_product`. `blocks_merge` sits in a + noise band ~±0.03; 0.55 is in the empty gap (their calibrate + claim). + + **Limits (theirs).** Grouping is message-based; a logic regression + fragments. Measured on `pallets/click`: 2 real regressions → 13 + failures → **10 clusters**. The "70 → 1" figure is an infra-cascade + property, not a general one. Signature = `apiName` + first 80 chars + of the normalized error. Error text is redacted in every output. + + **Four load-bearing mental models:** + + 1. **Cluster in code, judge labels, policy decides.** Same split as + OpenSmoke (scan every step; LLM autopsy only on flags): the + model estimates a *cause class*; the merge act is a table. + Internals are not a state machine; placement is a component + node (`notes.md` §49). + 2. **Decision-model as CI flaky-vs-real classifier.** The hole is + SDT criterion (false PASS on a real bug >> false BLOCK on + infra), not "make CI smarter." Pair with Harbor (frozen + artifacts × PASS/BLOCK labels; independent verifier) and + rh-guard (eval-integrity — do not let the agent game the tests + latch is classifying). Different holes; shared notes only. + 3. **Fail polarity is per surface.** Reporter never fails + Playwright (false block of the *run* loses evidence). `--gate` + fails closed on merge when a cluster is not confirmed infra. + Name the *act*. + 4. **Do not copy the reporter.** Thresholds, ledger path, and npm + wiring stay theirs. + + **Placement.** Environment / harness triage + (`applied-mappings.md` §3) + decision circuits (`mappings.md` §3) + + SDT (`mappings.md` §7). Pillar: SDT + Leveson sensor≠constraint. + Hole: triage / gate. Family: closed decision API. Eval path: none + published as Harbor this pass (demo + unit goldens). **Empirical** + as README behavior / demo. **Hypothesis** that the same cluster → + label → policy transfers to *your* runner. Cards: + `applied-mappings.md` §3; `mappings.md` §3 / §7 / §18; + `mixed-architecture.md`; `validation.md`; `faq.md`. No wrapper. + +2. **[`shitianfang/wakegate`](https://github.com/shitianfang/wakegate)** + (MIT; TypeScript; created 2026-09-18T22:13:11Z; 0★). Fail-open + **wake gate** before resuming a sleeping agent (Workers / Durable + Objects / Node). One typed question: given `waitingFor` plus an + optional event/observation, is this worth a full LLM turn? Code + owns sleep duration, skip counters, and force-wake. Skip only if + Jev answers **and** p(wake) < 0.2. Every other path wakes: + `fromUser`, nothing-to-judge, skip-limit (default 10), error / no + key / timeout (5 s), unsure band 0.2–0.5, p ≥ 0.5. + + **Eval (theirs; discount).** 21/21 on 21 hand-written scenarios + (11 wake / 10 sleep); p50 253 ms, p95 519 ms (n=21). Two of the + 11 correct wakes came only from the unsure band (calendar invite + 0.38; sold out 0.48). Same person wrote the scenarios and the + question. First yes/no wording scored 16/21 on this set. Current + three-way Choice picked on a separate 16-scenario **dev** set + (not in repo). **Smoke, not a benchmark.** Savings unmeasured. + + **Contrast (fail polarity, not products).** + [`phin-tech/pi-jev-approver`](https://github.com/phin-tech/pi-jev-approver) + fails **closed** without a key (tool remainder). jevgate **cannot + block** (allowlist proves; remainder fail-open). Compaction + (`notes.md` §50) fails closed to `keep_full` because *drop* is + irreversible. Wake *skip* is the irreversible act here (the agent + stays asleep), so the authorized skip is the rare, high-confidence + path; everything else wakes. Horvitz mixed-initiative / VOI: pay + for the LLM turn only if EV(decision) beats the token cost; a + regex / user-message / skip-limit already answers without a model + (meta-VOI). Event/observation are untrusted; `maxSkips` bounds + sleep, "neither is a security boundary" (their README). + + **Placement.** VOI / gather (`mappings.md` §6) + durable-agent + control (`mappings.md` §14, Hypothesis) + mixed-architecture + per-action fail table. Hole: gate / abstain. Family: closed + decision API. **Empirical** as README safety table. **Hypothesis** + as production savings. Do not copy wrangler/npm. Cards: + `mappings.md` §6 / §18; `mixed-architecture.md`; `faq.md`. + +3. **[`GreyssonEnterprises/s1-graphify-indexer`](https://github.com/GreyssonEnterprises/s1-graphify-indexer)** + (+ sibling [`s1-indexer`](https://github.com/GreyssonEnterprises/s1-indexer); + created 2026-09-18T22:23:04Z / 22:21:15Z; 0★; **license not on + GitHub this pass — do not invent**). Local-first semantic + indexer: small zero-shot System-1 models build a knowledge graph + of a repo; an LLM is used only on the ambiguous tail, and **only + when the System-1 backend actually loaded**. Default backend + `gliner2`. Stubs `jev` and `needle` ship `available=False` until + implemented. GitHub one-liner "10–50× faster than LLM-based + indexing" is a **target, not a measured speedup**. README: do not + treat it as a benchmark until the table is filled from one repo, + both backends, same machine. Table is TBD. + + If GLiNER2 cannot load: artifacts still written (file nodes from + the chunker, `run_status: degraded`), process exits nonzero, repo + is **not** dumped to `S1_LLM_CMD`. `query` tokenizes, matches node + names, walks two hops; no match prints `Insufficient evidence`; + **it does not invent edges**. `gliner2[local]` extra is heavier + than declared deps and is not pulled by default. + + **Mental model.** S1 zero-shot **locate/extract** (GLiNER spans / + relations) on the bulk; escalate-to-S2 only on low-confidence + remainder — same sandwich as allowlist ∩ remainder and as + dual-process `conf ≥ τ` → S1 else S2, here the expensive act is + *indexing tokens* not a chat reply. GLiNER-as-Jev-class cousin + (`judgment-class.md` locate vs decide): the indexer is not a Noul. + Same *family* as gliner25-compaction (encoder on code/text) with a + different hole (graph construction vs memory keep/drop). Do not + copy pip extras. + + **Placement.** Locate species + escalate-S2. Hole: perceive / + gather. **Hypothesis** as 10–50×. **Empirical** as degraded-load + and no-invent-edges README behavior. Cards: `judgment-class.md`; + `mixed-architecture.md`; `faq.md`. + +4. **[`VladyslavHontar/clear-head`](https://github.com/VladyslavHontar/clear-head)** + (MIT; Python; created 2026-09-18T22:15:08Z; 1★). Claude Code + **Stop** hook: Jev checks factual claims in the assistant's answer + against **what it actually read this session**. Anti-hallucinated- + done. Per claim, keyword-and-frequency retriever (not semantic) + sends matching tool-output *lines* — not whole files. Classifies + sentences as factual / proposal / recap / neither; then + supports / contradicts / doesn't-address. Blocks on contradicted, + or unsupported with **no** relevant evidence. Coverage below + `JEV_EVIDENCE_FLOOR` (default 0.3) is "nothing relevant found"; + above the floor, "Jev can't confirm a specific derived fact" is + usually their documented limit, not a block. `JEV_FIRM` default + 0.6: below that, logged, **never blocks**. + + **Limits (theirs).** Keyword retriever misses paraphrases and can + match stale session evidence on generic overlap (no recency / + topic-boundary). A true claim unread this session still flags + unsupported. Excerpts leave the machine. Do not copy `install.sh`. + + **Placement.** Done-check / claim–evidence entailment + (`agent-self-assessment.md`; `methods-catalog.md` NLI row). + Pointer family with jev-reviewer: the model judges against + retrieved lines, never against another model's prose. Hole: gate. + **Empirical** as README behavior. **Hypothesis** as transfer to + *your* transcript domain. Cards: `agent-self-assessment.md`; + `applied-mappings.md` §2; `faq.md`. + +5. **[`reification-labs/foreman`](https://github.com/reification-labs/foreman)** + (created 2026-09-18T22:46:57Z; 0★; **no license field in + `mix.exs` — do not invent**). GitHub description: "Parallel + specialist agents returning typed, calibrated answers behind a + single System Two foreman. Elixir/Phoenix, Jev/TypeSafe-native — + `{value, probability}` everywhere." **The checkout is a stock + Phoenix 1.8 scaffold.** README is the Phoenix generator starter; + `AGENTS.md` is Phoenix guidelines; `mix.exs` has Phoenix/Ecto/ + Bandit/Req — **no typesafe / jev dependency**. Fold as + **description-only greenfield**, not a measured product. Do not + invent an Elixir Jev API. + + **Distinguish** from the existing "foreman" *shape* in + `agent-self-assessment.md` (Kevthetech143/super-jev loop: + progress/stuck/complete → continue/stop/retry/verify; the model + estimates named probabilities; code owns hysteresis). Dual-process + cousin of [dual-process-ai](https://github.com/taro1985/dual-process-ai) + (`notes.md` §49): S1 specialists decide; S2 coordinates / writes. + Typed probability everywhere is the *class* claim, not a receipt + this pass. + + **Placement.** Mixed architecture / dual-process. **Watch / + description-only.** Cards: `mixed-architecture.md`; + `agent-self-assessment.md`; `mental-models.md`. + +6. **[`vinilana/jev-gateway-bench`](https://github.com/vinilana/jev-gateway-bench)** + (MIT; created 2026-09-18T22:29:58Z; 0★). Harbor-shaped bench for + sibling [`vinilana/jev-gateway`](https://github.com/vinilana/jev-gateway) + (MIT; product; fail-open if Jev down/slow/wrong key; never fails + the LLM request). Real coding agents (Codex / Claude Code) on + chess-engine tasks with Jev routing **on vs off**. Hidden verifier + (perft + targeted checks) the agent never sees. Chess chosen + because perft counts are published and one wrong rule changes + them. Tasks: `chess-engine` (build), `chess-bugfix` (five injected + bugs), `chess-san` (notation feature). Modes alternate; order + swaps between reps. + + **Preliminary one-run (author: first signal, not a measurement; + 2026-09-18; Codex 0.154 / `gpt-6-astra` / `jev-latest`; + `chess-bugfix`):** both 36/36 hidden checks; routing on 4 LLM req / + 76,678 in / 1,313 out / 35 s vs off 6 / 118,709 / 3,231 / 88 s; + Jev 4 calls ~$0.0008. Jev forced `exec` three times (p 0.91 / 0.99 + / 0.94) then switched tools off to answer (0.98). An earlier pair + the same day (pre token-metering fix) showed the same *shape* (4 + vs 6 requests; 40 s vs 69 s). With fewer than five runs per mode + the summary says so. A cheaper unsolved run is not a saving. A + wrongly forced tool can derail a turn. Do not copy npm/ports. + Public repo caveat: an agent with web access could find the + reference; tasks give no reason to look. + + **Placement.** Harbor/jevals practice (`validation.md`): taskset + (chess + hidden verifier) × harness (Codex/Claude) × runtime + (fresh gateway per run) × on/off treatment. Sibling of DMB + (class bake-off) and jev-testbench (collab arms). **Empirical** as + a *shape* and as one-run signal. **Hypothesis** as a cost/quality + claim. Cards: `validation.md`; `methods-catalog.md`; + `toolbox-mapping.md`. + +7. **[`LightningK0ala/jev-marshal`](https://github.com/LightningK0ala/jev-marshal)** + (created 2026-09-18T22:42:24Z). Description: "Repository rules for + pull requests, enforced by Jev." **Empty git repo** this pass + (default-branch 409). Fold as **Watch / description-only**. Cousin + of Abide / jev-pref / if-ai (policy-as-judgment on a PR). No + metrics, no fail polarity, no license to invent. + +8. **[`LilDojd/jevons`](https://github.com/LilDojd/jevons)** + (MIT; TypeScript; created 2026-09-18T22:43:34Z; 0★). Bounded **Pi** + execution supervisor. README 100% slop badge. **Not a second + coding agent:** Jev interprets evidence; ordinary code controls + freshness, limits, and permitted responses. Skill selection (≤3 + or none), failure recovery (actual completed tool outcomes; exact + repeats in code), review (chunk × rule), optional investigation, + verification (select among configured commands; never generates + them). Default recovery **shadow**. Steering (opt-in) delivers + fixed replan/ask-user guidance, **never generated commands**. + Optional pre-tool feedback is off by default. Distinguish from + `phin-tech/pi-jev-approver` (fail-closed remainder) and + `kevinpita/pi-jev-context` (sieve). Do not copy devenv/bun. + + **Placement.** Agent self-supervision lifecycle + (`agent-self-assessment.md`) as a bounded supervisor, not a + planner. Hole: gate / route. **Empirical** as README policy + (shadow default). **Hypothesis** as improved task completion — + "small fixture experiments are useful smoke tests, not evidence" + (theirs). Cards: `agent-self-assessment.md`; `mappings.md` §9; + `mixed-architecture.md`. + +### MED (brief) + +- **[`Victor-Casado/if-ai`](https://github.com/Victor-Casado/if-ai)** + (MIT; created 2026-09-18T22:09:14Z; 0★). Plain-English PR checks: + one condition, a required `min-confidence`, one GitHub Action. + Modes `pr-body` / `diff` / `per-file`. Pass only when true **and** + confidence ≥ threshold. Empty body/diff, timeout, API error → + **fail** (fail-closed on the Action). Two-option Choice (Noul has + no native confidence). PR text is data, not a security boundary. + Cousin of Abide / jev-pref / jev-marshal. Do not copy the Action + YAML. +- **[`alexsatch/omp-auto-mode`](https://github.com/alexsatch/omp-auto-mode)** + (MIT; created 2026-09-18T22:05:41Z; 0★). oh-my-pi plugin. README + is one line: classify tool calls `safe` / `unsafe` / `ask`. + Pre-action gate cousin. Description-thin. +- **[`jolehuit/jev-downloads-sorter`](https://github.com/jolehuit/jev-downloads-sorter)** + (MIT; created 2026-09-18T22:25:30Z; 0★). Device-loop Choice: + launchd `WatchPaths` on `~/Downloads`; one decision per file; + ~400 ms README examples. Closed folder catalog; never invents + folders; never overwrites; OpenRouter unreachable → extension-map + fallback. Do not copy `install.sh`. +- **[`LakshyaChaudhry/jev-label-desk`](https://github.com/LakshyaChaudhry/jev-label-desk)** + (created 2026-09-18T22:49:49Z; 0★). Description: weekend project + using Jev to automate trace/data labeling for a provided taxonomy. + README empty this pass. **Watch / description-only.** +- **[`flaviomartil/herdr-jev`](https://github.com/flaviomartil/herdr-jev)** + (created 2026-09-18T22:23:53Z; 0★; license not stated this pass). + Jev triage ~260 ms + triad orchestration (advisor / implementer / + reviewer). No key → local heuristic, 0 ms. Do not copy + `install.sh` or the model matrix. +- **jev-gateway siblings.** Product + [`vinilana/jev-gateway`](https://github.com/vinilana/jev-gateway) + (fail-open passthrough) + bench above. Claude Code is `hint` mode + (cannot force `tool_choice` with thinking / cache). Do not copy + ports. + +### Omni / Jev-omni + +Not multimodal. Usage: merge-gate / wake / claim-evidence as +text-state decisions **before** a generative turn. Archive + +landscape pointer. Archer still Watch. + +## 52. GLiNER2 Ultrafast — encoder backend of observe→score-among-candidates→code-acts (2026-09-18 ~16:56 Boise) + +America/Boise ~16:56 = 22:56 UTC. Docs-only fold into open PR #2 +(`cursor/augustus-store-envelope-00b4`). Not a competing PR. Archer +27B drop still **WATCH**. Identity lock vs `typesafe-ai` / `tenbin` / +`decision-first` holds. No wrapper, no `uv` / `.env` / Browser Harness +doctor how-to, no copied ports or Hub download scripts. No invented +metrics. **Not Jev. Not GLiNER2.5. Not multimodal.** Do not re-fold +§48 solari-reflex as a new product, §50 compaction as a species +rewrite, §51 CI merge-gate / wake / Harbor on/off, GLiGuard as a +safety clone, or jev-ultrafast's 7.1 s Flights +row as this demo. + +Backend-agnostic: this is an **observe → score among observed +candidates → code acts** placement (pointer/selection among a11y/DOM +controls; hybrid local decide + remote fill; `DONE` is loop +termination, not verified success). TypeSafe Jev is one backend for +that *job* (`browser-use/jev-ultrafast`, `hitakshiA/solari-reflex`); +local Fastino GLiNER2 is another. Same instinct as §50 compaction +(Jev Noul/Score vs GLiNER2.5 encoder): Augustus stays family-first. + +### HIGH + +1. **[`sahibzada-allahyar/gliner2-ultrafast`](https://github.com/sahibzada-allahyar/gliner2-ultrafast)** + (MIT; Python; created 2026-09-18T19:35:23Z; 12★ at capture). Fork / + adaptation of [`browser-use/jev-ultrafast`](https://github.com/browser-use/jev-ultrafast) + that swaps TypeSafe Jev for **local Fastino GLiNER2** + ([`fastino/gliner2-multi-v1`](https://huggingface.co/fastino/gliner2-multi-v1); + Apache-2.0 weights, 307M extractor, GLiNER2 not GLiNER2.5). Real + browser automation. **Work in progress.** **Not a prose planner. + Not a screenshot VLM. Not a hosted decision-model API.** + + Load-bearing loop (README + `docs/architecture.md`): + + ```text + goal → local GLiNER2 → requirements + page → observed controls → local matching + controller → browser action + → text helper (API) if typing needed + ``` + + GLiNER extracts requirement spans and **scores observed controls**. + Candidates come from DOM / accessible labels (`snapshot.js`); the + model chooses among them. It does **not** use screenshots. It does + **not** generate selectors or executable JavaScript. Code owns + requirement order, progress, calendar matching (English month + names, ISO, US-style numeric), form submit, freshness / visibility + / disabled / occlusion checks, and execution. Browser mutations + are not automatically retried after uncertain execution. The + inspector's scores are **not calibrated probabilities of task + success**. + + Default hybrid, not fully offline: local GLiNER2 for decide; + Mercury 2.5 via OpenRouter for typed field text (OpenAI-compatible + endpoint configurable). Websites and the text-model service need + the network. The browser uses an owned tab in an existing Chrome + profile. + + **`DONE` ≠ verified success.** Termination is heuristic. Loop + `DONE` reports that the agent stopped; applications must + independently inspect the actual result. Example verifiers run + *after* the loop and do not choose actions. Same Harbor-style + honesty as solari-reflex (`notes.md` §48). + + **Demo claims (theirs; not re-run; not a bake-off).** Live Google + Flights, one-way NYC→SFO on 9 Oct 2026; no ticket selected or + purchased. README / `docs/demo.md`: **12.20 s** to visible results + (12.201 s frame); complete action loop **13.785 s**; API usage + **~$0.0001** ($0.00010623 across three text-helper calls). Clock + starts after model loading, goal parsing, initial navigation, and + first page observation. Independent outcome verification is + *outside* that clock. Local compute and electricity excluded. + Their measurement doc: a demonstration, not a controlled + performance comparison or general reliability benchmark. Do not + invent a vs-Jev-Ultrafast table; do not merge with the atlas + 7.1 s / 9.5 s Flights figures (`notes.md` §49). + + **Four load-bearing mental models (architecture, not a plugin + catalog):** + + 1. **Same job, encoder backend.** Jev Ultrafast scores among + observed controls with TypeSafe Jev; this repo scores among + observed controls with local GLiNER2. The *hole* is + observe → score-among-candidates → code acts. Compaction + (`notes.md` §50) already taught that keep/drop is + backend-agnostic (Jev Noul/Score vs GLiNER2.5). Computer-use + selection is the same lesson on a different hole. Augustus + does not pick a vendor for the hole. + + 2. **Candidates from observation, not generation; not + screenshot multimodal.** Control set is a11y/DOM-derived. + Pointer/selection species with + [solari-reflex](https://github.com/hitakshiA/solari-reflex) + (structured obs, Jev, no screenshots; `notes.md` §48) and + [laya-mind2web-browser-agent](https://huggingface.co/ShaunSpark/laya-mind2web-browser-agent) + (Laya open head; operation + target index over a list of + interactive DOM elements; author-reported 74.3% on 68 + held-out — small n, not a ranking; Apache-2.0). Contrast + [blackwood-rlcd](https://huggingface.co/BlackwoodAI/blackwood-rlcd) + (screenshot + letters code already marked → Choice; CC BY-NC; + `notes.md` §46). Pixel-free / DOM-as-text is the preferred + computer-use placement when the environment is already + structured (`judgment-class.md` vision pattern 1; + `mental-models.md` §boundary). + + 3. **Hybrid local decide + remote fill.** Economics: tiny API + for TYPE text; local encoder for control scoring. Mixed + architecture, not dual-process-ai (that product is S1/S2 with + a confidence τ; routing accuracy unmeasured — `notes.md` + §49). Here the *split* is where compute lives: judgment on + the laptop, generation only where a string must be typed. + Code still owns actuators. A loop with nothing to type stays + local. + + 4. **Composition + independent outcome check.** Perception + (Browser Harness / DOM snapshot) → decision (GLiNER2) → + verified act (code). The post-run verifier is not the + decision policy. Same Harbor instinct as solari / jev-e2e: + score the *task*, not a self-report. + + **Limits (theirs, README + architecture; experimental).** Common + HTML and ARIA patterns; behavior varies with page structure and + goal wording. Date parsing is English-oriented. Not fully + offline. Optional traces can include page content. No published + Harbor taskset or calibration of control scores as P(success) — + do not invent them. + + **Siblings — complementary, do not merge.** + + - **`browser-use/jev-ultrafast`:** same job, Jev backend. Credit + in their README. Do not treat this demo clock as a head-to-head. + - **`hitakshiA/solari-reflex`:** same observe → decide → verified + act; Jev; Harbor-style table vs Codex on Solari (`notes.md` §48). + - **`m-newhauser/gliner25-compaction`:** Fastino encoder sibling + *class*, GLiNER2.5 (`fastino/gliner2.5-base-v1`), compaction / + keep-drop hole — not this checkpoint and not browser CU + (`notes.md` §50). + - **GLiGuard:** Fastino encoder sibling (safety-schema classify). + Not control scoring. + - **blackwood-rlcd:** screenshot multimodal decide. Different + input. Not this placement. + - **`24601/rh-guard`:** light note only — independent outcome + verification vs trusting `DONE`. Not reward-hack detection. + + **Placement.** Exact-text keep/drop among observed controls + (`applied-mappings.md` §2) + mixed architecture (code acts; LLM + writes TYPE only). Pillar: search/control (one substituted + classifier step) + runtime-assurance sandwich (freshness / + visibility / no generated selectors in code). Hole: perceive / + keep-drop / replace-one-classifier-step. Family: GLi\* encoder + (GLiNER2 local, `gliner2-multi-v1`). Fail-closed on actuation + (code validates the observed node). Eval path: their Flights + demonstration + post-run verifier; no class bake-off this pass. + **Empirical** as README / architecture behavior. **Hypothesis** + that the same envelope transfers to *your* sites. Cards: + `judgment-class.md` (primary); `mixed-architecture.md` (primary); + `applied-mappings.md` §2; `mappings.md` §9 / §12; `faq.md`; + `mental-models.md`; `validation.md`; `methods-catalog.md`; + `toolbox-mapping.md`; `agent-self-assessment.md`. No wrapper. + +### Omni / Jev-omni + +Not multimodal pixels. Strong **composition / open-weights decide** +exemplar for browser computer-use (local encoder + remote fill). +Archive + landscape. Harbor-style independent verify is already in +their framing. + +## 53. jev-pruner — evidence-preserving Bash stdout prune (2026-09-18 ~17:15 Boise) + +America/Boise ~17:15 = 23:15 UTC. Docs-only fold into open PR #2 +(`cursor/augustus-store-envelope-00b4`). Not a competing PR. Archer +27B drop still **WATCH**. Identity lock vs `typesafe-ai` / `tenbin` / +`decision-first` holds. No wrapper, no marketplace / Codex install +how-to, no copied `keepThreshold` as a class constant. No invented +metrics. **Not a summarizer. Not session compaction. Not GLiNER.** +Do not re-fold §50 gliner25-compaction as this product, §52 +observe→score-act, §48 jevprune/winnow as a new species, or +fast-jev-compaction as a duplicate. + +Family: **evidence-preserving reduce** (pointer/extractive, dropped +bytes recoverable). Same instinct as +[`tamaratran/fast-jev-compaction`](https://github.com/tamaratran/fast-jev-compaction) +and [`m-newhauser/gliner25-compaction`](https://github.com/m-newhauser/gliner25-compaction) +(`notes.md` §50). **Different job:** prune a just-run Bash result +*before* the main LLM sees it, not compact completed tool pairs +already in session history. **Different backend vs GLiNER2.5:** +TypeSafe Jev Noul, not an encoder retention Choice. Same author as +fast-jev-compaction; the READMEs say the two are independent and can +be installed together. Marketplace / Claude plugin id is still +`fast-jev-output` — name ≠ id. + +### HIGH + +1. **[`tamaratran/jev-pruner`](https://github.com/tamaratran/jev-pruner)** + (MIT; TypeScript; created 2026-09-18T03:00:58Z; 5★ at attached + capture, 7★ live this pass). Claude Code plugin: after Bash runs, + **before** the result is sent back to the main LLM, Jev + Noul-prunes stdout **without generating a summary**. Codex is an + **opt-in wrapper/skill**, not automatic `PostToolUse` + interception — the host cannot replace native shell output that + way. **Work in progress / measurement-native.** **Not a prose + compressor. Not a screenshot VLM. Not a GLiNER backend.** + + Load-bearing loop (README): + + ```text + Claude requests Bash → command runs → Jev prunes stdout → Claude receives result + (+ archive path for dropped spans) + ``` + + One Noul per chunk: “does any line in this chunk need to remain + available?” A single needed line protects the chunk. Chunks of + `chunkLines` lines (default 20), capped at 200 chunks. Categories + (build/test, search/excerpt) add guidance only — they never mark + a whole command disposable or change the keep threshold. + + **Four load-bearing mental models (architecture, not a plugin + catalog):** + + 1. **Evidence-preserving prune, not summarize.** Same extractive + honesty as gliner25-compaction / jev-reviewer / testimonial-miner: + the model **selects**; code **keeps verbatim chunks**. Dropped + text is archived under `.claude/fast-jev-output/` (Claude) or + `.jev-pruner/` (Codex) *before* the first scoring request; + markers point at the recovery path. A generator summary of + stdout is a different species. Credential-like commands/output + are **not** archived; markers tell the agent to re-run. That + check only skips local archive — it does **not** redact + secrets from Jev (`notes.md` this section). + + 2. **Hard envelope, then soft Noul.** Code proves pass-through + *before* Jev runs: ≤10,000 estimated tokens (`estimateTokens` + on raw stdout; `minTokens` can raise this, not lower it); + errors; JSON/XML/YAML/diff/binary; whole-document commands + (`cat`, `jq`, `git diff`, `git show`, `base64`, `openssl`). + Format detection beats a build/search category. Jev only + scores the residual noisy log. Same sandwich as + bitrate-advisor / jevgate / gliner25 mutation monitor + (`mappings.md` §12 / §18): structure first, remainder judged. + + 3. **Fail-safe keep original.** Archive write failure, Jev + failure, unfit state, incomplete scoring against every history + segment, first/last chunks, error/warning patterns → original + stdout untouched. The *reduction* is the irreversible act, so + uncertainty fails closed to keep-full — same polarity as + gliner25 `keep_full`, opposite slogan from Abide / jevgate + fail-open. Plugin-eval receipt of that envelope: Harbor's + `CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC` refuses the Jev + fetch, the hook falls back to original output, and the suite + **cannot exercise pruning** (evals README). That is the + fail-safe working, not a missing metric. + + 4. **Host capability shapes the product.** Claude: automatic + `tool.call` wrap of Bash `next()`. Codex CLI 0.152.1 cannot + replace native shell output from `PostToolUse`, so the product + is a wrapper + skill. Same judgment engine; different + insertion. Do not copy the wrapper. + + **Eval (theirs, not re-run; 2026-09-18).** Manual `trimOutput` + sweep, 3 runs/scenario, `jev-latest`: needles kept **24/24**; + mean reduction **83% (71–92%)** on scenarios meant to trim; + wrongly trimmed **0/12** pass-through; mean latency **240 ms**. + Wider sweeps: standard 8/8 / 83%; accuracy 36/36 / 87%; real + captures 10/10 / 54% (three correctly left whole); needle matrix + 9/9. README demonstration: a 76,379-char log whose 2,227-char + preview missed the error line became 4,013 chars of pruned + output that kept it — illustration, not a bake-off. Harbor + Terminal-Bench 2.0 adapter + six-run paired pilot is + **integration, not a full benchmark or significance test** (full + set is 89 tasks / 178 trials; no published full-run scores this + pass). Plugin `claude plugin eval` cases predate the 10k gate + and cannot reach Jev. Do not merge those tables. `keepThreshold` + default 0.5 is **their** knob. + + **Siblings — complementary, do not merge.** + + - **`tamaratran/fast-jev-compaction`:** same author; session + compaction of completed tool pairs (two Nouls). Different job. + - **`m-newhauser/gliner25-compaction`:** same evidence-preserving + *family*; GLiNER2.5 encoder backend; session compaction; + `shadowMode` default true (`notes.md` §50). + - **`ibrahemid/jevprune` / winnow:** per-line / per-block + relevance before context. Same *sieve* hole; this product adds + the 10k/format envelope + archive + Harbor harness. + - **`24601/rh-guard`:** light note only (fail-safe / envelope). + Not reward-hack detection. + + **Placement.** Context sieve + exact-text keep/drop + (`applied-mappings.md` §1–§2) + mixed architecture (code owns + envelope and archive; Jev scores residual chunks; LLM never + writes the kept bytes). Pillar: selective classification / SDT + (false drop >> false keep) + runtime-assurance sandwich. Hole: + sieve / keep-drop. Family: TypeSafe Jev (Noul). Fail-closed on + the reduction. Eval path: their manual sweep + in-repo Harbor + adapter (pilot ≠ full bench). **Empirical** as README / evals + README behavior. **Hypothesis** that the envelope transfers to + *your* command mix. Cards: `applied-mappings.md` §1 (primary); + `mixed-architecture.md`; `mappings.md` §12 / §18; `faq.md`; + `validation.md`; `judgment-class.md`; `mental-models.md`; + `methods-catalog.md`; `toolbox-mapping.md`; + `agent-self-assessment.md`. No wrapper. + +### Omni / Jev-omni + +Not multimodal. Usage: command-output sieve as a perception/memory +hygiene stage **before** the generative turn. Harbor-adjacent +eval harness in-repo. Archive + landscape pointer. + +## 54. Cua-S1 — specialist System One computer-use (form-v0 profile, source-only) (2026-09-18 ~17:21 Boise) + +America/Boise ~17:21 = 23:21 UTC. Docs-only fold into open PR #2 +(`cursor/augustus-store-envelope-00b4`). Not a competing PR. Archer +27B drop still **WATCH**. Identity lock vs `typesafe-ai` / `tenbin` / +`decision-first` holds. No wrapper, no `uv` / MCP / Driver how-to, no +copied factory env vars. No invented metrics. **Not TypeSafe Jev. +Not GLiNER. Not a general CUA. Not multimodal pixels-in. Source-only +— no weights, no checkpoint scores.** Do not re-fold §52 +gliner2-ultrafast as this product, §48 solari-reflex, §53 +jev-pruner, blackwood-rlcd as a screenshot cousin, or laya-mind2web +as a Laya DOM-index cousin. + +Family: **observe → score-among-candidates → code acts** (selection +head over observed elements; code owns execution order). Same hole +as jev-ultrafast / solari-reflex (Jev), gliner2-ultrafast (GLiNER2), +laya-mind2web (Laya). **Different naming:** CUA's "System One" is a +parallel research label, not a TypeSafe Jev contract. **Different +scope:** specialist form-oriented checkpoint profile +(`cua-s1-form-v0`), not a general computer-use agent. Weights +**Watch**. + +### HIGH + +1. **[`trycua/cua` `libs/cua-s1`](https://github.com/trycua/cua/tree/main/libs/cua-s1)** + (parent MIT; Python package `cua-s1` / `cua_s1`; ~23.3k★ parent + this pass). Research project for **small, specialist computer-use + models** with a defined task class. First checkpoint profile: + `cua-s1-form-v0` (form-oriented UI). This component ships model, + synth data, training, eval utilities, and optional Cua Driver + + MCP server **source**. It does **not** include or download + weights, datasets, demo binaries, or recordings. **No checkpoint + performance claim.** Future official weights may use separate + terms. **Not a TypeSafe Jev drop-in. Not a screenshot VLM. Not a + generator of selectors or field values.** + + Load-bearing loop (README + MODEL_CARD): + + ```text + snapshot / a11y elements + Label:value entities + → tinyx byte encoder + option-attention + → per element: fill | check | click | skip + → code orders execution (plan ≠ execute; dry-run default) + ``` + + Reference `tinyx`: byte-level transformer encoder + + **option-attention classification head**. Per observed interface + element, one option from a **fixed set**: fill with an entity + extracted from the source document, check, click, or skip. The + prototype scores elements independently. Document parser only + extracts `Label: value` pairs. **Code** turns selected options + into an execution order. Fill values are **selected**, not + generated. + + **Four load-bearing mental models (architecture, not a Driver + how-to):** + + 1. **Specialist S1 vs general agent.** Membership in the Cua-S1 + family does not imply general computer-use capability. A + checkpoint has a narrow task contract and checkpoint-specific + eval. Same philosophy as "Jev-class for a job," not omnimodal + AGI. Do not treat `form-v0` as evidence outside its evaluated + boundaries — and there is **no evaluated checkpoint** in this + source-only drop. + + 2. **Choice among observed elements / fixed actions.** Selection + head, not a generator of selectors or values. Family with + gliner2-ultrafast (score observed controls), solari-reflex + (structured observe → typed act), jev-ultrafast, laya-mind2web + (DOM indices). Contrast blackwood-rlcd (screenshot + marked + letters). If the fill entity is not already a `Label: value` + pair the parser holds, this card does not apply. + + 3. **Plan ≠ execute; dry-run default; fail-closed.** Planning and + execution are separate. Optional runtime defaults to dry run. + One unambiguous target window; snapshot-bound element tokens; + reobserve after each mutation. `execute` and `submit` are + independent opt-ins. Submit is narrow: at most one + high-confidence Button / AXButton whose normalized label is + exactly `Submit` or `Submit Form`. Fail-closed on missing + checkbox role/checked state; already-checked boxes skipped; + checked postcondition verified. PDF confined to allowed roots + (cwd default; production should use a dedicated directory). + Portable Cua Driver contract does not currently expose + `set_value` — fill **execution** fails closed unless the + connected runtime advertises token-based value mutation; + planning remains available. Inspect the dry-run plan before + enabling both execution flags. + + 4. **Not TypeSafe Jev.** Parallel "System One" naming in + computer-use research. No Choice/Score/Noul contract, no + `/v1/systemone` drop-in. Augustus stays family-first: the + *hole* is specialist decide among observed candidates under a + hard envelope. Backend-agnostic judgment class still applies. + + **Eval honesty (theirs; no scores this pass).** Included tests + exercise **implementation behavior, not checkpoint quality**. + Offline utilities *report* abstention, coverage, selective + accuracy, wrong actions, wrong targets, and unsafe actions when + the expected behavior was to abstain. Synthetic splits are + disjoint by form signature; model selection uses validation + rather than test. A future checkpoint **must** report exact + revisions, task set, environment, action space, independent + outcome verification, and failure categories. Responsible-use + text: do not treat model output or apparent task completion as + proof the action was correct. **Watch** for a `cua-s1-form-v0` + artifact drop. Do not invent metrics. + + **Siblings — complementary, do not merge.** + + - **gliner2-ultrafast / solari-reflex / jev-ultrafast / + laya-mind2web:** same observe→act *job*; Jev, GLiNER2, or Laya + backends with shipped loops. This is a source-only specialist + head. + - **blackwood-rlcd:** screenshot multimodal decide. Different + input. + - **`24601/rh-guard`:** light note only (dry-run / submit opt-in / + fail-closed state checks). Not reward-hack detection. + + **Placement.** Exact-text keep/drop among observed elements + (`applied-mappings.md` §2) + mixed architecture (code owns + envelope, dry-run, submit gate; model selects among candidates). + Pillar: search/control (one substituted classifier step) + + runtime-assurance sandwich (plan≠execute, fail-closed checkbox / + fill). Hole: perceive / keep-drop / replace-one-classifier-step. + Family: specialist encoder decide head (`tinyx` option-attention) + — **not** TypeSafe Jev, **not** GLiNER. Fail-closed on actuation. + Eval path: none published (source-only); metric *names* are + specified. **Empirical** as README / MODEL_CARD behavior. + **Hypothesis** that a future `form-v0` checkpoint fills the + profile. Cards: `judgment-class.md` (primary); + `mixed-architecture.md` (primary); `applied-mappings.md` §2; + `mappings.md` §9 / §12; `faq.md`; `mental-models.md`; + `validation.md`; `methods-catalog.md`; `toolbox-mapping.md`; + `agent-self-assessment.md`. No wrapper. + +### Omni / Jev-omni + +Not multimodal pixels. Strong **computer-use composition** signal: +perception (a11y/snapshots) → specialist decide → verified act. +Form specialist, not pixels-in. Weights TBD — Watch for +`cua-s1-form-v0`. Archive + landscape pointer. + +## 55. CUDA replica, decision-native RAG, verbatim recall, Ruby primitive, FHIR Harbor, AMBIGUOUS baselines (2026-09-18 ~17:48 Boise) + +America/Boise ~17:48 = 23:48 UTC. Docs-only fold into open PR #2 +(`cursor/augustus-store-envelope-00b4`). Not a competing PR. Archer +27B drop still **WATCH**. X MCP namespace flap continues; `since_id` +not advanced this pass. Attached archive path +`/workspace/jev-archive/2026-09-18/234740` is **not present locally** — +receipts are live GitHub/HF + the published report site. Identity lock +vs `typesafe-ai` / `tenbin` / `decision-first` holds. No wrapper, no +Windows CUDA/venv / gem / Rails / `uv` / `mcp add` / pip how-to, no +copied ports, thresholds as class constants, or invented metrics. + +Do **not** re-fold §50 GLiNER2.5 compaction, §51 latch/wakegate/ +s1-indexer/clear-head/foreman/jevons, §52 gliner2-ultrafast, §53 +jev-pruner, or §54 Cua-S1. Do not rewrite the student-b *species* +(already §33 / applied-mappings §1 / judgment-class LoRA table). + +Backend-agnostic: this hour is **local inference replica**, +**retrieve-wide → decide → evidence set**, **verbatim ledger + +scored recall**, **judgment as a language primitive**, **Harbor- +shaped healthcare measurement**, and **honest pre-registered +AMBIGUOUS eval** (cascade sign-flip; calibration theater). TypeSafe +Jev is the documented exemplar, not the monopoly. Augustus stays +family-first. + +### HIGH + +1. **[`Mintzs/jevify`](https://github.com/Mintzs/jevify)** + (Python; created 2026-09-18T23:41:21Z; 0★ at capture; **no LICENSE + file this pass — do not invent**). Experimental CUDA/PyTorch engine + for parallel classification, yes/no, and rubric scoring with a + shared context. Default model `Qwen/Qwen2.5-1.5B-Instruct`. Python + package `ora_decision_engine`; CLI `ora-decision`. README: **this + repository is independent of the Distillation project.** Default + `workflow.json` is a four-question **refund rubric, not a validated + policy**. Default `--answer-encoding letters`; `--answer-encoding + labels` is the literal-label comparison path. Single-token answers + keep selected-head scoring; explicit label encoding evaluates + multi-token labels with a cached prompt and batched known + continuations. **These are uncalibrated model likelihoods, not + measured correctness probabilities.** Optimizations named in + README: bounded CUDA graph replay, reused answer-head weights, + optional Triton RMSNorm/SwiGLU/RoPE, short-branch kernel, grouped + similar question lengths. Historical tensors live under ignored + `outputs/`; fresh checkout skips those integration tests. + + **Four load-bearing mental models:** + + 1. **Open local packed-logprob / CUDA replica class.** Same *job* + as [`ikermoel/open-alternative-jev`](https://github.com/ikermoel/open-alternative-jev) + (packed one-forward on an open LLM; not a Jev reproduction; + `notes.md` §49): shared context, score allowed continuations, + no generated prose. Family is constrained-AR / logprob surface, + **not** trained decision-only (Laya / kev / Archer Watch) and + **not** a teacher-copy LoRA (openjev-lm / student-b). + 2. **Uncalibrated likelihood ≠ Noul.** Softmax over A/B/C or + `false`/`true` is not a proper-scoring head. Do not threshold + it as P(permit) or as calibrated abstention. Temperature / + ECE on *your* labels if you use it. + 3. **Default refund workflow is not a policy.** Configure + questions separately from inputs; do not ship their example + as production refund logic. + 4. **Do not copy the Windows CUDA/venv how-to.** + + **Placement.** Judgment-class constrained-AR / packed-logprob + cousin (`judgment-class.md` when-to-use). Pillar: search/control + (one substituted classifier step). Hole: replace-one-classifier- + step / perceive. Family: constrained-AR surface on an open LLM. + Fail polarity: **not established** — likelihoods are not + decision scores. Eval path: historical tensors not in a fresh + clone. **Empirical** as README behavior. **Hypothesis** that + CUDA graphs / branch kernels transfer to *your* GPU. Cards: + `judgment-class.md` (primary); `faq.md`; `mixed-architecture.md`. + No wrapper. + +2. **[`emergency-lee/decision-native-rag-skills`](https://github.com/emergency-lee/decision-native-rag-skills)** + (MIT; HTML+skills; created 2026-09-18T23:29:20Z; 0★). Agent Skills + for migrating, evaluating, and designing RAG around a + **decision-native evidence pipeline** rather than fixed Top-K. + Tagline: retrieve broadly → decide explicitly → build an evidence + set → resolve conflicts → reason only over what matters. + **Provider-agnostic.** Jev / OpenJev are cheap semantic operators, + not a required SDK. Three skills: `rag-migrate` / `rag-evaluate` / + `rag-design`. **No bundled Python harness** — generate the smallest + fit-for-purpose harness inside the target project. Does **not** + claim a universal benchmark. Falsifiable hypothesis: at comparable + answer quality and safety, wide retrieval + explicit evidence + decisions can improve evidence recall and cut irrelevant/redundant + context vs fixed Top-K, inside an acceptable latency/cost envelope. + Default migration *gates* (starting targets, not promises): + required-evidence recall improve-or-hold; delivered precision + improve-or-tolerance; redundancy and unresolved contradiction + reduce; unsupported claims and provenance must not regress; p95 + and cost within SLO or explicit trade-off; live user/task success + before full rollout. Eval stages: offline frozen replay → shadow + → canary → A/B. A team should not ship because an offline LLM + judge prefers it. + + **Four load-bearing mental models (core Augustus RAG):** + + 1. **Embeddings stay candidate generators.** Similarity / rerank + is not relevance, sufficiency, redundancy, conflict, time, or + authority. Stop asking one Top-K cutoff to solve all six. + 2. **Retrieve wide → decide → evidence set → LLM.** Reason only + over kept evidence. Generation is downstream of an explicit + keep set. Same family as classifying-RAG-passages cookbook + + mappings §4, promoted from "rerank the shortlist" to + **evidence-set construction**. + 3. **Skills, not a measured harness.** No bundled Python, no + private corpus, no universal number. The hypothesis is + falsifiable on *your* system. + 4. **Do not copy skill files as a product.** + + **Placement.** Retrieval + bounded semantic reranking + (`mappings.md` §4) + applied moderation/ranking + (`applied-mappings.md` §4) + mixed architecture (code owns + evidence-set / conflict / provenance; model scores candidates). + Pillar: VOI (expand only if the evidence set is insufficient) + + MCDA (relevance / freshness / authority as named features). + Hole: rank / sieve / gather. Family: closed decision API *or* + any cheap semantic operator (backend-agnostic). Fail-open on + drop of a candidate (false drop loses evidence). Eval path: + rag-evaluate four stages; **Hypothesis** until a target-system + test runs. **Empirical** as README architecture. Cards: + `mappings.md` §4 (primary); `applied-mappings.md` §4; + `mixed-architecture.md`; `mental-models.md`; `faq.md`. No + wrapper. + +3. **[`Dharundp6/jev-carryforward`](https://github.com/Dharundp6/jev-carryforward)** + (MIT; TypeScript; created 2026-09-18T23:04:58Z; 1★; npm + `carryforward`). MCP session memory: `record` saves a fact + verbatim the moment it happens; `recall` scores non-rule entries + with Jev against the current task. **Nothing is summarised. + Nothing is deleted.** Kinds: `constraint` / `correction` / + `decision` / `measurement` / `thread`. Constraints and + corrections **always return in full** — Jev never votes on a + rule you set. Decisions / measurements / threads need a `ref`. + Provenance: `measured` / `decided` / `told` / `inferred`. + Thresholds (theirs, exported constants, not class constants): + p ≥ 0.60 full entry; 0.30–0.60 one line; below omit from the + brief (still on disk). Fail-open: no key / no task / scorer down + / rate-limited → **whole list** plus a line saying why. + Asker is swappable (`Asker.ask(state, questions)`). JSONL + append-only at `~/.carryforward/.jsonl`. Nine entries × + three tasks is a **hint, not proof**; no accuracy claim until a + proper test. Tests use a fake scorer; never the network. + + **Four load-bearing mental models:** + + 1. **Verbatim ledger, scored recall.** Pointer-not-generator on + *memory*: the model never rewrites the note. Sorting happens + at read, not write. Family with pi-jev-compaction / + testimonial-miner (select, copy, do not summarize). + 2. **Rules never judged.** Constraints/corrections are always- + keep in code. Same sandwich as jevgate's allowlist *proves* + / remainder judged — here the remainder is "is this still + live for the task?" + 3. **VOI / selective memory.** Pay for a scored brief iff it + beats dumping the whole ledger (fail-open dump is the safe + default). Horvitz: skip the irrelevant, never skip the rule. + 4. **Do not copy `claude mcp add` / SessionStart hooks.** + + **Placement.** Context sieve (`applied-mappings.md` §1) + VOI + (`mappings.md` §6). Pillar: VOI + selective classification. + Hole: sieve / gather. Family: closed decision API (Noul + "still live?"). Fail-open on scoring failure. Eval path: none + published (9×3 hint). **Empirical** as README behavior. + **Hypothesis** that scored recall beats dump-or-summary on + *your* session. Cards: `applied-mappings.md` §1 (primary); + `mappings.md` §6; `mixed-architecture.md`; + `agent-self-assessment.md`. No wrapper. + +4. **[`carldaws/hunch`](https://github.com/carldaws/hunch)** + (MIT; Ruby; created 2026-09-18T23:08:32Z; 0★). Probabilistic + control flow for Ruby: `if` / `case` / `<=>` for facts; Hunch + for judgment calls. English is the configuration. `chance` → + Noul (`almost_certain?` / `likely?` / `probable?` / named + levels); `pick` → Choice; `rate` → Score. Batch `Hunch.decide` + over one `given:`. Rails examples (theirs): validations, inbound + email routing, error triage, job retries, enum coercion, + comment moderation. **`rescue nil` on validations is deliberate + fail-open at save** — fail closed instead where it matters + (spam gate). Stub backend for tests. Same interface ≠ same + guarantees for a future LLM backend (Jev calibrated + typed + + milliseconds; an LLM backend is estimates, slower, dearer). + Cousin of [`southpolesteve/probably`](https://github.com/southpolesteve/probably) + (language whose *loop conditions* are Jev feelings) — Hunch is + a library in Ruby, not a new language. + + **Four load-bearing mental models:** + + 1. **Judgment as a language primitive.** `almost_certain?` / + `pick` / `rate` are control-flow, not a prompt. English-as- + config. Same instinct as probably-lang, one layer down. + 2. **Fail polarity is per action.** Validation fail-open + (`rescue nil`); spam *gate* should fail closed. Name the + act, not the slogan. + 3. **Stub is a backend.** Tests never need the network. + 4. **Do not copy gem / Rails how-to.** + + **Placement.** Decision circuits (`mappings.md` §3) + mixed + architecture. Pillar: EU / selective classification. Hole: + gate / route / replace-one-classifier-step. Family: closed + decision API. Fail polarity per call-site. Eval path: example + app tests against live model (theirs); stub for CI. + **Empirical** as README / example suite. Cards: + `mappings.md` §3 (primary); `mixed-architecture.md`; `faq.md`. + No wrapper. + +5. **[`si618/explore-typesafe-ai`](https://github.com/si618/explore-typesafe-ai)** + (Python; created 2026-09-18T23:48:44Z; 0★; **license not in + GitHub API this pass — do not invent**). FHIR clinical System + One (Jev) + Claude System Two on **100 synthetic Synthea** + patients. Report: + [si618.github.io/explore-typesafe-ai](https://si618.github.io/explore-typesafe-ai). + Three scenarios: NEWS2 huddle (Noul/Score/Choice); discharge + med recon (Choice fan-out, Noul, Score); post-discharge inbox + (Choice/Score/Noul + confidence gate). Labels committed + **before** any Jev run. 60 requests to `jev-1.13.0`. **Not + clinically validated.** Claude wrote reference labels, not + clinicians. 20 cases per scenario — wide uncertainty. + + **Report headline (theirs; not re-run):** NEWS2 alone under- + triaged 10/20; NEWS2 + Jev under-triaged 1/20; new-confusion + Noul 20/20. Discharge: 98% of 143 medication statuses; allergy + check 100%; duplicate/interaction checks weak (multi-hop) and + mostly escalate. Inbox: 7/20 auto-dispatched, all correctly; + every misroute caught by the confidence gate; prompt injection + did not steer routing. 65 of 403 judgments (16%) escalated to + blinded Claude Sonnet 5. Cost/speed: 403 judgments / 60 + requests; p50 329 ms/request; **$0.0038** total. + + **Four load-bearing mental models:** + + 1. **Harbor-shaped healthcare measurement.** Frozen synthetic + cohort, labels first, independent S2 review packet, code + owns NEWS2 / recon / routing. Capability demonstration, + not clinical safety evidence. + 2. **Code stays in charge.** Jev supplies inputs code cannot + compute (note meaning, brand names, new vs baseline + confusion). Thresholds re-policy without a new prompt. + 3. **Multi-hop over a list is still jagged.** Duplicate / + interaction checks escalate — decompose or don't ask. + 4. **Do not copy `uv` how-to.** Not a medical device. + + **Placement.** Validation Harbor (`validation.md`) + mixed + architecture (S1 decide / S2 review / code policy). Pillar: + SDT (under-triage cost >> over-triage) + Leveson + (sensor ≠ constraint). Hole: triage / gate / perceive. + Family: closed decision API. Fail-closed on actuation + (escalate / hold); **not clinically validated**. Eval path: + published report + committed labels. **Empirical** as that + named report. **Hypothesis** that the same split transfers + to real FHIR. Cards: `validation.md` (primary); + `mental-models.md`; `mixed-architecture.md`. No wrapper. + +6. **[`ickma2311/jev-baselines-eval`](https://github.com/ickma2311/jev-baselines-eval)** + (MIT; Python; created 2026-09-18T22:57:35Z; 0★). Pre-registered + independent eval of TypeSafe Jev vs nano-class LLM + (`gpt-5.4-nano`), frontier (`GPT-5.6 Terra`), and a supervised + encoder (`bge-small-en-v1.5` + logistic regression, 10,003 + Banking77 train, 9 ms laptop). Not affiliated; ~$1 API paid by + the author. **Both experiments returned AMBIGUOUS.** Same-day + errata, **three rounds** (calibration language, B0 escalation + numbers, encoder-vs-frontier arithmetic, missing cross-fit + accuracies, **threshold-margin sensitivity that flips the sign + of the headline cascade**, wrong parity explanation, latency + framing). Reviews in `reviews/` (GPT-6 Astra via Codex CLI); + author verified every quantitative finding from `results/`. + + **Numbers (theirs; recomputable from published JSONL):** + + - CLINC150 zero-shot n=200: Jev **0.870** vs nano **0.795** + (paired +7.5pp, 95% CI [+3.0, +12.5]) vs Terra **0.915**. + - Banking77 paired n=208: encoder **0.933** [0.899, 0.966] / + **9 ms** wins; vs Jev 0.832, paired encoder **+10.1pp + [+5.3, +15.4]**; encoder vs Terra **+5.8pp [+2.4, +9.6]**. + B0's own pre-registered verdict was also AMBIGUOUS. + - Cascade (B1 primary): at A_Terra − **1pp** (0.905), R_jev + **0.220** vs R_nano 0.485, Δ **+0.265**, CI [−0.530, +0.595] + → **AMBIGUOUS**. At **exact parity** R_jev **1.000** vs + nano 0.730 (Δ **−0.270**) — **sign flips**. Mechanism: Jev + confidence **exactly 1.0 on 102/200 items, 6 of which are + wrong** (only 1 of those 6 is one Terra gets right). No + threshold that *keeps any Jev answer* reaches parity + (t=1.0 → 0.490 escalation at 0.910; t=1.01 escalates + everything). Read the 1pp row as "cheap to get *close*", + never "at equal accuracy". + - Error-ranking AUROC: CLINC150 Jev 0.734 vs nano 0.816 + (paired CI includes zero); Banking77 reverse. **Neither + direction established.** This is *error ranking*, **not + ECE**. No ECE/reliability diagram in this report. + - Latency: recorded median call duration **~2.2×** shorter + for the Jev *configuration* than nano on the same 30 items + (0.42 s vs 0.92 s) — **not** the vendor 40–200×, **not** + isolated model inference speed, **serving-path not + model-speed**. Throughput under Vercel free-tier rate + limit is a different number (200-item run ~3.5 h). Same- + gateway control named and **not run**. + - Deviations disclosed: B0 n 300→208 (Terra 0.875 retained + vs 0.804 omitted); encoder added after B0 pre-reg; B1 run + despite B0's "ambiguous would not expand" stopping rule. + + **Four load-bearing mental models (jevals / Harbor practice + exemplar this hour):** + + 1. **Honest negative + pre-registration.** Kill/go printed + by the analysis scripts, including the one that failed. + AMBIGUOUS is a result. + 2. **Calibration theater.** Confidence = 1.0 on 102/200 + including 6 wrong. AUROC is not ECE. Do not say "better + calibrated" from error-ranking. A cascade at 1pp-below- + frontier is not a cascade at equal accuracy. + 3. **Encoder with labels still wins.** 10k labeled Banking77 + → 0.933 / 9 ms / $0. Test that baseline before paying + per call. Different information regime, not a like-for- + like model bake-off. + 4. **Serving-path ≠ model-speed.** 2.2× is two client-and- + service configurations. Do not invent 40–200× from this + repo. Do not copy pip. + + **Placement.** Validation Harbor / jevals (`validation.md` + primary). Pillar: SDT + calibration. Hole: measure / hill- + climb. Family: bake-off, not a product. **Empirical** as that + named report (AMBIGUOUS + errata). Cards: `validation.md`; + `faq.md`; `methods-catalog.md`; `toolbox-mapping.md`. No + wrapper. + +7. **[`SargeDev/jev-gate-student-b`](https://huggingface.co/SargeDev/jev-gate-student-b) + + [`jev-distill-corpus`](https://huggingface.co/datasets/SargeDev/jev-distill-corpus)** + — **light delta only.** HF card unchanged this pass vs §33: + LoRA r=16 α=32 on Qwen2.5-0.5B; P(relevant) from yes/no + logits; held-out n=60 MAE **0.187** / Pearson **0.791** / + agreement **90.0%** vs vanilla 0.536 / −0.067 / 38.3%; + ~59 ms RTX 3060; gate at 0.5; **fail-open on errors**; + teacher-copy, not independent gold; 148,160-row corpus. + Distillation of a System One *memory gate* remains the + species. Do not rewrite the LoRA table. Do not copy the + usage snippet. + +### STRONG MED (brief) + +- **[`fdemir/toolgate`](https://github.com/fdemir/toolgate)** + (MIT; TypeScript; created 2026-09-18T23:21:59Z; 0★). Pre-exec + tool gate: `allow` / `block` / `review` before execution. + Guard error or timeout **stops** (error distinct from a model + decision) — fail-closed on the *execution* act. Jev is a + probabilistic check, **not authorization**; keep permissions, + argument validation, and transaction limits. 72-case synthetic + starter dataset **not independently human-annotated**. Demos + prove execution wiring, not model accuracy. **Product**, not + the ndolinschi *vocabulary* already in + `agent-self-assessment.md` (that family used allow / ask_human + / deny). This repo's labels are allow / block / review. + `onReview` must obtain authenticated human approval, not ask + the agent to approve itself. Do not copy pnpm how-to. + Placement: `mappings.md` §18 + agent pre-action gate. + +- **[`masa-med-ai/typesafe-screening-mcp`](https://github.com/masa-med-ai/typesafe-screening-mcp)** + (MIT; Python; created 2026-09-18T23:45:47Z; 0★). PubMed + title/abstract screening MCP: `include` / `maybe` / `exclude`. + One Jev request per article (match Noul + relevance Score + + criterion Nouls); decision rule in code, sensitivity-first + (unmet inclusion never auto-excludes). Abstracts never enter + the LLM conversation. One real run (theirs): **326 hits ~ + 17 s ~ $0.014**. Thresholds **not calibrated** on labelled + data. Screening aid, not a systematic-review replacement. + Do not send patient/confidential text. Do not copy `uv` / + keychain how-to. + +- **[`laurentfabre/databricks-jev-pdf-lab`](https://github.com/laurentfabre/databricks-jev-pdf-lab)** + (Python; created 2026-09-18T23:47:15Z; 0★; **no OSS license + selected — public visibility is not a license**). Honest + negative: **no quality-equivalent, end-to-end Jev payoff + demonstrated** for Precision-Mode PDF extraction. Compact + metadata requests: 32.48% fewer input tokens but **26/236 + recommendations changed** (not equivalent-policy). Bounded + verifier 3/5 flags / 0/3 false alarms on four correlated + inspected cases — not calibrated acceptance. Selective-parse + rehearsal retained all 236 pages. Typed output is not truth. + Public snapshot cannot independently reproduce historical + accuracy. Do not treat this as a production router. + +- **[`yannip1234/codex-jev`](https://github.com/yannip1234/codex-jev)** + (Apache-2.0 via upstream Codex; Rust/Swift; created + 2026-09-18T23:49:00Z; 0★). Codex extractive compression + family: custom engine + desktop bridge + native macOS client. + Kept passages copied from source; API failure / timeout / + uncertainty / insufficient savings **preserve original**. + Manually sent official-app message reduced **~185 → 44 + estimated tokens**, `COMPACTION_OK` — **integration demo**, + not complete desktop compatibility. **Equal task accuracy and + lower total cost have not been established.** Savings are + estimated, not tokenizer-exact billing. Family with + fast-jev-compaction / jev-pruner / gliner25-compaction + (pointer, not summarizer). Do not copy Xcode/Rust build. + +- **[`kazuhideoki/jev-search`](https://github.com/kazuhideoki/jev-search)** + (Python; created 2026-09-18T23:43:52Z; 0★; **no LICENSE file + this pass**). Recursive semantic **file** search + fzf: + ripgrep enumerate → Jev match probability → fzf select. + **Not** [`superagents-lab/jev-search`](https://github.com/superagents-lab/jev-search) + (federated *web* search; Jev as query-understanding head and + result-ranking tail). File score is max over overlapping + chunks — **not** a calibrated whole-file probability; long + files may be favored. Failures/unevaluated chunks are not + treated as 0%. `--dry-run` needs no key. Do not copy `.env` + how-to. + +### Omni / Jev-omni / Archer + +Still **WATCH**. No Hub weights. X MCP flap; `since_id` not +advanced. This hour does not wait. Local CUDA replica (jevify) +is an open *inference class*, not that drop. Healthcare Harbor +(explore-typesafe-ai) is synthetic FHIR, not AU residency +weights. + +### Cross-links + +Cards: `judgment-class.md` (jevify uncalibrated replica; +student-b light); `applied-mappings.md` §1 (carryforward), +§4 (decision-native RAG); `mappings.md` §3 (hunch), §4 (RAG + +file-search vs web-search), §6 (carryforward VOI), §18 +(toolgate); `mixed-architecture.md` (fail table + gallery); +`validation.md` (jev-baselines-eval AMBIGUOUS + errata; +explore-typesafe-ai; pdf-lab negative); `faq.md`; +`mental-models.md`; `methods-catalog.md`; `toolbox-mapping.md`; +`agent-self-assessment.md`. No wrapper. + +## 56. Classify-first MCP + living applied-mappings atlas (2026-09-19 ~00:38 UTC / ~18:38 Boise) + +America/Boise ~18:38 = 2026-09-19T00:38Z. Docs-only fold into +open PR #2 (`cursor/augustus-store-envelope-00b4`). Not a +competing PR. Archer 27B drop still **WATCH**. Identity lock vs +`typesafe-ai` / `tenbin` / `decision-first` holds. No wrapper, +no MCP/npm how-to, no copied ports or key-file paths. No +invented metrics. Do not re-fold §50–§55. + +Two HIGH **usage / architecture** signals: a portable classify- +first MCP (same retrieve-wide → decide → evidence-set family as +decision-native RAG), and a living applied-mappings *atlas* +(class patterns from a curated showcase — not a 342-title hit +list). TypeSafe Jev is the documented exemplar, not the +monopoly. Augustus stays family-first. + +### HIGH + +1. **[`kbhuw/jev-sift`](https://github.com/kbhuw/jev-sift)** + (JavaScript; created 2026-09-18T00:13:31Z; 10★ this pass; + **no LICENSE file this pass — do not invent**; GitHub + `license` null). Portable agent plugin + stdio MCP: + **classify first, read selectively.** Tagline: let Jev decide + what the agent should look at next. Pass a query and a batch + of file paths, public webpage URLs, or inline text; the tool + loads content, sends it **directly to Jev**, and returns + compact relevance probabilities. The main agent only opens + items worth a closer look. Tool descriptions work too: score + a supplied description without executing the tool or + predicting an unseen result. Direct TypeSafe + `POST /v1/systemone` with `jev-latest`. Key via `JEV_API_KEY` + / `TYPESAFE_API_KEY` or a private key file (README names + `~/.config/jev-sift/api-key`; do not copy the path). Plugin + `0.2.0+codex.20260918200547`; package `0.2.0`; author Kush + Bhuwalka. Checked-in `dist/server.mjs` includes dependencies. + Core `classify` is framework-independent (inject `evaluate` / + `readText` / `readUrl`). Earlier generic prototype used + `CLASSIFY_*` / chat-completions — those settings are gone. + + **README / schema envelope (theirs; not a how-to):** up to + **50** items; **1–8** typed questions (boolean → Noul; + Choice 2–12 options; Score 2–10 levels) *or* a `query` + shorthand that returns `answers.relevant.probability`. + Exactly one of `text` / `path` / `url` per item; unique ids. + Concurrency **1–8**. File and extracted page text capped at + **60,000** JavaScript characters and flagged if truncated. + Web: **2 MB** / **20 s**, public-IP only (including redirect + targets; DNS pinned to the connection), HTTP(S) ports 80/443, + up to **three** redirects, no JavaScript, no login, no + browser cookies, PDFs / private-network / other binaries + unsupported. A successful fetch of login/challenge HTML is + not the intended article. Webpage fetches receive no Jev + credentials. Plugin does not persist source content or + results. Results retain input order and include per-item + errors, resolved source URLs, truncation flags, model id, + summed input-token usage. **Uncertain items should get a + closer look; errors and truncation are not evidence that an + item is irrelevant.** Tests cover mapping, ordering, + failures, cancellation, truncation, file boundaries, public- + URL validation, redirects, HTML extraction, and an isolated + bundled MCP exchange — **mocks, not an accuracy benchmark.** + Live smoke verifies connectivity only. Model quality on *your* + task still needs evaluation. + + **Four load-bearing mental models:** + + 1. **Retrieve-wide → decide → evidence-set (agent I/O).** + Same family as + [decision-native-rag-skills](https://github.com/emergency-lee/decision-native-rag-skills) + (`notes.md` §55): content goes to the judge **without + entering main agent context first** (paths/URLs). Inline + text the agent already read cannot recover that cost. The + main LLM reasons only over items worth a closer look. + Embeddings/file lists stay candidate generators. + 2. **VOI / context economics.** Uncertain → closer look. + Errors/truncation ≠ irrelevant. Pay for a full read iff + the relevance (or typed question) says it might change the + act. Webpage fetch still costs bandwidth — this saves the + *agent's* read, not the download. + 3. **Hard envelope on I/O.** 60k char, 2 MB / 20 s, public-IP + only, no JS/cookies/login, PDFs unsupported. Same sandwich + family as jev-pruner (size/format then Noul) and bitrate- + advisor (soft propose, code clamps). + 4. **Do not copy marketplace / `mcpServers` / key-file how- + to.** Transport tests ≠ accuracy. + + **Cousins, do not merge.** + [typesafe-screening-mcp](https://github.com/masa-med-ai/typesafe-screening-mcp) + — abstracts never enter the LLM conversation; include/maybe/ + exclude in code (`notes.md` §55). + [kazuhideoki/jev-search](https://github.com/kazuhideoki/jev-search) + — recursive *file* search + fzf, not this MCP (`notes.md` + §55). [jev-pruner](https://github.com/tamaratran/jev-pruner) + — prune Bash *after* it ran; this tool screens *before* the + agent reads. [carryforward](https://github.com/Dharundp6/jev-carryforward) + — scored recall over a ledger you already hold. + [jev-routing](https://github.com/nekowasabi/jev-routing) is a + **host adapter, not MCP**. Dual-orchestration topology A + (Jev-as-tool); the LLM still owns the outer loop. + + **Placement.** Context sieve (`applied-mappings.md` §1) + + retrieval (`mappings.md` §4) + mixed architecture (MCP as + topology A). Pillar: VOI + selective classification. Hole: + sieve / rank / gather. Family: closed decision API. Fail-open + on "open this file" (false drop loses evidence); truncation/ + error ≠ irrelevant. Eval path: none published (transport + tests). **Empirical** as README / schema behavior. + **Hypothesis** that classify-first beats dump-into-context on + *your* agent. Cards: `applied-mappings.md` §1 (primary); + `mappings.md` §4 / §6; `mixed-architecture.md`; `faq.md`. + No wrapper. + +2. **[jevable.com](https://jevable.com/)** — "Discover what + people build with Jev." Independent curated showcase (creator + Nikunj / `@nikunj` in site JSON-LD). **Claim (theirs, this + pass):** **342** curated projects (``, + `#result-count` sr-only, board-data `"total":342`). Homepage + JSON-LD `ItemList.numberOfItems` is **36** (first page / + featured). Board-data `"pageSize":36`, `"nextOffset":36`, + `"sort":"curated"`. Categories in the filter: Agents, Browser + extensions, Creative tools, Data & research, Developer tools, + Experiments, Finance, Games, Marketing, Productivity, + Robotics. HTTP 200 this pass (Railway). **No public API + discovered this pass.** Watch: refresh the claimed count and + category list; do not treat 342 as an Augustus census. + + **What it is for Augustus.** Living **applied-mappings + corpus**: how people apply judgment tools. Primary usage / + application atlas — **not** a 342-title hit list, **not** a + multimodal substrate, **not** a model. Prefer **class + patterns**. Maker demos are claims unless already measured in + notes. Cross-link exemplars already folded; do not invent + repos or clocks. + + **Class patterns (load-bearing; not a gallery dump):** + + 1. **Intent columns.** Spreadsheets recalculate numbers, not + meaning. Type a heading ("Urgency"); each row is scored + (~100 ms is **their** demo claim). Same *hole* as + dataframe semantic columns ([jevpandas](https://github.com/yalindogusahin/jevpandas) + / [jevframe](https://github.com/ktaletsk/jevframe), + `notes.md` §46 / §48) and the launch-week + `dabit3/jev-experiments` JUDGE/SCORE/CHOOSE formulas + (`docs/ecosystem.md`). Predictive app launcher (heading / + keystroke → intent, not alias/fuzzy/habit) is the same + pattern on a catalog. Pillar: MCDA. Hole: perceive / + rank. Weights and vetoes stay in code. + 2. **Score-among-observed.** Candidates already on the page + (a11y/DOM, action space, on-screen posts); the model + scores; **code** clicks / filters / removes. Showcase: + Browser Use Ultrafast (Flights **7 s / $0.0039** is the + *same* demo already in notes as ~7.1 s — do not merge + clocks with gliner2-ultrafast 12.20 s); ad blocker + (DOM element → ad/non-ad); Notte (new action space every + step); computer-use "100×" is a **claim**. Already + folded: jev-ultrafast / gliner2-ultrafast / solari-reflex + / cua-s1 / laya-mind2web (`notes.md` §4, §48, §52, §54). + **Your Signal** (Fabio Angela): score posts *already on + screen*, apply rules locally, reversible, BYOK, no + telemetry — same judge-once / re-policy family as Near + Here firehose (`applied-mappings.md` §4). + 3. **VOI gates.** Instant compaction (tamara: score tool + calls, drop irrelevant — **same job** as + [fast-jev-compaction](https://github.com/tamaratran/fast-jev-compaction) + / [jev-pruner](https://github.com/tamaratran/jev-pruner) / + [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction); + pointer, not summarizer). Prompt-difficulty classifier + before send (fast-mode offer) is a **route** gate, cousin + of [routeKit](https://github.com/rajdhakad9826/routeKit) + (`notes.md` §33) — Jev does not pick the LLM; code + offers. Gmail intent search: embeddings pull first, then + judge — decision-native RAG on mail (`notes.md` §55). + [jev-sift](https://github.com/kbhuw/jev-sift) (this + section) is the MCP of the same VOI: classify first. + 4. **Generative UI decide.** json-render + Jev: *your* + components, actions, design system; the model decides; + render is milliseconds. Cousin already folded: + [jev-agentworld-web-simulator](https://github.com/knowlet/jev-agentworld-web-simulator) + — Jev Choice for intent / layout; generator writes + documents; Zod + **deterministic** compiler emits json- + render spec; model cannot add components (`notes.md` + §48). Decision for control, generator for content. + 5. **Robotics text-state (not pixels).** MuJoCo robot-arm: + Jev does not accept images; it gets simplified geometry + and contacts **as text**; two-call split (what to do, + then how to move). MOSS: Jev picks the target; the + robot picks up the litter. Same observe→decide→act + *job* as computer-use, different body. Cousins: + [jev-drone](https://github.com/RomanSlack/jev-drone) + (500/50 Hz code, Jev advisory 2.5 Hz); Doom demo fed + structured JSON, not raw pixels (`notes.md` §1 / §4). + **Flag:** "drawing, one decision at a time" *claims* + pixel-parallel prediction — contrast the MuJoCo honesty + (text-state). Perception-then-judgment vs shared + multimodal (`notes.md` §39); Archer still Watch. + 6. **Draft-gate fail mode: silence as "safer".** Jev sat + between GPT and the user, killing drafts that broke + rules. Then Jev did not answer. The agent treated + **silence as safer** and stopped sending anything. Took a + second agent to unstick. **Your checker needs a + fail-open / heartbeat** when the judge is down — + missing verdict is not a block and is not a pass. Name + the irreversible act: *withholding the draft* is fail- + closed-by-absence. Contrast Abide `<0.5` silence (linter + stays quiet; the *edit proceeds*, `notes.md` §47) and + carryforward dump / wakegate wake-on-error. Stuck- + detector / done-check (`agent-self-assessment.md`) must + not treat no-answer as "hold forever." + + **Other patterns already in notes (confirm, don't invent):** + Higgsfield auto-routing is a **claim** (`notes.md` §44; + routeKit hole). Trading / signals → decisions: + [jev-trader](https://github.com/jarrodwatts/jev-trader). + Cambium first-class provider: keep in code what can be in + code (mixed-architecture slogan, not a new family). SEO + internal-link audit (maker claim: 45.1 s, 586 pages, **584** + links placed, **139** refused because nothing honestly fit, + $0.21) is Choice-with-`other` at corpus scale — wellposed / + kev NOTA (`notes.md` §45–§46); not re-run. Snack MCDA + (maker claim: 3,000 kids' snacks, multiple criteria, 28 s, + $0.11) is mapping §1 at catalog scale. ai-cli (yes/no / + choose / score from the shell) is a language-primitive + cousin of [hunch](https://github.com/carldaws/hunch) + (`notes.md` §55). Manhattan pathfinding / Sudoku playground: + **algorithm stays yours**; do not replace A* or a solver + with a Noul (`mappings.md` §9; ARC-AGI combinatorial ≠ + extractive, `notes.md` §49). + + **Jev-omni.** Archive the site snapshot as a community usage + atlas. Not multimodal substrate. Flag demos that claim + pixels vs text-state (drawing vs MuJoCo). Refresh count / + categories on hourly watch if useful. + + **Placement.** Applied-mappings atlas (this file + + `applied-mappings.md` / `mixed-architecture.md` gallery), + not a new species. **Empirical** as the public showcase + (342 is *their* count; we did not enumerate titles). + Maker clocks stay **claims** unless already a named receipt. + Cards: `applied-mappings.md`; `mappings.md` §1 / §4 / §6 / + §9; `mixed-architecture.md`; `faq.md`; `mental-models.md`; + `agent-self-assessment.md`; `question-design.md`. No + wrapper. No 342-row dump. + +### Omni / Jev-omni / Archer + +Still **WATCH**. No Hub weights. Showcase drawing-pixel claim +is **not** that drop. MuJoCo text-state is the honest robotics +posture until a multimodal decide ships (blackwood-rlcd is +screenshot-in, not arm-in). + +### Cross-links + +Cards: `applied-mappings.md` §1 (jev-sift classify-first), +§2 (score-among-observed atlas), §4 (intent search / Your +Signal); `mappings.md` §1 (intent columns / snack MCDA), §4 +(RAG family), §6 (VOI gates), §9 (robotics text-state; do not +replace A*); `mixed-architecture.md` (topology A MCP; generative +UI decide; draft-gate heartbeat); `faq.md`; `mental-models.md`; +`question-design.md` (SEO 139-refused as `other`); +`agent-self-assessment.md` (silence ≠ safer); `methods-catalog.md`; +`toolbox-mapping.md`. No wrapper. + +## 57. Stagehand experimental Jev stack — harness pick-and-copy (2026-09-19 ~00:48 UTC / ~18:48 Boise 2026-09-18) + +America/Boise ~18:48 = 00:48 UTC 2026-09-19. Docs-only fold into +open PR #2 (`cursor/augustus-store-envelope-00b4`). Not a competing +PR. Archer 27B drop still **WATCH**. Identity lock vs `typesafe-ai` +/ `tenbin` / `decision-first` holds. No wrapper, no SDK/init how-to, +no copied `experimentalJevAct` as a class constant. No invented +metrics — clocks below are **their PR bodies**, not re-run. Do not +re-fold §50–§56 as this product, jev-ultrafast / gliner2-ultrafast / +cua-s1 / solari as a new species, or blackwood-rlcd as the extract +path. + +**Placement.** Major harness **productization** of +observe→score-among-candidates→code-acts (and pointer-not-generator +extract) inside [browserbase/stagehand](https://github.com/browserbase/stagehand) +(MIT; Browserbase Inc.). Same *job* as jev-ultrafast / solari-reflex +(Jev backends), gliner2-ultrafast (GLiNER2), cua-s1 (option-attention, +not TypeSafe Jev), laya-mind2web (Laya DOM indices). This is the +harness, not a demo loop. TypeSafe Jev is the exemplar, not the +monopoly. **Experimental; all five PRs OPEN draft this pass.** +Watch merge of extract + act paths. + +### Stack (all OPEN draft; each PR targets its predecessor) + +Author [@miguelg719](https://github.com/miguelg719). Created +2026-09-17T06:40Z. Opt-in via `experimentalJevAct` / +`STAGEHAND_EXPERIMENTAL_JEV_ACT` (**not** public create config — +cross-language create contract unchanged). Do not copy the flag. + +1. **[#2951](https://github.com/browserbase/stagehand/pull/2951)** + (1/5) — report **editable** element ids alongside the a11y + snapshot (`plaintext` / `richtext` AX `editable`). Side channel + on the private snapshot type: outline text, xpath map, and url + map are **byte-for-byte unchanged**, so nothing the LLM sees (or + any cache key) moves. No consumer in this PR. +2. **[#2952](https://github.com/browserbase/stagehand/pull/2952)** + (2/5) — TypeSafe Jev client + candidate-picking library. No + wiring yet (tree-shaken). `/v1/systemone`; 8 s timeout; per + endpoint+key circuit breaker (auth 60 s, 3 consecutive failures + 30 s); errors carry the code only, never the body or key. Outline + → per-intent views (pointer / input / select / scroll / option / + broad / text / link). Pick: role view first, then named elements; + lists over 40 cut to the 30 sharing words with the instruction; + two questions per request — `best` (no "none") and `strict` (with + "none", vetoes above 0.9); unease holds while the next tier + tries; **ambiguity stops rather than guesses**; twins share a + vote; huge lists sharded by character budget. Args parsed in + code; `%variable%` values redacted before anything leaves the + process. +3. **[#2953](https://github.com/browserbase/stagehand/pull/2953)** + (3/5) — experimental Jev **decision tree for `act()`**. Intent + (no snapshot) → args in code → candidates + pick → existing + `performUnderstudyMethod` → deterministic checks (fill read-back, + native `