From ed9f037a642efe00d65427519de533ea868f8d2a Mon Sep 17 00:00:00 2001 From: Cursor Agent Date: Fri, 18 Sep 2026 19:06:47 +0000 Subject: [PATCH 01/43] Fold 12:58 Boise watch: store-index fork, hard envelopes, harness practice MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Archer drop still Watch — no architecture rewrite. Novel vs the 11:59 digest: sqlite-jev in-engine vs jevql CLI; bitrate-advisor and the JOB planner as soft judgment inside a hard envelope; jev-routing host adapter; jev-claw classify-then-policy. Already-folded HIGH get frames only (memory gate, receipts economics, encoder vs decoder replica). Docs-only; no Jev wrapper. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com> --- .agents/skills/augustus/SKILL.md | 8 +- .../augustus/references/applied-mappings.md | 9 +- .agents/skills/augustus/references/faq.md | 33 ++- .../augustus/references/judgment-class.md | 6 +- .../skills/augustus/references/mappings.md | 49 ++++- .../augustus/references/mental-models.md | 11 +- .../augustus/references/mixed-architecture.md | 34 ++- .../skills/augustus/references/validation.md | 6 +- CHANGELOG.md | 14 ++ README.md | 3 +- docs/ecosystem.md | 13 +- research/archive/findings.md | 33 +++ research/notes.md | 195 ++++++++++++++++++ research/refresh-log.md | 16 ++ research/sources.json | 56 ++++- 15 files changed, 459 insertions(+), 27 deletions(-) diff --git a/.agents/skills/augustus/SKILL.md b/.agents/skills/augustus/SKILL.md index 66487a6..4865cd3 100644 --- a/.agents/skills/augustus/SKILL.md +++ b/.agents/skills/augustus/SKILL.md @@ -1,6 +1,6 @@ --- name: augustus -description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, openjev-lm, Nimble, encoder DeBERTa, LoRA distill), announced open decision-model (Watch), constrained-AR (TypeAR, pcdServer), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"eval path\", \"jevals\", \"Harbor taskset\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge, TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, or \"formally verify with Jev\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." +description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, openjev-lm, Nimble, encoder DeBERTa, LoRA distill), announced open decision-model (Watch), constrained-AR (TypeAR, pcdServer), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"eval path\", \"jevals\", \"Harbor taskset\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge, TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, or \"formally verify with Jev\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." license: MIT metadata: version: 0.3.0 @@ -51,7 +51,9 @@ classical method you already trust, substitute it, classify the win this?", "is Jev the only model?", "is this only for software?", "formally verify with Jev / replace TLA+ / Dafny / DST", "Alloy vs Apalache", "GLiNER vs Jev", "LLM-as-judge", - "paraphrase brittleness", "allowlist then judge", or "TOCTOU-of-Noul": + "paraphrase brittleness", "allowlist then judge", "TOCTOU-of-Noul", + "Jev inside the database / sqlite-jev", or "Jev picks bitrate / join + order / the model": read `references/faq.md`, then `references/mental-models.md`, then `references/mixed-architecture.md`, then @@ -129,6 +131,8 @@ classical method you already trust, substitute it, classify the win | Selective classification / decision theory | Thresholds from action costs, abstention paths | `references/mappings.md#2-probabilistic-judgments--cost-sensitive-decisions` | | Decision tables / circuits / state machines | Judgment predicates, code owns transitions | `references/mappings.md#3-semantic-predicates--decision-circuits` | | Retrieve + expensive relevance fn | Bounded rerank of a retrieved shortlist | `references/mappings.md#4-retrieval--bounded-semantic-reranking` (independent TREC DL2019 benchmark: Jev zero-shot best MAP 0.4748, nDCG@10 0.683 vs tuned monoBERT 0.718 — competitive, not dominant) | +| Store as semantic index (SQL / SQLite / zoxide) | Cheap exact predicates first; typed questions on the remainder. In-engine extension (sqlite-jev) vs CLI rewrite (jevql) vs path index (joxide) | `references/mappings.md#4-retrieval--bounded-semantic-reranking` | +| Soft judgment inside a hard envelope | Model may only match the deterministic policy or be more conservative (bitrate ABR; query-planner override-when-confident) | `references/mappings.md#12-runtime-assurance-sandwich-hypothesis`; `references/mappings.md#18-structural-prove--soft-remainder-hypothesis-as-domain-general-empirical-as-named-shapes` | | Value of information / gather as an act | Pay for another observation only if EV(decision) improves more than cost; abstain from calling *any* model when a regex already answers (meta-VOI) | `references/mappings.md#6-value-of-information--gather-as-an-enumerated-act` (**Hypothesis** until a labeled act/outcome log; 149-row receipt is Empirical as a shape) | | Signal detection / ROC | Criterion and operating point from costs and base rate, not accuracy | `references/mappings.md#7-signal-detection--criterion-not-accuracy` (**Hypothesis** for non-SWE plots) | | Org / safety control structure | Sensor ≠ constraint (Leveson); STPA if the sensor lies | `references/mappings.md#8-control-structure--sensor--constraint-leveson` | diff --git a/.agents/skills/augustus/references/applied-mappings.md b/.agents/skills/augustus/references/applied-mappings.md index cedb1d9..91abe7c 100644 --- a/.agents/skills/augustus/references/applied-mappings.md +++ b/.agents/skills/augustus/references/applied-mappings.md @@ -47,7 +47,9 @@ Local teacher-copy for the same hole: [`SargeDev/jev-gate-student-b`](https://huggingface.co/SargeDev/jev-gate-student-b) (Qwen2.5-0.5B LoRA; P(relevant) from yes/no logits; 148,160-row [`jev-distill-corpus`](https://huggingface.co/datasets/SargeDev/jev-distill-corpus); -card: fail-open on errors). Agreement with Jev labels is not independent +card: fail-open on errors). That is System One as a **memory/context +gate**, not an action permit: vector recall → local yes/no → inject or +stub (`notes.md` §33, §44). Agreement with Jev labels is not independent gold (`notes.md` §33). Official cousin: classifying RAG passages cookbook (**Contract**). **Counterexample**: one Noul "is this log useful?" over 3k lines — that is nine judgments pretending to be one. **Test**: recall @@ -186,6 +188,11 @@ LlamaIndex selectors fail closed or a declared default. Jev estimates task *requirements*; code applies hard constraints and a deterministic cost/quality/latency policy — Jev does not pick the model (**Hypothesis** until measured on *your* catalog; `notes.md` §33). +[`trietphan/jev-claw`](https://github.com/trietphan/jev-claw) is the +same split for OpenClaw (classify axes; `decide()` maps the route; path +regex floors risk). [`nekowasabi/jev-routing`](https://github.com/nekowasabi/jev-routing) +is a host adapter, not an MCP plugin: compact, then one Choice + done, +then one schema (`notes.md` §44). [`TheoOliveira/pi-jev`](https://github.com/TheoOliveira/pi-jev) is the same selector hole inside Pi (tools + skills); fail-open to a keyword shortlist; **not** `kevinpita/pi-jev-context` (sieve). diff --git a/.agents/skills/augustus/references/faq.md b/.agents/skills/augustus/references/faq.md index e2ffc7c..d4a1edc 100644 --- a/.agents/skills/augustus/references/faq.md +++ b/.agents/skills/augustus/references/faq.md @@ -256,8 +256,10 @@ sentences, and that is the point ([Langfuse framing, 2026-09-18](https://x.com/langfuse/status/2100980004678971491)). When you need a paragraph rationale, a trace UI, or an annotation workflow, generation and the eval platform still own those seats. -Verbal LLM scores are uncalibrated. Do not thin this skill into a -Langfuse how-to. Mixed architecture: traces stay; the judge step can +Verbal LLM scores are uncalibrated. The Harbor/jevals-adjacent +practice is: shadow mode + fixtures that assert on the **action**, +not on prose (`jev-harness`, `validation.md`; `notes.md` §44). Do not +thin this skill into a Langfuse how-to. Mixed architecture: traces stay; the judge step can be a System One model. ## Allowlist first, then Jev? @@ -272,6 +274,33 @@ database query already answers, **do not call a model** (`wotai-dev/typesafe-jev-tools`, `notes.md` §42). That is meta-VOI, not a hook tutorial. +## Should Jev live inside the database? + +The *hole* is a semantic index over a structured store: cheap exact +predicates first, typed questions on the remainder. Two serving +choices, same hole (`mappings.md` §4; `notes.md` §44): + +- **In-engine extension** (`sqlite-jev`; cousin `pg-jev`): SQL sees + `jev()` / `jev_rows`. Convenient. The database process now has an + API key, a spend guard, and a residency problem. +- **Out-of-process CLI** (`jevql`): vanilla Postgres never sees + `jev()`. The rewrite layer owns the call. + +Neither is an index. Full-scan the post-filter remainder. Row contents +leave the store. zoxide/`joxide` is the same hole over paths. Do not +copy SQL. + +## Can Jev pick the bitrate, the join order, the model? + +Yes as a **proposal inside a hard envelope**, no as the actuator. +bitrate-advisor: Jev may only match the deterministic cap or be more +conservative; missing the model returns the policy's answer. +mmalisper's JOB planner: Postgres plans first; Jev overrides only when +confident (+12% geomean author-reported; join-order Choice alone was +2× slower). routeKit / jev-claw / Higgsfield: classify requirements; +code picks the generator. The envelope is load-bearing +(`mappings.md` §12, §15, §18). + ## Is confidence a trained score? No — not on Hume's reconstruction, and not as a new contract. Re-read diff --git a/.agents/skills/augustus/references/judgment-class.md b/.agents/skills/augustus/references/judgment-class.md index 567ea88..a03ea2b 100644 --- a/.agents/skills/augustus/references/judgment-class.md +++ b/.agents/skills/augustus/references/judgment-class.md @@ -405,7 +405,11 @@ is the generator, not a sixth surface. **Three open paths** (not three species, not extra when-to-use rows): encoder open-jev (DeBERTa, public gold); AR constrained decode (TypeAR Python/SGLang, pcdServer native GGUF); trained decision-only (Laya / -Nimble / Archer **Watch**). Pick from the hole. A constrained softmax +Nimble / Archer **Watch**). Encoder vs decoder **replicas** of that +third path: DeBERTa is public gold with an OOD drop; openjev-lm / +jev-gate LoRAs are teacher-copies with named receipts (overnight +6-vCPU, $0/call — economics, not a new species; `notes.md` §25, §44). +Pick from the hole. A constrained softmax is still not a Noul. Laya companion packaging this hour: [`laya-typed-decisions`](https://huggingface.co/convaiinnovations/laya-typed-decisions) (same 421.3M; acc 0.766 / Brier 0.066 on `LocalLLaMA/typed-decisions`, diff --git a/.agents/skills/augustus/references/mappings.md b/.agents/skills/augustus/references/mappings.md index a2fbadc..330b5cb 100644 --- a/.agents/skills/augustus/references/mappings.md +++ b/.agents/skills/augustus/references/mappings.md @@ -142,13 +142,18 @@ is not intrinsically wrong (offline, modest corpora) but it is not an index — per-query work still scales with candidates. Low latency ≠ no retrieval. **Store as the index (Empirical as a *shape*, 2026-09-18):** -[`kylemclaren/jevql`](https://github.com/kylemclaren/jevql) judges -schema-conditioned row objects; vanilla Postgres never sees `jev()`. -Cheap SQL first; the remainder is a typed Choice/Noul/Score over rows. -Row contents leave the database (same residency warning as AU health). +Cheap exact predicates first; typed questions on the remainder. Two +forks of the same hole: **in-engine extension** +([`mgaitan/sqlite-jev`](https://github.com/mgaitan/sqlite-jev), loadable +SQLite `jev_rows`; inspired by [`realZachi/pg-jev`](https://github.com/realZachi/pg-jev)) +vs **out-of-process CLI** ([`kylemclaren/jevql`](https://github.com/kylemclaren/jevql) +— vanilla Postgres never sees `jev()`). sqlite-jev is a semantic full +scan, not an index; `max_rows` is a spend guard; thresholds stay in SQL. [`ant4g0nist/joxide`](https://github.com/ant4g0nist/joxide): zoxide owns the directory index; Jev scores a shortlist; destinations are existing -local paths only; fail-open. `notes.md` §42. +local paths only; fail-open. Row contents leave the store (same +residency warning as AU health). Do not copy SQL, env, or CLI flags. +`notes.md` §42, §44. ## 5. Hierarchy → bounded heuristic search @@ -368,6 +373,17 @@ immediate win missed once reversed; Fool's-mate confidence 31%/37% so a hole on the constrained-AR surface. Not a strength rating. `notes.md` §42; `validation.md`. +**Query planner as the envelope (author-reported, 2026-09-18):** +[@mmalisper](https://x.com/mmalisper/status/2101001041903009987) on the +Join Order Benchmark. Jev picking join order was **2× slower**. +Cardinality estimates helped when outside context informed the plan; +when Jev was wrong, one query was ~10× slower. Hybrid: Postgres plans +first; Jev overrides **only when confident** → **+12% geomean**, no +dramatic slowdowns. A Jev call is 100s of ms, not yet practical on +every plan. The planner is the hard envelope; confidence is the gate; +fail-open to Postgres. **Hypothesis** until reproduced on *your* +workload. `notes.md` §44. + ## 10. Spec property pipeline (Hypothesis) **Method**: NL/ADR → candidate properties → human strengthens → @@ -424,9 +440,19 @@ monitor = RV / ptLTL / named invariant # exact, compiled act = code, only if monitor admits ``` +**Named live-stream shape (Empirical as a *shape*, 2026-09-18):** +[`affirmitv/bitrate-advisor`](https://github.com/affirmitv/bitrate-advisor) +— Jev proposes ABR rungs; deterministic policy (probe × headroom, +history percentiles, loss/queue/thermal/battery) is the monitor. Jev +may only match that envelope or be more conservative. Missing the +model returns the policy's answer. Author-measured three states +(~$0.00004, 0.25–0.39 s) are a receipt for the *shape*, not a codec +benchmark. `notes.md` §44. + **Counterexample:** "the model was confident" as the monitor. **Test:** inject a monitor-violating trace the Noul would have admitted; -the sandwich must refuse. **Hypothesis.** Links: `mental-models.md` +the sandwich must refuse. **Hypothesis** as domain-general; bitrate is +Empirical as the named envelope. Links: `mental-models.md` conformal; `formal-methods.md` help list. ## 13. DST multiverse triage (Hypothesis) @@ -480,7 +506,11 @@ policy = starvation/fairness rules in code ``` **Example (Hypothesis):** incident-commander assignment; grant-panel -paper allocation; GPU scheduling. **Counterexample:** Choice over +paper allocation; GPU scheduling. **Empirical as a *shape*:** +[`affirmitv/bitrate-advisor`](https://github.com/affirmitv/bitrate-advisor) +— Jev's rung is the soft affinity; probe/history/thermal caps are the +solver; the model cannot violate them (`notes.md` §44). +**Counterexample:** Choice over assignees that ignores load. **Test:** a feasible assignment the solver finds that the Score alone would skip because it "felt" worse; hard constraints never yield. Links: `mental-models.md` §OR. @@ -565,7 +595,10 @@ doc-router 9 OCR-misses vs 28 for rules-only. [`poponline63/hermes-jev-north-star`](https://github.com/poponline63/hermes-jev-north-star): deterministic shell checks first; empty evidence refuses to judge; then one Jev call on the remainder. Empty state was self-contradictory — -that is why the refuse-empty rule exists. **Beyond SWE (Hypothesis):** +that is why the refuse-empty rule exists. +[`affirmitv/bitrate-advisor`](https://github.com/affirmitv/bitrate-advisor) +is the same sandwich on a live encoder: policy proves the cap; Jev +judges only inside it (`notes.md` §44). **Beyond SWE (Hypothesis):** recipe book ∩ "does this leftover look done?"; labor-law allowlist ∩ hiring-fit Noul; SPF/DKIM pass ∩ phishing Noul on the body. **Counterexample:** Jev on `/bin/ls` as the first tier. **Test:** planted writers never diff --git a/.agents/skills/augustus/references/mental-models.md b/.agents/skills/augustus/references/mental-models.md index ed99861..d2a6996 100644 --- a/.agents/skills/augustus/references/mental-models.md +++ b/.agents/skills/augustus/references/mental-models.md @@ -184,7 +184,10 @@ Contract. The *placement* (gather as an enumerated act) is the method; the calculator is **Hypothesis** until you log act/outcome pairs. Mapping card: `mappings.md` §6. Paying *zero* because a regex already answers is also VOI — abstain from calling any model -(`typesafe-jev-tools`, `notes.md` §42). +(`typesafe-jev-tools`, `notes.md` §42). A harness that *picks which +primitive to run* (JevML's claim: PCA / MCMC / diffusion / NCA) is the +same gate one layer down: maybe none of them (`notes.md` §44). +Hypothesis until that picker has a labeled log. **Transfers:** "ask a second question" / "retrieve one more candidate" / "run the expensive LLM" only when VOI clears the cost. Cheap fan-out @@ -247,11 +250,15 @@ probe. | Business | sales stages | "is this still a real opp?" | amount, close date in CRM | | Life | cook / rest / check | "does this look done?" | thermometer (probe) | | Org | incident command | "is this still contained?" | head-count, location | +| Infra | Postgres query planner | override join/card when confident | the stock planner (fail-open) | +| Live media | ABR rung / resolution | "which ladder step?" | probe × headroom, thermal, battery | Rejected: bandits without observed rewards; Jev as the planner that picks its next tool in a loop (`boundary-audit.md`); PufferLib Ocean scores as a capability claim (`formal-methods.md` DST trio). Mapping -card for the cross-domain loop: `mappings.md` §9. +card for the cross-domain loop: `mappings.md` §9. Soft judgment +inside a hard envelope: bitrate-advisor (ABR) and mmalisper's JOB +hybrid (Postgres plans first) — `notes.md` §44. ## Signal detection diff --git a/.agents/skills/augustus/references/mixed-architecture.md b/.agents/skills/augustus/references/mixed-architecture.md index b45bd76..6321a6b 100644 --- a/.agents/skills/augustus/references/mixed-architecture.md +++ b/.agents/skills/augustus/references/mixed-architecture.md @@ -211,6 +211,19 @@ Live ecosystem (examples of the *shape*, not SDKs to copy): - `rajdhakad9826/routeKit` — Jev estimates requirements; a deterministic policy picks the model. Jev does not choose the LLM. **Hypothesis** until your catalog (`notes.md` §33). +- [`trietphan/jev-claw`](https://github.com/trietphan/jev-claw) — same + hole for OpenClaw: Jev classifies task/complexity/risk; `decide()` in + code maps to a route; a sensitive-path regex floors risk. Confidence + is the **min** across heads. 11 offline policy tests; live 10/10 is + the author's 10 samples (`notes.md` §44). +- [`nekowasabi/jev-routing`](https://github.com/nekowasabi/jev-routing) + — host **adapter** (Go binary), not an MCP server, in front of + Claude Code / Codex / Grok Build. Compacts tool results, then one + Choice + done-Noul, then one schema. Adding it via `mcp add` makes + the catalog worse. No key → on-device classifier. Do not copy ports. +- Higgsfield API auto-routing ([tweet](https://x.com/higgsfield_ai/status/2101022473248727177)) + is the same hole on a video/image catalog. **Claim**, no labeled + receipt. - Function-calling cookbook (**Contract**): function *names* and closed-set args as questions; code still validates the call. @@ -238,7 +251,10 @@ SREGym-Lite is topology A: Jev ranks next tests/evidence; the agent still runs them and still diagnoses (`notes.md` §33). [`runta-dev/jot`](https://github.com/runta-dev/jot) is topology B with a *closed* tool catalog (the host executes; the calculator does the -math). "First general-purpose System One agent" is a claim. **Does not:** the decision model as the planner — neither inventing tools +math). "First general-purpose System One agent" is a claim. +[`nekowasabi/jev-routing`](https://github.com/nekowasabi/jev-routing) is +a host adapter in front of an existing coding CLI (not topology B, not +MCP): it peels the catalog *before* the generator sees it. **Does not:** the decision model as the planner — neither inventing tools nor picking its own next tool in a loop (standing red flag, above and in `boundary-audit.md`); skipping schemas so the model "just knows"; treating a workflow AST as a proof. The outer loop stays with the LLM or @@ -286,7 +302,11 @@ Related placements: Claude CLI vs 1.3s Jev on a 24-row filter). Log would-do until evals pass. Selective abstention (`mappings.md` §2): low confidence is `review`, not a guess. Coppe on SREGym regressions: inspect whether - confidence was high on the wrong Choice (`notes.md` §33). + confidence was high on the wrong Choice (`notes.md` §33). This hour + the *practice* is first-class: eval CLI asserts on the **action**, + not on prose; recipes span alerts / RTB / sports-bet / prediction + markets (`notes.md` §44). LLM-as-judge is not the primary System One + score (`faq.md`). Do not copy the client. - **Hybrid countable + judgment rules** — `DanRWilloughby/snifftest`: deterministic tells score 1.00; judgment rules flag only outside the unsure band. Explicit: a reading near 0.5 is *no judgment*, never a pass. @@ -323,7 +343,13 @@ decision-design card. Do not clone APIs from READMEs. | Preference lint | Per-rule Noul/Choice on a diff | Rule text, outcome map | jev-pref, JevLint | | Context / log prune | Per-line or per-block relevance | Always-keep set, recall keys | jevprune, winnow | | Exact hunk staging | Per-hunk include/exclude/mixed | `git diff`, atomic apply | git-jev-stage | -| Semantic `WHERE` | Noul/`jev_prob` over a row | SQL, indexes, LIMIT | `kylemclaren/jevql` (CLI rewrites; DB sees ordinary SQL) | +| Semantic `WHERE` | Noul/`jev_prob` over a row | SQL, indexes, LIMIT | jevql (CLI; DB sees ordinary SQL); sqlite-jev (in-engine extension) | +| Formula / query embedding | JUDGE as a function | Spreadsheet/SQL engine | judge-sheets, jevql, sqlite-jev | +| Soft ABR / live encoder | Choice over a ladder | Probe × headroom, thermal, battery | bitrate-advisor | +| Voice → typed act | Choice/Noul on a transcript | ASR producer; macOS actions | jev-voice-control (README stub) | +| Host-adapter routing | Choice next-tool + done-Noul | Shrink `tools[]`; compaction | jev-routing (not MCP) | +| Multi-model route | Classify axes; policy maps | Escalation `if`, path regex | jev-claw, routeKit | +| Finish-line gate | Noul/Score/Choice on evidence | Deterministic shell checks first | hermes-jev-north-star | | Home automation read | Choice/Score/Noul as an entity | Automations, device I/O | `AboveColin/HA-Jev` | | Browser loop without generation | Action Choice over visible elements | Perception, constraints, click | lizard-agent | | Android / macOS computer-use | Choice over prevalidated candidates | UI tree / AX / OmniParser; no generated coordinates | jev-mobile, jev-macos-loop | @@ -332,9 +358,7 @@ decision-design card. Do not clone APIs from READMEs. | Closed-catalog System One shell | Choice over host tools | Execute, arithmetic, credentials | jot | | Jump-by-description | Noul relevance on a local shortlist | zoxide index, local paths only | joxide | | Game move | Choice over legal actions | Rules, legality, win check | jev-plays-games | -| Finish-line gate | Noul/Score/Choice on evidence | Deterministic shell checks first | hermes-jev-north-star | | Analyst attention cascade | Step-level silent-failure Nouls | Grouping, LLM autopsy | OpenSmoke | -| Formula / query embedding | JUDGE as a function | Spreadsheet/SQL engine | judge-sheets, jevql | On-device / Home Assistant / mobile are newly-feasible via the economics inversion, not proven ports of every app. Named placements this hour diff --git a/.agents/skills/augustus/references/validation.md b/.agents/skills/augustus/references/validation.md index 165d13f..3ad635f 100644 --- a/.agents/skills/augustus/references/validation.md +++ b/.agents/skills/augustus/references/validation.md @@ -105,7 +105,11 @@ Adoption pattern (AntonioCoppe/jev-harness): run the Jev judgment in parallel wi the live system and only **log what you would have done** (policy + gate applied) until behavioral evals over replayed fixtures pass; then flip to enforcement. Assert on the *action* (block/warn/pass), not on free text. This is the safe path for any -confidence gate added to an existing pipeline. +confidence gate added to an existing pipeline. Harbor/jevals-adjacent +practice, not a second eval product: LLM-as-judge is not the primary +System One score (`faq.md`). Recipes in that repo (alerts, RTB, sports-bet, +prediction-markets) are existence proofs of the same substrate across +business and life, not SWE-only (`notes.md` §44). Do not copy the client. ## Frontmatter (by agents and by Jev rankers) diff --git a/CHANGELOG.md b/CHANGELOG.md index ec703d1..486b417 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -99,6 +99,20 @@ Pegged against [`typesafe-ai/skills` v0.5.7](https://github.com/typesafe-ai/skil two-layer finish gate. pi-jev (not pi-jev-context). jev-plays-games option-order probe. joxide jump-by-description. laya-typed-decisions companion packaging. No wrapper. +- Hourly ~12:58 Boise fold (`research/notes.md` §44): Archer still + Watch (no architecture rewrite). Store-index fork: in-engine + ([sqlite-jev](https://github.com/mgaitan/sqlite-jev), pg-jev cousin) + vs CLI rewrite (jevql). Soft judgment inside a hard envelope + ([bitrate-advisor](https://github.com/affirmitv/bitrate-advisor); + mmalisper JOB planner +12% geomean, author-reported; join-order + Choice alone was 2× slower). Distill-to-device as a *memory* gate + (jev-gate, already §33). Encoder vs decoder open-replica receipts + (openjev-lm $0/call overnight CPU). jev-harness as Harbor-adjacent + practice (assert on action). Host adapter + ([jev-routing](https://github.com/nekowasabi/jev-routing), not MCP); + OpenClaw typed routing ([jev-claw](https://github.com/trietphan/jev-claw)). + Voice-control and JevML are README stubs. Higgsfield auto-routing is + a claim. No wrapper. - Effect-oriented loops (`notes.md` §28, `mappings.md` §19): Ward's ZIO client keeps Jev as the outer Choice and the handler as the effect. Not Effect.ts. GLiNER author: GLiNER2 "like jev" is GLiGuard diff --git a/README.md b/README.md index 47d4592..6ba4742 100644 --- a/README.md +++ b/README.md @@ -50,7 +50,8 @@ never launder a Noul as a proof. exact-text keep/drop, env triage, moderation/ranking, skill routing - `.agents/skills/augustus/references/faq.md` — "just classification", stack replacement, Jev vs open head vs encoder vs LoRA vs constrained AR, - GLiNER vs GLiClass vs CLIP, LLM-as-judge, not-another-how-to + GLiNER vs GLiClass vs CLIP, LLM-as-judge, in-engine vs CLI store, + hard envelope (bitrate / planner), not-another-how-to - `.agents/skills/augustus/references/mappings.md` — classical-method mappings with boundaries, counterexamples, acceptance tests (including Hypothesis cards §6–§19 — promote only with a test that ran) diff --git a/docs/ecosystem.md b/docs/ecosystem.md index 3637daa..98038b9 100644 --- a/docs/ecosystem.md +++ b/docs/ecosystem.md @@ -25,7 +25,7 @@ weekdays. Jev is the densest public corpus, not the class monopoly. ### Agent harnesses & self-supervision - **Kevthetech143/super-jev** — domain-independent loop: observe → questions → decide → **permit (independent of confidence)** → execute (idempotency key) → verify → JSONL replay. -- **AntonioCoppe/jev-harness** — policy + confidence gate + shadow mode + offline eval CLI; 24-row filter 48.9s (Claude CLI) vs 1.3s Jev. Selective abstention. +- **AntonioCoppe/jev-harness** — policy + confidence gate + shadow mode + offline eval CLI asserting on the **action**; 24-row filter 48.9s (Claude CLI) vs 1.3s Jev. Harbor/jevals-adjacent practice. `notes.md` §33, §44. - **Friedjof/jev-mobile** — durable Android worker + Mobile MCP; Jev sees prevalidated candidates only. `notes.md` §33. - **jcpsimmons/jev-macos-loop** — Apple-silicon computer-use; local OmniParser/OCR/AX; text-only Jev. Finder demo independently verified. - **rajdhakad9826/routeKit** — Jev estimates task requirements; policy engine selects the LLM. Jev does not pick the model. @@ -48,7 +48,7 @@ Per-keystroke launchers (104ms median, sequence-tagged staleness), firehose mode ### Local / open heads & GLi\* species - **GLiNER / GLiNER2.5 / GLiClass** — species map: locate spans vs categorize the sequence vs local multi-head (fastino-ai GLiNER2.5 CPU-first). Peer of Jev, not a footnote. `references/judgment-class.md`. Author primary source: GLiNER2 "like jev" is schema-conditioned categorize (GLiGuard), not a Noul (`notes.md` §28). 36× Browser Use claim is a tweet (`notes.md` §25). - **GLiGuard** (fastino-ai) — 0.3B GLiNER2 encoder, checkpoint `fastino/gliguard-LLMGuardrails-300M`. One bidirectional pass over a safety schema. Same interface shape as batched questions; different objective. Not a Jev weight clone. `judgment-class.md`; `notes.md` §30. -- **DECRUX9812/openjev-lm** — Qwen2.5-0.5B+LoRA distilled from hosted Jev answers; 65/70 = 92.9% on 70 hand-labelled rows (one annotator, one domain, one seed) overnight on 6 vCPU. Its 98.1% on fresh rows is teacher *agreement*, not gold. +- **DECRUX9812/openjev-lm** — Qwen2.5-0.5B+LoRA distilled from hosted Jev answers; 65/70 = 92.9% on 70 hand-labelled rows (one annotator, one domain, one seed) overnight on 6 vCPU, $0/call. Its 98.1% on fresh rows is teacher *agreement*, not gold. Receipts pattern: `notes.md` §25, §44. - **convaiinnovations/laya** — open Choice/Score/Noul head, text-only, 512 tok. Companion packaging this hour: [`laya-typed-decisions`](https://huggingface.co/convaiinnovations/laya-typed-decisions) (421.3M, acc 0.766 / Brier 0.066 unverified). `notes.md` §18, §42. - **zmtomorrow/TypeAR** — constrained autoregressive decoding surface: typed fields on a pretrained open model, no retraining. Not a proper-scoring decision head. `research/notes.md` §32. - **stephanj/pcdServer** — native Parallel Constrained Decoder (C++20, llama.cpp GGUF, Apple+Linux). TypeAR-class serving: 2–256 enums, 1–63 parallel fields; softmax over allowed values is not a Noul. `notes.md` §42. @@ -77,6 +77,12 @@ Per-keystroke launchers (104ms median, sequence-tagged staleness), firehose mode - **TheoOliveira/pi-jev** — Pi semantic tool/skill routing. Not kevinpita/pi-jev-context. - **vtrivedy/jev-plays-games** — Choice over legal moves; probabilities ≠ win odds. - **ant4g0nist/joxide** — zoxide index, Jev shortlist, local paths only. +- **mgaitan/sqlite-jev** — batched NL judgments as a SQLite loadable extension (`jev_rows`). In-engine sibling of jevql's CLI rewrite; inspired by pg-jev. Semantic full scan, not an index. `notes.md` §44. +- **affirmitv/bitrate-advisor** — live ABR: Jev proposes, deterministic policy clamps (never bolder). Missing the model returns policy. `notes.md` §44. +- **nekowasabi/jev-routing** — Go host adapter for Claude Code / Codex / Grok Build. Not MCP, not npx. `notes.md` §44. +- **trietphan/jev-claw** — OpenClaw typed routing: Jev classifies, `decide()` in code. `notes.md` §44. +- **chris-wozniczek/jev-voice-control** — Speech → Jev → macOS actions. README-only this pass. Hypothesis. `notes.md` §44. +- **gamesonrblx/JevML** — claimed PCA/MCMC/diffusion/NCA primitives + a picker. README-only. Hypothesis. `notes.md` §44. ### Skills & tooling - **typesafe-ai/skills** — official skill (contracts/patterns). @@ -93,7 +99,8 @@ owns control. Movers that sharpened the card: `git-jev-stage` (exact hunk Choice), `jevprune` (per-line relevance with an always-keep set), `llama-index-jev` (rerank fails open / select fails closed), `jev-pref` (AGENTS.md as criteria), `lizard-agent` (no LLM when nothing needs writing), -`jevql` (judgment as SQL `WHERE`), OpenSmoke (Jev over every step, LLM only +`jevql` (judgment as SQL `WHERE`; CLI so Postgres never sees `jev()`), +`sqlite-jev` (in-engine SQLite extension; same hole), OpenSmoke (Jev over every step, LLM only on flags). Neighbor skills `tenbin` and `decision-first` are *not* Augustus clones — they own lint/eval and try-Jev-first habit. Entropy allocator (**Hypothesis**): cheap typed scorers for low- and medium-entropy diff --git a/research/archive/findings.md b/research/archive/findings.md index 39efe52..3dd3f48 100644 --- a/research/archive/findings.md +++ b/research/archive/findings.md @@ -648,5 +648,38 @@ are producers not perceive; two call shapes removed from the optimizer card; $0.042/MTok tagged vendor-stated; GodsBoy 94.4% tagged exploratory. `notes.md` §43. +## Batch #27 (2026-09-18, ~12:58 Boise hourly) + +Note: `research/notes.md` §44. Docs-only. Archer still Watch (no +architecture rewrite). Do not rehash §42 HIGH. + +- **sqlite-jev (Contract as README):** in-engine SQLite extension; + batched `jev_rows`; sibling *pattern* to jevql, different serving + (DB sees `jev()`). Inspired by pg-jev. Semantic full scan, not an + index. License file absent. 0★. +- **bitrate-advisor (Empirical as a shape):** Jev proposes ABR; policy + is the envelope; never bolder. Missing model → policy answer. + Three-state receipt author-reported. +- **jev-routing (Contract as README):** Go host adapter, not MCP, for + Claude/Codex/Grok. Compact then one Choice + done. +- **jev-claw (Empirical as author's 10/10 + 11 offline tests):** Jev + classifies; `decide()` maps; path regex floors risk; confidence is + min. +- **jev-harness practice:** already §33; this hour assert-on-action + as Harbor-adjacent substrate; recipes across business/life. +- **openjev-lm / jev-gate / DeBERTa / mini-jev-runs / tree-cap / + jev-pref:** frames only (receipts economics; memory gate; encoder vs + decoder replica). No rewrite. +- **jev-voice-control / JevML:** README-only stubs. Hypothesis. +- **X:** mmalisper JOB hybrid +12% geomean, join-order 2× slower, + fail-open to Postgres (author-reported). Higgsfield GenAI auto-route + is a claim. + +Cross-repo addition: (as) structured-store semantic index has an +in-engine vs CLI fork; (at) soft judgment inside a hard envelope +(ABR, planner); (au) distill-to-device as a context sieve, not only +an action gate. + + diff --git a/research/notes.md b/research/notes.md index f14c821..0957583 100644 --- a/research/notes.md +++ b/research/notes.md @@ -2121,3 +2121,198 @@ Kept, and patched in the cards: Not patched, standing risk: doctrine is copied across cards and the next fold is where copies drift; `notes.md` §1 is a one-day pin, not a live contract. + +## 44. Hourly fold ~12:58 America/Boise (2026-09-18) — store siblings, hard envelopes, distill-to-device, harness practice + +Window: America/Boise ~12:58 ≈ 18:58 UTC. Watch run 185349. Novel +versus §42 (~11:59). Docs-only. No Jev wrapper, no serving-stack +how-to, no copied SQL/`advise()`/proxy ports/`predict()`. Local +`/workspace/jev-archive` path for this run was not present; GitHub +READMEs, Hub cards, and X posts fetched live. HTTP 200 on cited URLs. +**Archer drop still WATCH** — no architecture rewrite this hour. User +watch says ~65% done, Qwen3.8-27B multimodal no-audio still expected +~19 Sep Boise. This pass did not retrieve a new Archer status tweet +(rate-limit / query constraint); Hub was not treated as a landing. + +Do **not** rehash pcdServer, jevql's already-folded CLI frame, +OpenSmoke, jev-mode, jot, typesafe-jev-tools, openevals, +hermes-north-star. Already-folded HIGH from §33 (jev-harness, +openjev-lm, jev-gate-student-b, open-jev-deberta, mini-jev-runs, +jev-tree-choice-cap, jev-pref) get a *frame*, not a second card. + +Six frames, then the artifacts. + +**(a) Structured-store semantic index.** Cheap exact predicates first; +typed questions on the remainder. Two *forks* of the same hole: +**in-engine extension** (`mgaitan/sqlite-jev`; inspired by +`realZachi/pg-jev`) vs **out-of-process CLI** (`kylemclaren/jevql` — +vanilla Postgres never sees `jev()`). zoxide (`joxide`, §42) is the +same hole over a path index. A semantic full scan is not an index. +Row contents leave the store: same residency warning as AU health +(§33). + +**(b) Distill-to-device memory/context gates.** System One as a +**sieve**, not only an action permit. Already §33 / `applied-mappings.md` +§1: `jev-gate-student-b` P(relevant) from yes/no logits. This hour the +meta is: vector recall → local yes/no gate → inject or stub. Fail-open +on errors. Teacher-copy ≠ gold. + +**(c) Encoder vs AR open replicas.** Already the three-open-paths cut +(§42). Encoder DeBERTa = public gold, OOD drop; decoder LoRA +(openjev-lm, jev-gate) = teacher-copy; constrained AR (TypeAR / +pcdServer) = softmax over allowed tokens, not a Noul. Do not pick a +path until meta-VOI says a model is needed at all. + +**(d) Soft judgment inside a hard safety envelope.** The model may +only **match the envelope or be more conservative**. Deterministic +policy is load-bearing; missing the model must still be safe. +bitrate-advisor is the named live-stream shape. mmalisper's JOB +planner is the named search/control shape (Postgres plans first; Jev +overrides only when confident). + +**(e) Perception → decision.** ASR / vision produce schema'd state; +System One decides; code acts. SAM and ASR are **producers**, not the +perceive species (§43). `jev-voice-control` is a README-only stub of +that pipeline. + +**(f) Harbor/jevals-style harness + shadow + confidence.** Assert on +the **action**, not on free text. LLM-as-judge is not the primary +System One score (`faq.md`). jev-harness already named this; this hour +the practice is first-class: recipes across alerts / RTB / sports-bet / +prediction-markets, offline fixtures, shadow until evals pass. + +### HIGH + +1. **[`mgaitan/sqlite-jev`](https://github.com/mgaitan/sqlite-jev)** + (created 18:25Z, C, license file absent, 0★). Loadable SQLite + extension: batched NL Noul/Choice/Score over row objects + (`jev_rows` virtual table; up to 40 rows per shared state). + Inspired by [`realZachi/pg-jev`](https://github.com/realZachi/pg-jev) + (in-engine Postgres, 152★ this pass — pointer, not a second card). + Sibling *pattern* to jevql, **not** the same serving choice: here + the database *does* see `jev()` as SQL. README limits: semantic + full scan, not an index; deterministic SQLite filters first; + `max_rows` is a spend guard; row contents go to TypeSafe. Thresholds + stay in SQL so they can rise with false-positive cost. Mock-server + tests never call TypeSafe. Do not copy `.load`, env, or SQL + signatures. Card: `mappings.md` §4. + +2. **[`AntonioCoppe/jev-harness`](https://github.com/AntonioCoppe/jev-harness)** + — already §33 (policy, confidence gate, shadow, 24-row 48.9s Claude + CLI → 1.3s Jev). This hour the *practice*: Harbor/jevals-adjacent + measurement substrate. Eval CLI asserts on the **action**, not on + prose. Recipes span alerts, NL row-filter, high-freq reflex (order / + RTB / fraud), prediction-market and sports-bet gates — business and + life, not SWE-only. Explicitly **not** a `/compact` replacement. + Do not copy the client. `validation.md`; gallery. + +3. **[`DECRUX9812/openjev-lm`](https://github.com/DECRUX9812/openjev-lm)** + — already §25. This hour the *receipts pattern*: overnight 6-vCPU, + $0/call, two independently written harnesses both 65/70 = 92.9% on + the same 70 hand-gold rows; 98.1% on 106 later postings is teacher + **agreement**, not gold; rare-class cells are one-row wide. Encoder + vs this decoder LoRA is the (c) fork, not a new architecture. Do + not copy train commands. + +4. **`SargeDev/jev-gate-student-b` + `jev-distill-corpus`** — already + §33. This hour the meta: **context sieve / memory gate**, not only + an action gate. Vector recall → P(relevant) yes/no logits → inject + or stub. Fail-open. Teacher-copy. `applied-mappings.md` §1. + +5. **`com-kotobalabs/open-jev-deberta-v3-large`** — already §33 + (434M, Banking77/SST5/BoolQ, in-domain ECE 0.022, OOD acc 0.690). + This hour: the **encoder** arm of open-replica, as opposed to + decoder LoRA (openjev-lm / jev-gate) and constrained AR. Public + gold, not a Jev teacher. Do not overwrite those numbers. + +### MED + +6. **[`affirmitv/bitrate-advisor`](https://github.com/affirmitv/bitrate-advisor)** + (created 15:27Z, MIT, TypeScript, 0★). Live-stream ABR: Jev + proposes initial rung / ceiling / resolution / next step from + telemetry + venue/carrier history; **deterministic policy fences + it in**. Jev may be as bold as the measured network allows and as + cautious as it likes, **never bolder**. Without an API key the same + call returns the policy's answer (`source: "policy"`). Power plan + (finish-the-game battery/thermal) is deterministic; Jev may only + make it more conservative. Author-measured 2026-09-18, three + states, OpenRouter billed ~$0.000041–0.000044, 0.25–0.39 s; ~$0.015 + per hour of stream at one decision / 10 s — **author-reported**, + not re-run. Domain: youth-sports phone streams (life / business), + not a codec tutorial. Empirical as a *shape* for §12 / §15 / §18. + Do not copy `advise()` or keys. + +7. **[`chris-wozniczek/jev-voice-control`](https://github.com/chris-wozniczek/jev-voice-control)** + (created 18:47Z, license absent, README-only this pass, 0★). + Speech → Jev typed decisions → macOS actions. Menu-bar Swift + **claim**. Hypothesis as a product; useful as the perception→decision + pipeline with ASR as producer (§39 / §43). Do not invent a Swift + API. + +8. **`doeixd/jev-pref`** — already the preference-lint contract + (YOU define the rule / Jev classifies / code maps outcome). No + rewrite. This hour it sits next to bitrate's envelope: policy-as- + prefs is the same ownership split on a diff instead of a live + stream. + +9. **[`nekowasabi/jev-routing`](https://github.com/nekowasabi/jev-routing)** + (created 08:13Z, MIT, Go, 0★). Host **adapter**, not an MCP + server: one Go binary in front of Claude Code / Codex / Grok + Build. Compacts tool results (verbatim drop/truncate, same + contract as fast-jev-compaction), then one Choice (next tool) + + Noul (done), then shrinks `tools[]` to **one schema**. README: + adding this via `mcp add` makes the catalog *worse*. No key → + on-device classifier. Do not copy ports, env, or install. + Host-adapter breadth, not a new species. `applied-mappings.md` §5. + +10. **[`trietphan/jev-claw`](https://github.com/trietphan/jev-claw)** + (created 18:44Z, MIT, JS, 0★). Typed routing for OpenClaw + agents. **Jev classifies** (task_type / complexity / risk / + second-opinion); **`decide()` in code** maps to a route. + Sensitive-path regex floors risk even if Jev underrates a + migration. Confidence is the **minimum** across classifications, + not the average. 11 offline policy tests (no network). Live eval + 10/10 on the author's 10 samples — **author-reported**, tiny. + Same hole as routeKit (§33): Jev does not get to skip the + escalation `if`. Do not copy the eight route names as doctrine. + +11. **[`gamesonrblx/JevML`](https://github.com/gamesonrblx/JevML)** + (created 18:47Z, license absent, README-only, 0★). "PCA / MCMC / + text-diffusion / NCA primitives + a harness that picks the right + tool." Meta-VOI adjacent: which primitive, if any. Hypothesis. + Do not invent those APIs. + +12. **`Mikhail/mini-jev-runs`** — already §33 (27.9k option-logit + runs, frozen Qwen3-4B, no token generated). Calibration / + constrained-decode corpus. No rewrite. + +13. **`reachjalil/jev-tree-choice-cap`** — already §33 / mappings §5. + Hierarchical Choice under the 255 cap. No rewrite. + +### X discourse this hour (verified) + +- [@mmalisper](https://x.com/mmalisper/status/2101001041903009987) + (17:30:57Z) and thread: Jev-assisted Postgres query planner on the + Join Order Benchmark. Join-order Choice **2× slower** (defaulted to + smallest table). Cardinality estimates helped when outside context + informed the plan; when Jev was wrong, one query was an **order of + magnitude slower**. Hybrid: Postgres plans first; Jev overrides + **only when confident** → **+12% geomean**, no dramatic slowdowns. + Downside: a Jev call is 100s of ms, not yet practical on every + plan. **Author-reported**, not re-run. Frame (d) + search/control + (§9): the planner is the envelope; confidence is the gate. + Fail-open to Postgres. + +- [@higgsfield_ai](https://x.com/higgsfield_ai/status/2101022473248727177) + (18:56:07Z) and [demo](https://x.com/higgsfield_ai/status/2101022133753430365) + (18:54:46Z): "perfect use case" — Jev auto-routes GenAI (video/image) + models for cost/speed/quality on the Higgsfield API. **Claim**, no + labeled catalog receipt. Same hole as routeKit / jev-claw: + classify requirements, policy picks the generator. Hypothesis + until *your* catalog. + +Cards: `mappings.md` §4 / §9 / §12 / §15 / §18; `mixed-architecture.md` +gallery + host-adapter note; `applied-mappings.md` §1 / §5; +`validation.md` harness practice; `mental-models.md` envelope + +planner; FAQ in-engine vs CLI; `judgment-class.md` encoder vs decoder +replica (no new species). No wrapper. diff --git a/research/refresh-log.md b/research/refresh-log.md index 03d6d9b..b95a391 100644 --- a/research/refresh-log.md +++ b/research/refresh-log.md @@ -362,4 +362,20 @@ - Dropped: Harbor wording matches the verifiers v1 post; mini-jev-runs dual use; translation-invariant claim; SREGym arithmetic. +## 2026-09-18 18:58 UTC — ~12:58 Boise hourly fold + +- America/Boise ~12:58. Watch run 185349. Docs-only. Archer still + **WATCH** (no architecture rewrite; ~65% done is user watch, not a + retrieved status tweet this pass). Local jev-archive path absent; + READMEs and X fetched live. +- Novel vs §42: sqlite-jev in-engine store fork; bitrate-advisor hard + envelope; jev-routing host adapter; jev-claw OpenClaw routing; + mmalisper JOB planner thread; Higgsfield auto-route claim. +- Frames only (already folded): jev-harness practice, openjev-lm + receipts, jev-gate as memory sieve, encoder vs decoder replica. + README stubs: jev-voice-control, JevML. Skip rewrite: jev-pref, + mini-jev-runs, jev-tree-choice-cap. +- notes.md §44; sources.json; findings.md batch #27. No wrapper. + + diff --git a/research/sources.json b/research/sources.json index 69c1e84..a2cfeb3 100644 --- a/research/sources.json +++ b/research/sources.json @@ -1,6 +1,6 @@ { "refresh_cadence": "hourly", - "retrieved": "2026-09-18T17:55Z", + "retrieved": "2026-09-18T18:58Z", "sources": [ { "kind": "docs", @@ -1177,6 +1177,60 @@ "title": "brunoqgalvao: Pareto cost/perf takeaway", "url": "https://x.com/brunoqgalvao/status/2101008125356626019", "note": "2026-09-18T17:59:06Z. Closes earlier 4-task ~1200-item thread. Author-reported. notes.md \u00a742." + }, + { + "kind": "github", + "title": "mgaitan/sqlite-jev", + "url": "https://github.com/mgaitan/sqlite-jev", + "note": "In-engine SQLite extension; batched jev_rows. Sibling pattern to jevql, not the CLI rewrite. License absent. notes.md \u00a744." + }, + { + "kind": "github", + "title": "realZachi/pg-jev", + "url": "https://github.com/realZachi/pg-jev", + "note": "In-engine Postgres cousin sqlite-jev is inspired by. Pointer only. notes.md \u00a744." + }, + { + "kind": "github", + "title": "affirmitv/bitrate-advisor", + "url": "https://github.com/affirmitv/bitrate-advisor", + "note": "Live ABR: Jev proposes, deterministic envelope clamps. MIT. notes.md \u00a744." + }, + { + "kind": "github", + "title": "chris-wozniczek/jev-voice-control", + "url": "https://github.com/chris-wozniczek/jev-voice-control", + "note": "Speech \u2192 Jev \u2192 macOS. README-only this pass. Hypothesis. notes.md \u00a744." + }, + { + "kind": "github", + "title": "nekowasabi/jev-routing", + "url": "https://github.com/nekowasabi/jev-routing", + "note": "Go host adapter for Claude/Codex/Grok. Not MCP. MIT. notes.md \u00a744." + }, + { + "kind": "github", + "title": "trietphan/jev-claw", + "url": "https://github.com/trietphan/jev-claw", + "note": "OpenClaw typed routing: classify then decide() in code. MIT. notes.md \u00a744." + }, + { + "kind": "github", + "title": "gamesonrblx/JevML", + "url": "https://github.com/gamesonrblx/JevML", + "note": "Claimed PCA/MCMC/diffusion/NCA + picker. README-only. Hypothesis. notes.md \u00a744." + }, + { + "kind": "social", + "title": "mmalisper: Jev-assisted Postgres query planner +12% JOB", + "url": "https://x.com/mmalisper/status/2101001041903009987", + "note": "2026-09-18T17:30:57Z. Hybrid: PG plans first; Jev overrides when confident. Join-order Choice 2x slower. Author-reported. notes.md \u00a744." + }, + { + "kind": "social", + "title": "higgsfield_ai: GenAI model auto-routing", + "url": "https://x.com/higgsfield_ai/status/2101022473248727177", + "note": "2026-09-18T18:56:07Z. Claim, no labeled catalog. Same hole as routeKit. notes.md \u00a744." } ] } From efe46943cfcaae6c53ec9486971a0a8cbe3f4254 Mon Sep 17 00:00:00 2001 From: Cursor Agent Date: Fri, 18 Sep 2026 20:04:01 +0000 Subject: [PATCH 02/43] Fold kev as the runnable Archer reconstruction on the trained path Place jaredpalmer/kev beside Laya / Nimble / Archer Watch: Qwen2.5-0.5B LoRA + pointer, Apache-2.0, POST /v1/systemone drop-in. Public gold, not a Jev teacher. Contrast vs TypeAR, encoder DeBERTa, proprietary Jev. Cite README isolation/ECE/acc/permute/IIA/forgery; laptop-local development/eval, not a knowledge substitute. Docs-only; no serve how-to. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com> --- .agents/skills/augustus/SKILL.md | 10 +- .../augustus/references/applied-mappings.md | 2 +- .agents/skills/augustus/references/faq.md | 28 ++++-- .../augustus/references/formal-methods.md | 3 +- .../augustus/references/judgment-class.md | 97 +++++++++++++++---- .../augustus/references/mental-models.md | 3 +- .../augustus/references/mixed-architecture.md | 9 +- .../references/optimizer-integration.md | 4 +- .../skills/augustus/references/validation.md | 21 +++- CHANGELOG.md | 11 +++ README.md | 4 +- docs/ecosystem.md | 3 +- research/archive/findings.md | 25 +++++ research/notes.md | 81 ++++++++++++++++ research/refresh-log.md | 14 +++ research/sources.json | 20 +++- 16 files changed, 287 insertions(+), 48 deletions(-) diff --git a/.agents/skills/augustus/SKILL.md b/.agents/skills/augustus/SKILL.md index 4865cd3..0a73aab 100644 --- a/.agents/skills/augustus/SKILL.md +++ b/.agents/skills/augustus/SKILL.md @@ -1,6 +1,6 @@ --- name: augustus -description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, openjev-lm, Nimble, encoder DeBERTa, LoRA distill), announced open decision-model (Watch), constrained-AR (TypeAR, pcdServer), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"eval path\", \"jevals\", \"Harbor taskset\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge, TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, or \"formally verify with Jev\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." +description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill), announced open decision-model (Watch), constrained-AR (TypeAR, pcdServer), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"eval path\", \"jevals\", \"Harbor taskset\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge, TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, or \"formally verify with Jev\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." license: MIT metadata: version: 0.3.0 @@ -23,8 +23,8 @@ Noul), not the monopoly. This skill owns **where judgment belongs**; the official `typesafe-ai` skill plus the live docs own Jev integration contracts — read them before writing Jev API code. Neighbor skills `tenbin` (lint/measure) and `decision-first` (try-Jev-first habit) own -their jobs. Do not collapse into a TypeSafe how-to, a Laya install, or -a GLiClass or GLiNER tutorial. +their jobs. Do not collapse into a TypeSafe how-to, a Laya install, a kev +serve, or a GLiClass or GLiNER tutorial. Pick the **pillar** from the hole (expected utility, VOI, MCDA, signal detection, search/control, org/safety, formal methods), then the @@ -113,8 +113,8 @@ classical method you already trust, substitute it, classify the win | Familiar method | Judgment shape | Detail | |---|---|---| | Mental models across domains (not SWE-only) | EU, abstention, VOI, MCDA, SDT, search/control, Leveson, NATM/Norman/snap-fit | `references/mental-models.md` | -| Judgment-model class (Jev is exemplar, not monopoly) | Species: decide / locate (GLiNER) / categorize (GLiClass) / rank / perceive; open heads include encoder DeBERTa, LoRA distill, openjev-lm | `references/judgment-class.md` | -| Open weights vs constrained decoding vs proprietary API | Three open paths: encoder open-jev / AR constrained decode (TypeAR + pcdServer) / trained decision-only (Laya, Nimble, Archer Watch). Softmax over allowed tokens ≠ Noul | `references/judgment-class.md` (when-to-use table) | +| Judgment-model class (Jev is exemplar, not monopoly) | Species: decide / locate (GLiNER) / categorize (GLiClass) / rank / perceive; open heads include encoder DeBERTa, LoRA distill, openjev-lm, kev | `references/judgment-class.md` | +| Open weights vs constrained decoding vs proprietary API | Three open paths: encoder open-jev / AR constrained decode (TypeAR + pcdServer) / trained decision-only (Laya, Nimble, kev, Archer Watch). Softmax over allowed tokens ≠ Noul | `references/judgment-class.md` (when-to-use table) | | Entropy as allocator (low / medium / high) | Typed low+medium decisions → System One marginals; high-entropy synthesis → frontier decoder. Product rhetoric, not a meter. **Hypothesis** | `references/judgment-class.md` | | Formal / semi-formal (proof vs judgment) | Sensor vs constraint vs searchlight; Alloy vs Apalache; DST trio; TOCTOU-of-Noul, AI×FM | `references/formal-methods.md` (one-screen: `references/formal-semi-formal.md`) | | Mixed architecture (judgment model + LLM) | Provider judges, LLM writes, code owns control; not a stack replacement | `references/mixed-architecture.md` | diff --git a/.agents/skills/augustus/references/applied-mappings.md b/.agents/skills/augustus/references/applied-mappings.md index 91abe7c..889fb82 100644 --- a/.agents/skills/augustus/references/applied-mappings.md +++ b/.agents/skills/augustus/references/applied-mappings.md @@ -3,7 +3,7 @@ These cards are *where a judgment-class model sits* in running software. They are family-agnostic: the **typed judgment provider** is TypeSafe Jev by default (live docs / `typesafe-ai`); an open Choice/Score/Noul head -(e.g. Laya) is a substitute you must self-eval (`research/notes.md` §18); +(e.g. Laya, kev) is a substitute you must self-eval (`research/notes.md` §18, §45); GLiNER (locate) / GLiClass (categorize) / listwise rankers / vision scorers are cousin species with different objectives (`judgment-class.md`). Do not copy request fields from this file. diff --git a/.agents/skills/augustus/references/faq.md b/.agents/skills/augustus/references/faq.md index d4a1edc..e5689ba 100644 --- a/.agents/skills/augustus/references/faq.md +++ b/.agents/skills/augustus/references/faq.md @@ -60,10 +60,11 @@ needs a paragraph, the generator re-enters. No. Augustus designs for the whole class of fast/cheap categorization-classification-scoring models. TypeSafe Jev is the documented exemplar (typed Choice / Score / Noul, live docs). Neighbors -in the class — open System-1 / decision-model heads (Laya, openjev-lm, -encoder DeBERTa, LoRA distill; Hume's 27B drop is Watch), constrained-AR -(TypeAR), GLiNER/GLiClass encoder family (locate vs categorize vs local -multi-head), listwise/pairwise rankers, vision scorers — are substitutes +in the class — open System-1 / decision-model heads (Laya, kev, +openjev-lm, encoder DeBERTa, LoRA distill; Hume's 27B drop is Watch), +constrained-AR (TypeAR), GLiNER/GLiClass encoder family (locate vs +categorize vs local multi-head), listwise/pairwise rankers, vision +scorers — are substitutes or cousins. Pick the family from the hole, then the vendor (`judgment-class.md` species map and when-to-use table). `typesafe-ai` still owns *Jev* contracts; other families own their own @@ -85,19 +86,28 @@ on fresh rows measures *agreement with the teacher*, not gold. Self-eval on your own independent labels before you treat it as a decision API (`notes.md` §25). +[jaredpalmer/kev](https://github.com/jaredpalmer/kev) is the laptop-local +System One **API drop-in** on that same open path: Qwen2.5-0.5B LoRA + +pointer, public gold not a Jev teacher, official SDK with a `base_url` +change. Use it for development and eval. Do not use 0.5B ID ECE as a +knowledge or frontier substitute (`notes.md` §45). + ## Open weights vs Jev vs constrained decoding vs encoder vs LoRA? Five surfaces, not one family (`judgment-class.md` when-to-use table). Three *open* paths sit beside proprietary Jev: **encoder** open-jev (DeBERTa, public gold), **AR constrained decode** (TypeAR; native [pcdServer](https://github.com/stephanj/pcdServer) GGUF serving), -**trained decision-only** (Laya / Nimble / Archer Watch). Proprietary Jev is the documented decision API; you do not hold the +**trained decision-only** (Laya / Nimble / **kev** / Archer Watch). Proprietary Jev is the documented decision API; you do not hold the weights, so checks around the boundary stay black-box -(`formal-methods.md`). A trained decision-only open head (Laya, +(`formal-methods.md`). A trained decision-only open head (Laya, kev, openjev-lm, encoder DeBERTa, a LoRA student) copies the Choice / Score / Noul *shape* and moves eval onto you. Distills trained on Jev's *answers* (openjev-lm, jev-gate-student-b) are teacher-copies — read -agreement separately from gold. Encoder open-jev +agreement separately from gold. **kev** is not that distill: CE on +public labelled outcomes, pointer readout, isolation probes, ID ECE +0.065 (0.031 after temperature scaling) on 1,350 questions — still +self-eval, still not OOD (`notes.md` §45). Encoder open-jev ([DeBERTa-v3-large](https://huggingface.co/com-kotobalabs/open-jev-deberta-v3-large)) was trained on public gold, not Jev; in-domain ECE 0.022, OOD acc 0.854→0.690. Constrained autoregressive decoding (TypeAR; README names @@ -134,7 +144,9 @@ Hole first, logo last. These are **species**, not aliases limits (e.g. 255-way Choice) are Jev's, not the class's. Distilled open heads (openjev-lm, jev-gate-student-b) copy the *teacher*, not independent gold. Encoder open-jev (DeBERTa) is the same *shape* on - public gold — still self-eval, especially OOD. Hume's 27B drop is + public gold — still self-eval, especially OOD. **kev** is the + causal-decoder + pointer productization of Archer's reconstruction on + public gold (API-compatible; not a teacher-copy). Hume's 27B drop is Watch. When-to-use axes: `judgment-class.md`. - **Cross-encoder / listwise ranker:** order of a retrieved shortlist. Fail **open** (keep retrieval order). Translation-invariant listwise diff --git a/.agents/skills/augustus/references/formal-methods.md b/.agents/skills/augustus/references/formal-methods.md index 1f20888..2753eca 100644 --- a/.agents/skills/augustus/references/formal-methods.md +++ b/.agents/skills/augustus/references/formal-methods.md @@ -59,7 +59,8 @@ deployment control, not a feud with TypeSafe. Holding those weights, when they exist, still does not discharge a proof. A constrained-AR softmax (TypeAR, pcdServer) is still a sensor: it is not a discharged proof because the next token stayed in a declared set -(`notes.md` §42). +(`notes.md` §42). A local kev pointer-softmax is the same sensor on the +trained decision-only path (`notes.md` §45). Existing grammar: composition-algebra position 9 (verifier) — verdicts are evidence, not enforcement. Position 3 (gate) — a filter is not diff --git a/.agents/skills/augustus/references/judgment-class.md b/.agents/skills/augustus/references/judgment-class.md index a03ea2b..0fa7252 100644 --- a/.agents/skills/augustus/references/judgment-class.md +++ b/.agents/skills/augustus/references/judgment-class.md @@ -36,7 +36,7 @@ taxonomy with enough of *your* data (XGBoost still wins there — | Family | What it optimizes | Typical output | Use when | Watch | |---|---|---|---|---| | **Closed decision API** (TypeSafe Jev) | Calibrated decision (proper-scoring / RLCD lineage) | Choice / Score / Noul + distributions | Default when you need act/abstain, fan-out, documented envelope | Cloud, pin version, re-measure on your data. AU health data-residency is a reason *not* to pick this family (`notes.md` §33) | -| **Open System-1 / decision-model head** (Laya, openjev, LightJev, openjev-lm, Nimble, Hume **Watch**) | Same *shape* as Jev, you host it | Same primitives or logits-as-options | Air-gap, $0/token, inspectable weights, deployment control | Self-eval duty; Laya text-only, 512 tok; vendor vs-Jev tables are claims (`notes.md` §18). A distill learns the *teacher's* answers: openjev-lm and jev-gate-student-b (`notes.md` §25, §33). Nimble is an open LoRA recipe on hard labels, not a Jev distill (`notes.md` §35). Hume's 27B dense drop is **Watch**, not a Hub checkpoint. He prefers the class name **decision models** over "system one" | +| **Open System-1 / decision-model head** (Laya, openjev, LightJev, openjev-lm, Nimble, **kev**, Hume **Watch**) | Same *shape* as Jev, you host it | Same primitives or logits-as-options | Air-gap, $0/token, inspectable weights, deployment control | Self-eval duty; Laya text-only, 512 tok; vendor vs-Jev tables are claims (`notes.md` §18). A distill learns the *teacher's* answers: openjev-lm and jev-gate-student-b (`notes.md` §25, §33). Nimble is an open LoRA recipe on hard labels, not a Jev distill (`notes.md` §35). **kev** is a shipped Qwen2.5-0.5B LoRA + pointer readout of Archer's reconstruction — public gold, not a Jev teacher; ID ECE only (`notes.md` §45). Hume's 27B dense drop is **Watch**, not a Hub checkpoint. He prefers the class name **decision models** over "system one" | | **Encoder open-jev** (DeBERTa-v3-large) | Same *shape*, bidirectional encoder, public gold (not a Jev teacher) | Choice / Score / Noul from one pass | Self-host decide without a decoder; 512 tok | In-domain ECE 0.022; OOD acc 0.854→0.690. English / three public domains. `notes.md` §33 | | **Constrained-AR surface** (TypeAR, **pcdServer**; not a species) | Next-token constraint on a pretrained generator | Distribution over allowed values | Typed fields without retraining; later fields must see earlier answers; local GGUF serving | Different objective from a proper-scoring head. TypeAR README enums ≤16; pcdServer 2–256 / 1–63 parallel fields. No abstention primitive. Compute-graph card below (`notes.md` §31, §32, §42). Public logit dump: mini-jev-runs | | **GLi\* encoder family** (GLiNER locate / GLiClass categorize / GLiNER2.5 local multi-head / GLiGuard safety schema) | One-pass labels-in-encoder; spans, sequence labels, a safety schema, or both | Spans + types; per-label sigmoid/softmax; optional relations/records | Laptop/local; large or changing label sets; "what's *in* the text" vs "what *is* the text" vs "which safety labels fire" | Affinities are not automatically a gateable P(permit). GLiGuard is not a Jev weight clone. Species map below. Not a Jev how-to and not a GLiNER or GLiGuard install | @@ -54,7 +54,7 @@ code owns side effects." They are not aliases. ```text locate GLiNER (span NER) what's *in* the text categorize GLiClass / GLiGuard what the text is; which schema labels fire -decide Jev / Laya / openjev Choice / Score / Noul over a state +decide Jev / Laya / openjev / kev Choice / Score / Noul over a state rank listwise / cross-encoder order a retrieved shortlist perceive CLIP / SigLIP / region Choice score candidates you extracted ``` @@ -85,9 +85,11 @@ below, next to the when-to-use table. objective. That is Jev's product claim. Open heads copy the *shape*; distillation copies the *teacher* (openjev-lm, jev-gate-student-b). Encoder open-jev (DeBERTa) copies the shape on public gold and still - owes OOD self-eval. A constrained autoregressive decode can emit a - label and still not be this species — compute-graph card below. - Hume's announced 27B dense drop is Watch. + owes OOD self-eval. **kev** copies the Archer readout (block-causal + isolation, pointer head, CE vs labelled public gold) and still owes + *your* ECE — ID numbers are not OOD. A constrained autoregressive + decode can emit a label and still not be this species — compute-graph + card below. Hume's announced 27B dense drop is Watch. - **Categorize (safety schema).** [GLiGuard](https://github.com/fastino-ai/GLiGuard) ([arXiv 2605.07982](https://arxiv.org/abs/2605.07982); Zaratiana, Newhauser, Hurn-Maloney, Lewis, Fastino): a GLiNER2 encoder that @@ -245,13 +247,14 @@ capability shift, independent of vendor: 5. **Open heads and GLi\* make the control plane local.** Air-gap / on-device / laptop (GLiNER2.5 74M–287M CPU-first; openjev-lm 0.5B LoRA overnight on 6 vCPU; encoder open-jev DeBERTa-v3-large 434M; - jev-gate-student-b 0.5B LoRA memory gate) become newly feasible *if* + jev-gate-student-b 0.5B LoRA memory gate; **kev** 0.5B LoRA + pointer, + ~160 ms / 6 questions on an M5) become newly feasible *if* you accept self-eval and envelope limits. They do not make calibration optional. Distilling a hosted teacher is not independent gold. Hume's 27B dense drop is the large-local Watch, not a third how-to. -6. **Cross-modal is still thin.** Discourse, GLiNER/GLiClass, Laya, and - the encoder open-jev are text-first. Vision is a scoring pattern +6. **Cross-modal is still thin.** Discourse, GLiNER/GLiClass, Laya, kev, + and the encoder open-jev are text-first. Vision is a scoring pattern (above), not a shipped omni decision API. Hume reports that a multimodal *base* plus text post-training generalizes to images with little intentional multimodal training — a Watch claim, not a @@ -376,14 +379,16 @@ program. | Need | Place | Do not | |---|---|---| | Calibrated p(y\|x) over a closed set | Trained decision-only head (Jev, or an open head you have proper-scored and measured on your labels) | Threshold a generated "90%", an affinity you have not calibrated, TypeAR constrained scores, or a LoRA student's agreement with the teacher | +| Laptop-local System One API for development / eval | **kev** — trained decision-only readout; official SDK with a `base_url` change (`notes.md` §45) | Treat 0.5B ID ECE as a knowledge or frontier substitute, or as OOD calibration | | Dependent sequential decisions | Constrained AR that conditions later steps on earlier answers (TypeAR sequential), or code-owned transitions and a new request per stage | Treat sibling questions on one request as if they attend each other | -| Open multimodal self-host / data-residency | Hume's announced **decision-model** drop **when it ships** (Qwen3.8 27B **dense**, 265k, multimodal, no audio; one forward pass locally once AR is removed; MoE next then shrink). Driver: healthcare AU residency, not anti-TypeSafe | Ship on "smarter than Jev." That is his early claim, against his own order-sensitivity and in-distribution calibration warnings. **WATCH** — no Hub weights this pass. Laya remains text-only. jev-visual is region Choice, not this drop | +| Open multimodal self-host / data-residency | Hume's announced **decision-model** drop **when it ships** (Qwen3.8 27B **dense**, 265k, multimodal, no audio; one forward pass locally once AR is removed; MoE next then shrink). Driver: healthcare AU residency, not anti-TypeSafe | Ship on "smarter than Jev." That is his early claim, against his own order-sensitivity and in-distribution calibration warnings. **WATCH** — no Hub weights this pass. Laya remains text-only. kev is text-only. jev-visual is region Choice, not this drop | | Image-in now, different graph | Diffusion structured reads that already accept images on a Jev-shaped interface ([djev-spark](https://github.com/mmastrac/djev-spark)) | Wait on the row above for image-in, or treat this graph as a proof it beats a decision head | ### When to use which decision surface -Five *surfaces*, not five species — plus an open recipe and a diffusion -graph that are not extra species either. Pick from the hole and these +Five *surfaces*, not five species — plus open recipes (Nimble; **kev** +as the runnable Archer reconstruction) and a diffusion graph that are +not extra species either. Pick from the hole and these axes; do not start from a logo. Hume prefers the class name **decision models** over "system one" ([tweet](https://x.com/4rcherhume/status/2100604161821979134)). This @@ -400,23 +405,28 @@ is the generator, not a sixth surface. | **Encoder open-jev** (DeBERTa-v3-large 434M) | Public gold, CE+Brier, val temperature. In-domain ECE 0.022 / acc 0.854; OOD acc 0.690 / ECE 0.035. **Not** a Jev teacher-copy | One pass over state + all questions; 512 tok | Author: 28 ms / 10 questions H100; 1.8 s / 4q M1 Max CPU | apache-2.0, self-host | Text | Jev-shaped 255 / Score 2–10 / Noul; 512 ctx | | **Tiny LoRA distill** (jev-gate-student-b) | Teacher-copy. P(relevant) from yes/no logits. Held-out n=60 vs vanilla 0.5B; 148,160-row corpus | Memory-gating / context sieve; **fail-open** on errors | Qwen2.5-0.5B LoRA; ~59 ms RTX 3060 | Local, apache-2.0 | Text | Binary relevance | | **Nimble** (open LoRA recipe, not a distill) | Hard synthetic labels. They say temperature was not tuned to correctness rates. 324-row agreement is their receipt, not an ECE (`notes.md` §35) | Not a gather primitive | Their latency table, not re-run | Self-host the adapter. Model card Apache-2.0; repo license absent | Text only | Enum ≤26; 2,048 tokens | +| **kev** (Qwen2.5-0.5B LoRA + pointer; Apache-2.0) | Public gold, CE. Held-out ECE 0.065 (0.031 after T=1.47); acc 0.799 on 1,350 ID questions. Isolation exact. **Not** a Jev teacher-copy (`notes.md` §45) | Laptop-local System One drop-in for development/eval; independent questions, one prefill | ~160 ms / 6 questions; ~1h45m train on M5; 38 MB adapter | Self-host; official `typesafe-sdk` with `base_url` | Text. Not multimodal. 0.5B knowledge | noul / choice 2–255 / score | | **Diffusion structured reads** (djev-spark) | Interface claim only. **Hypothesis** it beats a decision head on your labels (`notes.md` §36) | Optional sequential chunks, text-only | Their GX10 tables, not a class benchmark | DGX Spark container. Do not copy the route | Images are an extension; think and sequential reject images | README criteria, not copied here | **Three open paths** (not three species, not extra when-to-use rows): encoder open-jev (DeBERTa, public gold); AR constrained decode (TypeAR Python/SGLang, pcdServer native GGUF); trained decision-only (Laya / -Nimble / Archer **Watch**). Encoder vs decoder **replicas** of that -third path: DeBERTa is public gold with an OOD drop; openjev-lm / -jev-gate LoRAs are teacher-copies with named receipts (overnight -6-vCPU, $0/call — economics, not a new species; `notes.md` §25, §44). -Pick from the hole. A constrained softmax +Nimble / **kev** / Archer **Watch**). **kev** is the cleanest *runnable* +productization of Archer's reconstruction on that third path +(API-compatible; `notes.md` §45). Watch stays Watch. Encoder vs decoder +**replicas** of that third path: DeBERTa is public gold with an OOD +drop; openjev-lm / jev-gate LoRAs are teacher-copies with named receipts +(overnight 6-vCPU, $0/call — economics, not a new species; `notes.md` +§25, §44). kev is public gold on the same 0.5B backbone as openjev-lm, +not a teacher-copy. Pick from the hole. A constrained softmax is still not a Noul. Laya companion packaging this hour: [`laya-typed-decisions`](https://huggingface.co/convaiinnovations/laya-typed-decisions) (same 421.3M; acc 0.766 / Brier 0.066 on `LocalLLaMA/typed-decisions`, unverified — do not overwrite `notes.md` §18). Reject: TypeAR or pcdServer scores as fail-closed P(permit); a LoRA student's -agreement with Jev as independent gold; shipping on "smarter than Jev"; +agreement with Jev as independent gold; kev's ID ECE as a license to skip +a held-out test on *your* workflow; shipping on "smarter than Jev"; thresholding [`jp-sns-jev7-estimator`](https://huggingface.co/kokuren/jp-sns-jev7-estimator) teacher scores as P(toxic) — the card says they are **not** calibrated, and `threat` F1@0.5 is 0.0000 on their table (`notes.md` §33). Domain-local @@ -425,7 +435,7 @@ universal ranking. Diffusion beating a decision head is Hypothesis. Detail: `research/notes.md` §33 (surfaces), §34 (marginals), §35 (Nimble), §36 (diffusion), §38 (entropy allocator, Hypothesis), -§42 (pcdServer serving, meta-VOI, games). Before +§42 (pcdServer serving, meta-VOI, games), §45 (kev). Before adopting a surface, the bake-off is a jevals-shaped suite and, for a product loop, a Harbor taskset (`validation.md`, Eval & hill-climb). The stage pipeline into that decision is the same file @@ -503,7 +513,7 @@ The GitHub repo has no license file — do not call the repo Apache-2.0. (302/324), untuned Qwen3.8-27B 84.88% (275/324), base Qwen3.5-9B 66.36% (215/324). **Empirical** as that named receipt, not a ranking. Synthetic labels, six source families, 162 pairs. -- **Bake-off candidate** beside Laya, openjev-lm, and TypeAR. +- **Bake-off candidate** beside Laya, openjev-lm, TypeAR, and kev. Adoption still requires the eval path (`validation.md`, Eval & hill-climb). Archer stays Watch. Not a jevals how-to. @@ -515,6 +525,55 @@ the standing Choice-conditional-on-offered-set boundary (`SKILL.md`). Add "no match" when coverage is open. A high probability is not a correctness guarantee (their sentence). +### kev — runnable Archer reconstruction (not a distill) + +[jaredpalmer/kev](https://github.com/jaredpalmer/kev) (Apache-2.0; +README, MODEL_CARD, LICENSE, and release +[v0.1.0](https://github.com/jaredpalmer/kev/releases/tag/v0.1.0) HTTP +200). LoRA + pointer readout on Qwen2.5-0.5B. Shared state, isolated +questions under a block-causal mask, one prefill, no decode. +Architecture follows [Archer Hume's reconstruction](https://archerhume.com/posts/jevs-architecture-unmasked). +Speaks TypeSafe `POST /v1/systemone`; official `typesafe-sdk` works +with a `base_url` change. Weights `kev-0.5b` (38 MB) on that release. +Not a how-to: do not copy serve flags, ports, or train commands. + +**Place it on the trained decision-only open path** next to Laya / +Nimble / Archer Watch. It is the cleanest *runnable* productization of +that reconstruction (API-compatible). Watch stays Watch: 27B, +multimodal, no Hub weights this pass. Not `Kevthetech143/super-jev`. + +**Contrast** (same I/O shape, different graph / objective / duty): + +- **vs TypeAR / pcdServer.** Constrained AR decode; softmax over + allowed tokens is not a Noul. kev does not decode. Use TypeAR when a + later field must see an earlier answer. +- **vs encoder open-jev (DeBERTa).** Bidirectional encoder, public + gold, OOD drop measured (acc 0.854→0.690). kev is a causal decoder + + pointer; OOD is unmeasured for this checkpoint. +- **vs proprietary Jev.** Documented cloud decision API, ~32k envelope, + in-dist ECE 0.0313 with OOD collapse. kev is laptop-local, 0.5B + knowledge, ID calibration only. +- **vs openjev-lm / jev-gate.** Those LoRAs copy a Jev *teacher*. kev + trains CE on public labelled outcomes. Same 0.5B backbone, different + gold. +- **vs Nimble.** 9B contrastive hard labels, no measured ECE, enum + ≤26. kev publishes ECE and isolation probes. + +**Evidence (README; not re-run).** Isolation exact: packed vs separate +max Δ 3.7e-6; secret-in-sibling p=0.03 vs in-state 0.99. Held-out ECE +0.065 (0.031 after temperature scaling); overall acc 0.799 on 1,350 ID +questions. Permute argmax flips 7.4%; IIA log-odds shift mean 0.13; +boundary forgery held. Honest limits: 0.5B knowledge; ID calibration +only; not multimodal. Model card: research prototype, not production, +not Jev. + +**When to use.** Laptop-local System One API drop-in for development +and eval. Not a knowledge or frontier substitute. **Bake-off +candidate** on the jevals/Harbor path (`validation.md`); mechanism +tests (isolation, permute, IIA, boundary forgery) mirror Archer probes +— they falsify the reconstruction, they do not prove kev = Jev. +`notes.md` §45. + ### Constrained-AR surface (not `decide`) [TypeAR](https://github.com/zmtomorrow/TypeAR) ("Type-Safe Decoding for diff --git a/.agents/skills/augustus/references/mental-models.md b/.agents/skills/augustus/references/mental-models.md index d2a6996..5a70e2e 100644 --- a/.agents/skills/augustus/references/mental-models.md +++ b/.agents/skills/augustus/references/mental-models.md @@ -117,7 +117,8 @@ prefers the class name **decision models** over "system one" (`notes.md` §33); this file still says System One when quoting TypeSafe. Three open paths, not three species: encoder open-jev, AR constrained decode (TypeAR + pcdServer), trained decision-only (Laya / Nimble / -Archer Watch). A constrained softmax is still not a Noul (`notes.md` §42). +kev / Archer Watch). kev is the runnable Archer reconstruction on that +third path; Watch stays Watch. A constrained softmax is still not a Noul (`notes.md` §42, §45). **Readout versus a token; IIA is a property.** A direct probability and a generated "91%" are different objects; the format calibrates neither diff --git a/.agents/skills/augustus/references/mixed-architecture.md b/.agents/skills/augustus/references/mixed-architecture.md index 6321a6b..75cf209 100644 --- a/.agents/skills/augustus/references/mixed-architecture.md +++ b/.agents/skills/augustus/references/mixed-architecture.md @@ -372,14 +372,17 @@ Reproduce/open heads (`rongxinzy/LightJev`, openjev family, [`convaiinnovations/laya`](https://huggingface.co/convaiinnovations/laya), encoder [`open-jev-deberta-v3-large`](https://huggingface.co/com-kotobalabs/open-jev-deberta-v3-large), LoRA [`jev-gate-student-b`](https://huggingface.co/SargeDev/jev-gate-student-b), -companion packaging [`laya-typed-decisions`](https://huggingface.co/convaiinnovations/laya-typed-decisions)) +companion packaging [`laya-typed-decisions`](https://huggingface.co/convaiinnovations/laya-typed-decisions), +[`jaredpalmer/kev`](https://github.com/jaredpalmer/kev)) are evidence that the *interface* (Choice/Score/Noul, or yes/no logits as P(relevant)) is the transferable part — not a request to implement a backbone or a second API skill. Laya: self-hostable, text-only, 512 tokens/question; vendor benches vs Jev are **claims**. Encoder open-jev: -public gold, OOD drop. LoRA student: teacher-copy. Hume's 27B +public gold, OOD drop. LoRA student: teacher-copy. **kev**: public gold, +pointer readout, System One API drop-in; ID ECE only; not a teacher-copy +(`notes.md` §45). Hume's 27B decision-model drop is **Watch**. Closed calibrated API vs open weights -is a self-eval tradeoff (`research/notes.md` §18, §33). When-to-use +is a self-eval tradeoff (`research/notes.md` §18, §33, §45). When-to-use axes: `judgment-class.md`. TypeSafe remains the documented *exemplar*, not the class monopoly. GLiNER (locate) / GLiClass (categorize) / GLiNER2.5 (local multi-head), listwise, and vision families: diff --git a/.agents/skills/augustus/references/optimizer-integration.md b/.agents/skills/augustus/references/optimizer-integration.md index ebf3b7f..4185274 100644 --- a/.agents/skills/augustus/references/optimizer-integration.md +++ b/.agents/skills/augustus/references/optimizer-integration.md @@ -37,7 +37,9 @@ prompt loop: schema, criteria, and policy thresholds, scored on labeled eval (jevals) — not a search over a decoder. Open recipes such as Nimble: climb data curation and LoRA, measured on holdout ECE and agreement. Nimble's published holdout is agreement on synthetic -labels, not a measured ECE (`judgment-class.md`). +labels, not a measured ECE (`judgment-class.md`). kev: climb LoRA / +public-gold labels; the published ID ECE is a receipt, not your +workflow (`notes.md` §45). No call shape in this paragraph. The adapter notes below stay names of seats, not a request you copy. diff --git a/.agents/skills/augustus/references/validation.md b/.agents/skills/augustus/references/validation.md index 3ad635f..af12e0a 100644 --- a/.agents/skills/augustus/references/validation.md +++ b/.agents/skills/augustus/references/validation.md @@ -33,6 +33,11 @@ rejected two correct mates (`notes.md` §42). 149-row cousin: [`typesafe-jev-tools`](https://github.com/wotai-dev/typesafe-jev-tools) — Jev confidence monotonic vs Haiku invert in 0.80–0.95; do not copy the hook. +Open reconstruction cousin: [`jaredpalmer/kev`](https://github.com/jaredpalmer/kev) +— isolation packed vs separate max Δ 3.7e-6; secret-in-sibling p=0.03 +vs in-state 0.99; permute argmax flips 7.4%; IIA log-odds shift mean +0.13; boundary forgery held. Those tests mirror Archer probes; they do +not prove kev = Jev (`notes.md` §45). ## Offline eval: selective binary decisions @@ -52,9 +57,12 @@ the offline Brier / reliability / cost evaluator. Pointer only the Harbor substrate, and the one composition table are **Eval & hill-climb** below — do not restate them here. A bake-off candidate on that same labeled-case surface, beside Laya, openjev-lm, -and TypeAR, is [Bespoke Nimble](https://github.com/bespokelabsai/nimble) +TypeAR, and [kev](https://github.com/jaredpalmer/kev), is +[Bespoke Nimble](https://github.com/bespokelabsai/nimble) — an open LoRA recipe, not a Jev distill; their 324-example holdout is -a named receipt, not a ranking (`research/notes.md` §35). Same +a named receipt, not a ranking (`research/notes.md` §35). kev is the +runnable Archer-reconstruction candidate on the same surface (public +gold, measured ID ECE, not a teacher-copy; `notes.md` §45). Same acceptance-test *surface*, different UI: [jeiel85/jevscope](https://github.com/jeiel85/jevscope) (local-first visual debugger + JSONL regression; policy buckets are JevScope-derived, @@ -304,15 +312,18 @@ row is enough. ECE above is wanted, not a Nimble result. ### Bake-off mandate -Before adopting proprietary Jev vs Laya vs TypeAR vs Nimble vs Archer +Before adopting proprietary Jev vs Laya vs TypeAR vs Nimble vs kev vs Archer vs openjev-lm, run a jevals-shaped labeled suite (or an equivalent with this hygiene) and, for a product loop, a Harbor taskset. A design card with no eval path is incomplete. Archer weights are still a **Watch** — not on the Hub as of 2026-09-18 -(`notes.md` §31–§33). That bake-off is future, not Empirical. "A 9B +(`notes.md` §31–§33). That bake-off is future, not Empirical. kev is +the shipped 0.5B reconstruction on the trained decision-only path, not +that drop (`notes.md` §45). "A 9B LoRA is enough versus Jev" stays **Hypothesis** (`notes.md` §35). -openjev-lm is the name of that distill. Nimble is not a Jev distill +openjev-lm is the name of that distill. kev is not a Jev distill. +Nimble is not a Jev distill (model card Apache-2.0; GitHub LICENSE was 404). Meijer: marginals, not a PPL, not Kleisli (`notes.md` §34). djev-spark is a third compute graph, not the winner of this bake-off (`notes.md` §36). Do not diff --git a/CHANGELOG.md b/CHANGELOG.md index 486b417..21142f1 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -113,6 +113,17 @@ Pegged against [`typesafe-ai/skills` v0.5.7](https://github.com/typesafe-ai/skil OpenClaw typed routing ([jev-claw](https://github.com/trietphan/jev-claw)). Voice-control and JevML are README stubs. Higgsfield auto-routing is a claim. No wrapper. +- kev (`jaredpalmer/kev`, `research/notes.md` §45): runnable Archer + reconstruction on the trained decision-only open path next to Laya / + Nimble / Watch. Qwen2.5-0.5B LoRA + pointer, Apache-2.0, `POST + /v1/systemone` drop-in. Isolation exact (packed vs separate max Δ + 3.7e-6; secret-in-sibling p=0.03 vs in-state 0.99). Held-out ECE + 0.065 (0.031 after temp scale); acc 0.799 on 1,350 ID questions. + Permute argmax flips 7.4%; IIA log-odds shift mean 0.13; boundary + forgery held. Laptop-local System One for development/eval; not a + knowledge/frontier substitute; not a Jev teacher-copy. Contrast vs + TypeAR, encoder DeBERTa, proprietary Jev. jevals/Harbor bake-off + candidate. No serve how-to. - Effect-oriented loops (`notes.md` §28, `mappings.md` §19): Ward's ZIO client keeps Jev as the outer Choice and the handler as the effect. Not Effect.ts. GLiNER author: GLiNER2 "like jev" is GLiGuard diff --git a/README.md b/README.md index 6ba4742..1bc5b14 100644 --- a/README.md +++ b/README.md @@ -29,7 +29,7 @@ never launder a Noul as a proof. frames (EU, abstention, VOI, MCDA, SDT, search/control, Leveson, NATM/snap-fit/Norman); not SWE-only - `.agents/skills/augustus/references/judgment-class.md` — the class (Jev - exemplar, not monopoly): open heads (Laya, encoder DeBERTa, LoRA + exemplar, not monopoly): open heads (Laya, kev, encoder DeBERTa, LoRA distill), constrained-AR (TypeAR, pcdServer), announced decision-model (Watch), GLiNER/GLiClass species (locate vs categorize vs local multi-head), listwise vs decision objectives, vision scoring, when-to-use axes, @@ -49,7 +49,7 @@ never launder a Noul as a proof. - `.agents/skills/augustus/references/applied-mappings.md` — context sieve, exact-text keep/drop, env triage, moderation/ranking, skill routing - `.agents/skills/augustus/references/faq.md` — "just classification", - stack replacement, Jev vs open head vs encoder vs LoRA vs constrained AR, + stack replacement, Jev vs open head vs encoder vs LoRA vs constrained AR vs kev, GLiNER vs GLiClass vs CLIP, LLM-as-judge, in-engine vs CLI store, hard envelope (bitrate / planner), not-another-how-to - `.agents/skills/augustus/references/mappings.md` — classical-method diff --git a/docs/ecosystem.md b/docs/ecosystem.md index 98038b9..f004a0d 100644 --- a/docs/ecosystem.md +++ b/docs/ecosystem.md @@ -50,12 +50,13 @@ Per-keystroke launchers (104ms median, sequence-tagged staleness), firehose mode - **GLiGuard** (fastino-ai) — 0.3B GLiNER2 encoder, checkpoint `fastino/gliguard-LLMGuardrails-300M`. One bidirectional pass over a safety schema. Same interface shape as batched questions; different objective. Not a Jev weight clone. `judgment-class.md`; `notes.md` §30. - **DECRUX9812/openjev-lm** — Qwen2.5-0.5B+LoRA distilled from hosted Jev answers; 65/70 = 92.9% on 70 hand-labelled rows (one annotator, one domain, one seed) overnight on 6 vCPU, $0/call. Its 98.1% on fresh rows is teacher *agreement*, not gold. Receipts pattern: `notes.md` §25, §44. - **convaiinnovations/laya** — open Choice/Score/Noul head, text-only, 512 tok. Companion packaging this hour: [`laya-typed-decisions`](https://huggingface.co/convaiinnovations/laya-typed-decisions) (421.3M, acc 0.766 / Brier 0.066 unverified). `notes.md` §18, §42. +- **jaredpalmer/kev** — Qwen2.5-0.5B LoRA + pointer readout; Apache-2.0; `kev-0.5b` 38 MB on GitHub release v0.1.0. Runnable Archer reconstruction (`POST /v1/systemone`). Public gold, not a Jev teacher. Isolation exact; ID ECE 0.065 (0.031 after T); acc 0.799 / 1,350. Not a knowledge/frontier substitute. `notes.md` §45. - **zmtomorrow/TypeAR** — constrained autoregressive decoding surface: typed fields on a pretrained open model, no retraining. Not a proper-scoring decision head. `research/notes.md` §32. - **stephanj/pcdServer** — native Parallel Constrained Decoder (C++20, llama.cpp GGUF, Apple+Linux). TypeAR-class serving: 2–256 enums, 1–63 parallel fields; softmax over allowed values is not a Noul. `notes.md` §42. - **com-kotobalabs/open-jev-deberta-v3-large** — encoder open-jev, DeBERTa-v3-large 434M, apache-2.0, public gold (not a Jev teacher). In-domain ECE 0.022; OOD acc 0.854→0.690. `notes.md` §33. - **Mikhail/mini-jev-runs** — 27.9k schema-driven decisions; one forward pass; answer from next-token logits; no token generated. Calibration / constrained-decoding gold. `notes.md` §33. - **kokuren/jp-sns-jev7-estimator** — JP SNS seven-axis ONNX distill; teacher scores, not calibrated probabilities; `threat` F1@0.5 = 0. Domain-local categorize. -- **Archer Hume open decision-model** — **Watch.** Qwen3.8 27B dense, 265k, multimodal no audio; one forward pass locally once AR is removed. Driver: AU healthcare data-residency. No Hub weights this pass. `notes.md` §31–§33. +- **Archer Hume open decision-model** — **Watch.** Qwen3.8 27B dense, 265k, multimodal no audio; one forward pass locally once AR is removed. Driver: AU healthcare data-residency. No Hub weights this pass. `notes.md` §31–§33. Runnable *architecture* productization (0.5B, not that drop): **jaredpalmer/kev**. `notes.md` §45. - **bespokelabsai/nimble** — open recipe: contrastive hard labels, not a Jev distill. Model card Apache-2.0 LoRA `bespokelabs/Bespoke-Nimble-9B` on Qwen3.5-9B (repo license absent). Their 324-row holdout is a named receipt, not a ranking. `research/notes.md` §35. - **mmastrac/djev-spark** — DiffusionGemma 26B-A4B NVFP4, Jev-shaped decisions, images as an extension. Third compute graph. Interface claim, not a win over a decision head. `research/notes.md` §36. - **Perception then judgment** — SAM 3.1 (masks and tracks) or ASR (a transcript) are upstream producers, not the perceive species. System One on that state is decide. Composition, not native omni. Information dies at the interface. `research/notes.md` §39. diff --git a/research/archive/findings.md b/research/archive/findings.md index 3dd3f48..d154caf 100644 --- a/research/archive/findings.md +++ b/research/archive/findings.md @@ -680,6 +680,31 @@ in-engine vs CLI fork; (at) soft judgment inside a hard envelope (ABR, planner); (au) distill-to-device as a context sieve, not only an action gate. +## Batch #28 (2026-09-18) — kev runnable Archer reconstruction + +Note: `research/notes.md` §45. Docs-only. Folded into PR #2, not a +second PR. Archer 27B drop still Watch. + +- **kev (Empirical as named ID receipt; Contract as README/API).** + `jaredpalmer/kev`, Apache-2.0, 24★ this pass. Qwen2.5-0.5B LoRA + + pointer; `POST /v1/systemone`; typesafe-sdk `base_url`. Public gold, + not a Jev teacher. Isolation exact (Δ 3.7e-6; sibling p=0.03 vs + state 0.99). ECE 0.065 / 0.031 after T; acc 0.799 / 1,350 ID. + Permute 7.4%; IIA mean 0.13; boundary forgery held. 0.5B knowledge; + ID calibration only; not multimodal. +- **Place:** trained decision-only open path next to Laya / Nimble / + Watch. Cleanest *runnable* productization of Archer's reconstruction. +- **Contrast:** TypeAR (constrained AR ≠ Noul) vs encoder DeBERTa + (OOD measured) vs proprietary Jev vs openjev-lm (teacher-copy). +- **Eval:** jevals/Harbor bake-off candidate; mechanism tests mirror + Archer probes. +- **When-to-use:** laptop-local System One for development/eval; not a + knowledge/frontier substitute. No serve how-to. + +Cross-repo addition: (av) the trained decision-only path now has a +shipped API-compatible reconstruction (kev); Watch remains the 27B +announcement. + diff --git a/research/notes.md b/research/notes.md index 0957583..519f0ff 100644 --- a/research/notes.md +++ b/research/notes.md @@ -2316,3 +2316,84 @@ gallery + host-adapter note; `applied-mappings.md` §1 / §5; `validation.md` harness practice; `mental-models.md` envelope + planner; FAQ in-engine vs CLI; `judgment-class.md` encoder vs decoder replica (no new species). No wrapper. + +## 45. kev — runnable Archer reconstruction, not a distill (2026-09-18) + +HTTP 200 this pass: +[jaredpalmer/kev](https://github.com/jaredpalmer/kev) README (raw +`main`), [MODEL_CARD.md](https://github.com/jaredpalmer/kev/blob/main/MODEL_CARD.md), +[LICENSE](https://github.com/jaredpalmer/kev/blob/main/LICENSE) +(Apache-2.0; GitHub API `license.key` apache-2.0), +[release v0.1.0](https://github.com/jaredpalmer/kev/releases/tag/v0.1.0), +[Archer Hume, *Jev's Architecture Unmasked*](https://archerhume.com/posts/jevs-architecture-unmasked) +(with and without trailing slash), TypeSafe +[System One API](https://docs.typesafe.ai/api). + +GitHub API this pass: created 2026-09-17T20:49:39Z; pushed +2026-09-18T19:50:58Z; 24★; language TypeScript (playground); default +branch `main`. Description: "tiny Jev-like model built on top of +Qwen2.5-0.5B you can train and run on your MacBook." Not +`Kevthetech143/super-jev`. + +**What it is (Contract as README + model card).** Jared Palmer, Apache-2.0 +adapter/head (Qwen2.5-0.5B under the Qwen license). LoRA (r=16) + a +pointer readout on that 0.5B causal backbone. Typed questions in, +calibrated probabilities out, one prefill pass, no decode. State and +questions packed into one sequence; a block-causal mask lets each +question see the document and never a sibling. The pointer head scores +each option against the question's `` token and softmaxes. +Trained with cross-entropy against labelled outcomes. Architecture +explicitly follows Hume's reconstruction (§31): shared state, isolated +questions, pointer head, CE vs labelled outcomes. API follows TypeSafe +`POST /v1/systemone`; official `typesafe-sdk` works with a `base_url` +change. Weights `kev-0.5b` (38 MB: LoRA adapter, readout head, +tokenizer) on GitHub release v0.1.0; base model from the Hub on first +load. Trains ~1h45m on an Apple M5; ~160 ms for a six-question +request. Model card: research prototype, not production, not Jev. +Do not copy serve flags, ports, or train commands into skill cards. + +**Not a distill.** Six public datasets converted to TypeSafe-shaped +requests (Banking77, AG News, MNLI, BoolQ, SST-5, Yelp): 9,000 records, +13,500 questions, two epochs. No LLM-generated data. Distinct from +openjev-lm / jev-gate on the *same* 0.5B backbone: those copy a hosted +Jev teacher. Distinct from Nimble: 9B contrastive synthetic labels, no +measured ECE. Distinct from encoder open-jev: bidirectional DeBERTa, +OOD drop measured. Distinct from TypeAR / pcdServer: those decode a +constrained next token. Distinct from proprietary Jev: closed weights, +~32k envelope. Distinct from Archer Watch: 27B announced, multimodal, +no Hub weights this pass. + +**Evidence (README / model card; not re-run; Empirical as their named +receipt).** Isolation exact: packed vs separate max Δ **3.7e-6**; +secret-in-sibling / absent / in-state **p = 0.03 / 0.03 / 0.99**. +Held-out ECE **0.065** (10 bins) on 1,350 ID questions; **0.031** after +one-parameter temperature scaling (T=1.47). Overall acc **0.799**. +Permute (4 orders, Choice K ≥ 3): argmax flips **7.4%**. IIA: log-odds +shift from one irrelevant option **mean 0.13**, p90 0.34. Boundary +forgery: option count unchanged, forged option p ≤ 0.09. Per-source +cells (acc / ECE) stay in the model card; do not promote them into a +ranking against Jev. + +**Honest limits (their words).** 0.5B knowledge (on the TypeSafe docs' +structured-criteria example kev picks `return_policy` where Jev picks +`return_status`). Calibration is in-distribution; ECE on the training +datasets says nothing about a new workflow. Not multimodal. Trained at +384 state / 1,024 branch tokens; serving caps at 8,192 vs Jev ~32k. +Score confidence is a stand-in; TypeSafe has not published theirs. +Choice `confidence` uses `(p_max − 1/K) / (1 − 1/K)` — the same +arithmetic Hume reconstructed in the official adapter (§31), as *their* +API derivation, not a TypeSafe contract. + +**Placement.** (a) trained decision-only open path next to Laya / Nimble +/ Archer Watch; (b) cleanest *runnable* productization of Archer's +reconstruction (API-compatible); (c) contrast vs TypeAR (constrained +AR decode) vs encoder open-jev vs proprietary Jev; (d) jevals/Harbor +bake-off candidate; mechanism tests mirror Archer probes — they +falsify the reconstruction, they do not prove kev = Jev; (e) when to +use: laptop-local System One API drop-in for development/eval; not a +knowledge/frontier substitute. **Empirical** as a public repo + named +ID receipt. **Hypothesis** that it substitutes for Jev on *your* +labels. Cards: `judgment-class.md`; FAQ; `validation.md`; +`mental-models.md`; `mixed-architecture.md`; `formal-methods.md` +(pointer-softmax is still a sensor); `optimizer-integration.md`. +No wrapper. diff --git a/research/refresh-log.md b/research/refresh-log.md index b95a391..77b1a3e 100644 --- a/research/refresh-log.md +++ b/research/refresh-log.md @@ -377,5 +377,19 @@ mini-jev-runs, jev-tree-choice-cap. - notes.md §44; sources.json; findings.md batch #27. No wrapper. +## 2026-09-18 19:54 UTC — kev (runnable Archer reconstruction) + +- Folded into open PR #2 (`cursor/augustus-store-envelope-00b4`); no + duplicate PR. Docs-only. Archer 27B drop still **WATCH**. +- Source: [`jaredpalmer/kev`](https://github.com/jaredpalmer/kev) + README + MODEL_CARD + LICENSE Apache-2.0 + release v0.1.0, all HTTP + 200. 24★. Isolation / ECE / acc / permute / IIA / forgery cited from + README; not re-run. Not a Jev distill (public gold CE). +- Cards: `judgment-class.md` (family, holes, when-to-use, dedicated + card); FAQ; `validation.md` bake-off + mechanism tests; + `mental-models.md`; `mixed-architecture.md`; `formal-methods.md`; + `optimizer-integration.md`; SKILL.md path + identity lock. +- notes.md §45; sources.json; findings.md batch #28. No wrapper. + diff --git a/research/sources.json b/research/sources.json index a2cfeb3..5b7397f 100644 --- a/research/sources.json +++ b/research/sources.json @@ -1,6 +1,6 @@ { "refresh_cadence": "hourly", - "retrieved": "2026-09-18T18:58Z", + "retrieved": "2026-09-18T19:54Z", "sources": [ { "kind": "docs", @@ -1231,6 +1231,24 @@ "title": "higgsfield_ai: GenAI model auto-routing", "url": "https://x.com/higgsfield_ai/status/2101022473248727177", "note": "2026-09-18T18:56:07Z. Claim, no labeled catalog. Same hole as routeKit. notes.md \u00a744." + }, + { + "kind": "github", + "title": "jaredpalmer/kev", + "url": "https://github.com/jaredpalmer/kev", + "note": "Runnable Archer reconstruction: Qwen2.5-0.5B LoRA + pointer, POST /v1/systemone, Apache-2.0. Public gold, not a Jev teacher. notes.md \u00a745." + }, + { + "kind": "github", + "title": "jaredpalmer/kev release v0.1.0 (kev-0.5b)", + "url": "https://github.com/jaredpalmer/kev/releases/tag/v0.1.0", + "note": "38 MB LoRA + pointer head weights. Held-out acc 0.799 / ECE 0.065. notes.md \u00a745." + }, + { + "kind": "github", + "title": "jaredpalmer/kev MODEL_CARD.md", + "url": "https://github.com/jaredpalmer/kev/blob/main/MODEL_CARD.md", + "note": "Research prototype; ID ECE 0.065 / 0.031 after T=1.47; isolation probes. notes.md \u00a745." } ] } From 7ce50433e49efc475f023defe7ae60c86ed48652 Mon Sep 17 00:00:00 2001 From: Cursor Agent Date: Fri, 18 Sep 2026 20:12:45 +0000 Subject: [PATCH 03/43] Fold 14:03 Boise watch: open multimodal RLCD, bake-off, decision-token LoRA Cite Hub/GitHub receipts only. Archer 27B remains Watch. Skill cards carry mental models (omni decide without waiting, ECE/NLL/Brier bake-off, decision-token training, allowlist-as-proof, forced Choice, S1 keeps control), not a hit list or serve how-to. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com> --- .agents/skills/augustus/SKILL.md | 23 +- .../references/agent-self-assessment.md | 6 +- .../augustus/references/applied-mappings.md | 5 + .agents/skills/augustus/references/faq.md | 50 +++- .../augustus/references/formal-methods.md | 4 +- .../augustus/references/judgment-class.md | 133 +++++++---- .../skills/augustus/references/mappings.md | 18 +- .../augustus/references/mental-models.md | 10 +- .../augustus/references/methods-catalog.md | 3 +- .../augustus/references/mixed-architecture.md | 27 ++- .../references/optimizer-integration.md | 4 +- .../augustus/references/question-design.md | 4 +- .../augustus/references/toolbox-mapping.md | 4 +- .../skills/augustus/references/validation.md | 41 +++- CHANGELOG.md | 17 ++ README.md | 8 +- docs/ecosystem.md | 15 +- research/archive/findings.md | 37 +++ research/notes.md | 221 ++++++++++++++++++ research/refresh-log.md | 20 ++ research/sources.json | 56 ++++- 21 files changed, 613 insertions(+), 93 deletions(-) diff --git a/.agents/skills/augustus/SKILL.md b/.agents/skills/augustus/SKILL.md index 0a73aab..1c059f1 100644 --- a/.agents/skills/augustus/SKILL.md +++ b/.agents/skills/augustus/SKILL.md @@ -1,6 +1,6 @@ --- name: augustus -description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill), announced open decision-model (Watch), constrained-AR (TypeAR, pcdServer), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"eval path\", \"jevals\", \"Harbor taskset\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge, TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, or \"formally verify with Jev\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." +description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill, open multimodal RLCD / blackwood), announced open decision-model (Watch — still not landed), constrained-AR (TypeAR, pcdServer, decision-token LoRA), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"wait for Archer vs open multimodal RLCD\", \"screenshot/DOM candidates → typed Choice\", \"eval path\", \"jevals\", \"Harbor taskset\", \"shared bake-off ECE/NLL/Brier\", \"LLM-as-judge is not the System One score\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge (allowlist *proves*; fail-open cannot block), \"missing other → confident wrong Choice\", \"lint the Jev request\", \"S1 reflex keeps control / optional S2 one-use advice\", TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, or \"formally verify with Jev\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." license: MIT metadata: version: 0.3.0 @@ -24,7 +24,7 @@ the official `typesafe-ai` skill plus the live docs own Jev integration contracts — read them before writing Jev API code. Neighbor skills `tenbin` (lint/measure) and `decision-first` (try-Jev-first habit) own their jobs. Do not collapse into a TypeSafe how-to, a Laya install, a kev -serve, or a GLiClass or GLiNER tutorial. +serve, a blackwood vLLM how-to, or a GLiClass or GLiNER tutorial. Pick the **pillar** from the hole (expected utility, VOI, MCDA, signal detection, search/control, org/safety, formal methods), then the @@ -52,8 +52,9 @@ classical method you already trust, substitute it, classify the win software?", "formally verify with Jev / replace TLA+ / Dafny / DST", "Alloy vs Apalache", "GLiNER vs Jev", "LLM-as-judge", "paraphrase brittleness", "allowlist then judge", "TOCTOU-of-Noul", - "Jev inside the database / sqlite-jev", or "Jev picks bitrate / join - order / the model": + "Jev inside the database / sqlite-jev", "Jev picks bitrate / join + order / the model", "wait for Archer", "lint the request / missing + other", or "screenshot Choice / omni System One": read `references/faq.md`, then `references/mental-models.md`, then `references/mixed-architecture.md`, then @@ -114,7 +115,7 @@ classical method you already trust, substitute it, classify the win |---|---|---| | Mental models across domains (not SWE-only) | EU, abstention, VOI, MCDA, SDT, search/control, Leveson, NATM/Norman/snap-fit | `references/mental-models.md` | | Judgment-model class (Jev is exemplar, not monopoly) | Species: decide / locate (GLiNER) / categorize (GLiClass) / rank / perceive; open heads include encoder DeBERTa, LoRA distill, openjev-lm, kev | `references/judgment-class.md` | -| Open weights vs constrained decoding vs proprietary API | Three open paths: encoder open-jev / AR constrained decode (TypeAR + pcdServer) / trained decision-only (Laya, Nimble, kev, Archer Watch). Softmax over allowed tokens ≠ Noul | `references/judgment-class.md` (when-to-use table) | +| Open weights vs constrained decoding vs proprietary API | Three open paths: encoder open-jev / AR constrained decode (TypeAR + pcdServer; decision-token LoRA) / trained decision-only (Laya, Nimble, kev, **blackwood-rlcd** multimodal now, Archer Watch still Watch). Softmax over allowed tokens ≠ Noul | `references/judgment-class.md` (when-to-use table) | | Entropy as allocator (low / medium / high) | Typed low+medium decisions → System One marginals; high-entropy synthesis → frontier decoder. Product rhetoric, not a meter. **Hypothesis** | `references/judgment-class.md` | | Formal / semi-formal (proof vs judgment) | Sensor vs constraint vs searchlight; Alloy vs Apalache; DST trio; TOCTOU-of-Noul, AI×FM | `references/formal-methods.md` (one-screen: `references/formal-semi-formal.md`) | | Mixed architecture (judgment model + LLM) | Provider judges, LLM writes, code owns control; not a stack replacement | `references/mixed-architecture.md` | @@ -145,16 +146,16 @@ classical method you already trust, substitute it, classify the win | Assignment hybrid | Soft affinity + hard solver | `references/mappings.md#15-assignment-hybrid--soft-affinity--hard-solver-hypothesis` (**Hypothesis**) | | Situated density (Shirky) | Aggressive soft loops only inside a named community | `references/mappings.md#16-situated-density-shirky-hypothesis` (**Hypothesis**) | | Input brittleness / paraphrase stability | Synonymous wording that swings p → abstain or rewrite | `references/mappings.md#17-input-brittleness--sensitivity-calibration-selective-abstention-hypothesis` (**Hypothesis**) | -| Structural prove ∩ soft remainder | Allowlist/text-layer/law first; judge only leftovers | `references/mappings.md#18-structural-prove--soft-remainder-hypothesis-as-domain-general-empirical-as-named-shapes` (**Hypothesis**; jevgate/OCR shapes Empirical) | +| Structural prove ∩ soft remainder | Allowlist *proves* the easy verbs; judge only unlisted leftovers; fail-open (cannot block) | `references/mappings.md#18-structural-prove--soft-remainder-hypothesis-as-domain-general-empirical-as-named-shapes` (**Hypothesis**; jevgate/OCR shapes Empirical) | | Effect-oriented state-machine loops | Soft predicates on transitions; code owns the transition | `references/mappings.md#19-effect-oriented-state-machine-loops-hypothesis` (**Hypothesis**; ZIO client, not Effect.ts) | -| Agent self-supervision / on-track detection | Pre-gate → output judge → done-check → supervisor nouls | `references/agent-self-assessment.md` | +| Agent self-supervision / on-track detection | Pre-gate → output judge → done-check → supervisor nouls. S1 reflex keeps control; optional S2 is one-use advice | `references/agent-self-assessment.md` | | Optimizer/program frameworks (Ax, DSPy) | Typed fields → one provider request; judge metrics; threshold discipline. Ax and DSPy climb LM-program knobs only | `references/optimizer-integration.md` | -| Perception → decision pipeline / measure / hill-climb | Stages with a versioned state contract; frozen taskset; DSPy/Ax only on the LM-program slice. **Hypothesis**. Same section as the row below | `references/validation.md#eval--hill-climb` | -| Eval & hill-climb | Decision-stage jevals hygiene; Harbor taskset × harness × runtime; one score-composition table | `references/validation.md#eval--hill-climb` | +| Perception → decision pipeline / measure / hill-climb | Stages with a versioned state contract; frozen taskset; DSPy/Ax only on the LM-program slice. Specialist composition stays **Hypothesis**; open multimodal decide (blackwood-rlcd) is a named receipt. Same section as the row below | `references/validation.md#eval--hill-climb` | +| Eval & hill-climb | Decision-stage jevals hygiene; Harbor taskset × harness × runtime; one score-composition table. Shared bake-off exemplar: open-jev-laya-bench (ECE/NLL/Brier; LLM-as-judge is not the score) | `references/validation.md#eval--hill-climb` | | (meta) Finding new mappings & applications | Toolbox sweep: judgment-shaped component of a known method, substituted + falsified | `references/toolbox-mapping.md` | | Named methods / operators / theorems | Substitution tiers: operand-judgments, preconditioned theorems, non-substitutable | `references/methods-catalog.md` | | (meta) Where a judgment model sits relative to any construct | 11 positions + logical-operator rules + position×construct traversal as the application generator | `references/composition-algebra.md` | -| Question mechanics & debugging | Instruction/criteria/state shape, budgets, diagnosis table, revision discipline | `references/question-design.md` | +| Question mechanics & debugging | Instruction/criteria/state shape, budgets, diagnosis table, revision discipline. Missing `other` → confident wrong Choice (confidence gating cannot catch); lint the request (`wellposed` recipe; `tenbin` owns the skill) | `references/question-design.md` | | Heuristic search over a taxonomy | Parallel beam over Choice distributions | `references/mappings.md#5-hierarchy--bounded-heuristic-search` | | Existing-system insertion / code-smell audit | Opportunity map, fit test, smallest boundary, policy centralization | `references/boundary-audit.md` | @@ -206,7 +207,7 @@ Desired behavior and non-judgment baseline: Semantic judgment(s) and what each output means: Pillar (EU / VOI / MCDA / SDT / search / safety / formal): Hole (sieve / keep-drop / triage / rank / route / gate / perceive / abstain / gather): -Family (closed decision API / open head / encoder open-jev / constrained-AR surface / GLiNER locate / GLiClass categorize / listwise ranker / vision scorer): +Family (closed decision API / open head / open multimodal RLCD / encoder open-jev / constrained-AR surface / GLiNER locate / GLiClass categorize / listwise ranker / vision scorer): Evidence/candidate source and known coverage gaps: Deterministic policy, constraints, and action ownership: Batchable vs genuinely dependent steps: diff --git a/.agents/skills/augustus/references/agent-self-assessment.md b/.agents/skills/augustus/references/agent-self-assessment.md index d36cbda..05f055d 100644 --- a/.agents/skills/augustus/references/agent-self-assessment.md +++ b/.agents/skills/augustus/references/agent-self-assessment.md @@ -30,7 +30,11 @@ relying: foreman, pi-jev, pi-warden, winnow, fast-jev-compaction, jev-judgment. `tests_sufficient`, `worker_stuck`, `work_off_track`, `ready_to_finish`; deterministic policy with hysteresis (retry counts, verification history) gates continue/stop/retry/verify. The model never - commands; it estimates named probabilities. + commands; it estimates named probabilities. Same split as + [jev-reflex-autonomy-lab](https://github.com/khordoo/jev-reflex-autonomy-lab): + **S1 keeps control**; optional S2 is one-use advice on low confidence + and does not fly the drone (`notes.md` §46). Experimental viz, not a + production supervisor. 6. **Context economy**: the context-sieve card (`references/applied-mappings.md#1-context-sieve`). Judge every large tool result with one relevance Noul before it enters context. Hide diff --git a/.agents/skills/augustus/references/applied-mappings.md b/.agents/skills/augustus/references/applied-mappings.md index 889fb82..52b0ec6 100644 --- a/.agents/skills/augustus/references/applied-mappings.md +++ b/.agents/skills/augustus/references/applied-mappings.md @@ -81,6 +81,11 @@ the subset operation you already had hunks unstaged; lines never split; atomic apply after confirm. Line-by-line search cookbook (**Contract**): score existing line ids, do not generate ids. lizard-agent: pick among visible elements; answers are *located*. +Omni cousin this hour: [blackwood-rlcd](https://huggingface.co/BlackwoodAI/blackwood-rlcd) +picks among **letters drawn on the screenshot**; code still clicks +(`notes.md` §46). Text-only cousin: [jev-e2e](https://github.com/perixtar/jev-e2e) +— Jev selects observed controls; Playwright independently checks; +a confident model cannot substitute for checked expectations. **Counterexample**: "write the patch that matches this sentence" — that is generation. **Test**: every kept byte occurs in the input; mixed never auto-included; snapshot stale → abort, don't guess. diff --git a/.agents/skills/augustus/references/faq.md b/.agents/skills/augustus/references/faq.md index e5689ba..f300df2 100644 --- a/.agents/skills/augustus/references/faq.md +++ b/.agents/skills/augustus/references/faq.md @@ -61,7 +61,7 @@ No. Augustus designs for the whole class of fast/cheap categorization-classification-scoring models. TypeSafe Jev is the documented exemplar (typed Choice / Score / Noul, live docs). Neighbors in the class — open System-1 / decision-model heads (Laya, kev, -openjev-lm, encoder DeBERTa, LoRA distill; Hume's 27B drop is Watch), +blackwood-rlcd, openjev-lm, encoder DeBERTa, LoRA distill; Hume's 27B drop is Watch), constrained-AR (TypeAR), GLiNER/GLiClass encoder family (locate vs categorize vs local multi-head), listwise/pairwise rankers, vision scorers — are substitutes @@ -98,10 +98,11 @@ Five surfaces, not one family (`judgment-class.md` when-to-use table). Three *open* paths sit beside proprietary Jev: **encoder** open-jev (DeBERTa, public gold), **AR constrained decode** (TypeAR; native [pcdServer](https://github.com/stephanj/pcdServer) GGUF serving), -**trained decision-only** (Laya / Nimble / **kev** / Archer Watch). Proprietary Jev is the documented decision API; you do not hold the +**trained decision-only** (Laya / Nimble / **kev** / **blackwood-rlcd** / +Archer Watch). Proprietary Jev is the documented decision API; you do not hold the weights, so checks around the boundary stay black-box (`formal-methods.md`). A trained decision-only open head (Laya, kev, -openjev-lm, encoder DeBERTa, a LoRA student) copies the Choice / Score / +openjev-lm, encoder DeBERTa, a LoRA student, **blackwood-rlcd**) copies the Choice / Score / Noul *shape* and moves eval onto you. Distills trained on Jev's *answers* (openjev-lm, jev-gate-student-b) are teacher-copies — read agreement separately from gold. **kev** is not that distill: CE on @@ -119,7 +120,11 @@ brittleness; compose with abstention and an allowlist gate (`mappings.md` §2, §17, §18). Hume's announced open **decision-model** (27B dense, multimodal, AU healthcare residency — not anti-TypeSafe) is **WATCH** until weights, license, and evals exist (`notes.md` §31, -§33). Constrained decoding is §32; native serving is §42. Public logit dump for the +§33). Open multimodal *decide* that already shipped: +[blackwood-rlcd](https://huggingface.co/BlackwoodAI/blackwood-rlcd) +(CC BY-NC; not that drop; `notes.md` §46). Constrained decoding is §32; native serving is §42. Decision-token QLoRA on that graph: +[Foodoo1/Qwen3-14B-RLCD-Decision-LoRA](https://huggingface.co/Foodoo1/Qwen3-14B-RLCD-Decision-LoRA) +(train the decision token, not prose; synthetic fraud receipt). Public logit dump for the read-the-letter graph: mini-jev-runs. "Smarter than Jev" is a claim. He prefers "decision models" over "system one"; this skill still quotes TypeSafe's name for the exemplar. Before you pick any of those paths, @@ -146,7 +151,8 @@ Hole first, logo last. These are **species**, not aliases independent gold. Encoder open-jev (DeBERTa) is the same *shape* on public gold — still self-eval, especially OOD. **kev** is the causal-decoder + pointer productization of Archer's reconstruction on - public gold (API-compatible; not a teacher-copy). Hume's 27B drop is + public gold (API-compatible; not a teacher-copy). **blackwood-rlcd** is + the open multimodal decide head (CC BY-NC; not Archer Watch). Hume's 27B drop is Watch. When-to-use axes: `judgment-class.md`. - **Cross-encoder / listwise ranker:** order of a retrieved shortlist. Fail **open** (keep retrieval order). Translation-invariant listwise @@ -270,7 +276,10 @@ When you need a paragraph rationale, a trace UI, or an annotation workflow, generation and the eval platform still own those seats. Verbal LLM scores are uncalibrated. The Harbor/jevals-adjacent practice is: shadow mode + fixtures that assert on the **action**, -not on prose (`jev-harness`, `validation.md`; `notes.md` §44). Do not +not on prose (`jev-harness`, `validation.md`; `notes.md` §44). Shared +bake-off this hour scores **ECE / NLL / Brier**, not an LLM paragraph +([open-jev-laya-bench](https://huggingface.co/datasets/pngwn/open-jev-laya-bench); +`notes.md` §46). Do not thin this skill into a Langfuse how-to. Mixed architecture: traces stay; the judge step can be a System One model. @@ -284,7 +293,22 @@ first so a comment can talk it into a write is the rejected design. The same three-way test, one hour later: if a regex, a DNS lookup, or a database query already answers, **do not call a model** (`wotai-dev/typesafe-jev-tools`, `notes.md` §42). That is meta-VOI, not -a hook tutorial. +a hook tutorial. This hour's wording of the same sandwich: the allowlist +**proves** read-only verbs; Jev judges only unlisted leftovers; +fail-open (cannot block) (`notes.md` §46). + +## Can confidence gating catch a forced wrong Choice? + +No. Choice probabilities are conditional on the offered set. If coverage +is open and you omit `other`, the model must pick a listed option — and +the distribution can peak at **1.00 on the wrong label**. Downstream +confidence gates see a healthy answer. Lint the *request* (missing +escape hatch, broken state paths) before you trust the number. Recipe: +[wellposed](https://github.com/suraj-phanindra/wellposed) (unsubscribe +email → `"support issue"` at 1.00 without `other`; overlapping options +collapse to 0.19 — that failure is loud). `tenbin` still owns the +design-time lint *skill*; Augustus owns the placement. +`question-design.md`; `notes.md` §46. ## Should Jev live inside the database? @@ -313,6 +337,18 @@ confident (+12% geomean author-reported; join-order Choice alone was code picks the generator. The envelope is load-bearing (`mappings.md` §12, §15, §18). +## Wait for Archer to ship omni System One? + +No. Archer's 27B dense drop is still **Watch** (no Hub weights this +pass; user watch ~2026-09-19). Omni perception→decision already has an +open model: [`BlackwoodAI/blackwood-rlcd`](https://huggingface.co/BlackwoodAI/blackwood-rlcd) +(CC BY-NC; Jev-compatible shim; screenshot + marked candidates → +Choice). Soft judgment over pixel candidates inside deterministic +code. Jev still leads general *text* (0.850 vs 0.786 on their 8,456-item +table). Specialist composition (SAM / OCR → text → Jev) remains valid. +Do not wait, and do not treat screenshot-vs-Jev-text as the same input. +`judgment-class.md`; `notes.md` §46. + ## Is confidence a trained score? No — not on Hume's reconstruction, and not as a new contract. Re-read diff --git a/.agents/skills/augustus/references/formal-methods.md b/.agents/skills/augustus/references/formal-methods.md index 2753eca..f520d28 100644 --- a/.agents/skills/augustus/references/formal-methods.md +++ b/.agents/skills/augustus/references/formal-methods.md @@ -207,7 +207,9 @@ A Noul is still not a proof that the property holds, and a clean PBT run is not one either. A perception-to-decision handoff is a contract surface — the schema of objects or utterances, not the pixels or the waveform: property-test that interface, and do not pretend the Noul is -over raw pixels or raw audio. Hill-climb of that handoff: +over raw pixels or raw audio. A shared multimodal *decide* head +(blackwood-rlcd) still judges **marked candidates**, not an open click; +the act stays in code (`notes.md` §46). Hill-climb of that handoff: `validation.md`. ## 4. Deterministic simulation testing (semi-formal trio) diff --git a/.agents/skills/augustus/references/judgment-class.md b/.agents/skills/augustus/references/judgment-class.md index 0fa7252..625c73c 100644 --- a/.agents/skills/augustus/references/judgment-class.md +++ b/.agents/skills/augustus/references/judgment-class.md @@ -36,9 +36,9 @@ taxonomy with enough of *your* data (XGBoost still wins there — | Family | What it optimizes | Typical output | Use when | Watch | |---|---|---|---|---| | **Closed decision API** (TypeSafe Jev) | Calibrated decision (proper-scoring / RLCD lineage) | Choice / Score / Noul + distributions | Default when you need act/abstain, fan-out, documented envelope | Cloud, pin version, re-measure on your data. AU health data-residency is a reason *not* to pick this family (`notes.md` §33) | -| **Open System-1 / decision-model head** (Laya, openjev, LightJev, openjev-lm, Nimble, **kev**, Hume **Watch**) | Same *shape* as Jev, you host it | Same primitives or logits-as-options | Air-gap, $0/token, inspectable weights, deployment control | Self-eval duty; Laya text-only, 512 tok; vendor vs-Jev tables are claims (`notes.md` §18). A distill learns the *teacher's* answers: openjev-lm and jev-gate-student-b (`notes.md` §25, §33). Nimble is an open LoRA recipe on hard labels, not a Jev distill (`notes.md` §35). **kev** is a shipped Qwen2.5-0.5B LoRA + pointer readout of Archer's reconstruction — public gold, not a Jev teacher; ID ECE only (`notes.md` §45). Hume's 27B dense drop is **Watch**, not a Hub checkpoint. He prefers the class name **decision models** over "system one" | +| **Open System-1 / decision-model head** (Laya, openjev, LightJev, openjev-lm, Nimble, **kev**, **blackwood-rlcd**, Hume **Watch**) | Same *shape* as Jev, you host it | Same primitives or logits-as-options | Air-gap, $0/token, inspectable weights, deployment control; **image-in now** (blackwood) without waiting for Archer | Self-eval duty; Laya text-only, 512 tok; vendor vs-Jev tables are claims (`notes.md` §18). A distill learns the *teacher's* answers: openjev-lm and jev-gate-student-b (`notes.md` §25, §33). Nimble is an open LoRA recipe on hard labels, not a Jev distill (`notes.md` §35). **kev** is a shipped Qwen2.5-0.5B LoRA + pointer readout of Archer's reconstruction — public gold, not a Jev teacher; ID ECE only (`notes.md` §45). **blackwood-rlcd** is open multimodal RLCD (CC BY-NC), Jev-compatible shim; Jev still leads general text (`notes.md` §46). Hume's 27B dense drop is **Watch**, not a Hub checkpoint. He prefers the class name **decision models** over "system one" | | **Encoder open-jev** (DeBERTa-v3-large) | Same *shape*, bidirectional encoder, public gold (not a Jev teacher) | Choice / Score / Noul from one pass | Self-host decide without a decoder; 512 tok | In-domain ECE 0.022; OOD acc 0.854→0.690. English / three public domains. `notes.md` §33 | -| **Constrained-AR surface** (TypeAR, **pcdServer**; not a species) | Next-token constraint on a pretrained generator | Distribution over allowed values | Typed fields without retraining; later fields must see earlier answers; local GGUF serving | Different objective from a proper-scoring head. TypeAR README enums ≤16; pcdServer 2–256 / 1–63 parallel fields. No abstention primitive. Compute-graph card below (`notes.md` §31, §32, §42). Public logit dump: mini-jev-runs | +| **Constrained-AR surface** (TypeAR, **pcdServer**; decision-token LoRA; not a species) | Next-token constraint on a pretrained generator | Distribution over allowed values | Typed fields without retraining; later fields must see earlier answers; local GGUF serving; train the *decision token* if you LoRA | Different objective from a proper-scoring head. TypeAR README enums ≤16; pcdServer 2–256 / 1–63 parallel fields. No abstention primitive. Decision-token QLoRA: `Foodoo1/Qwen3-14B-RLCD-Decision-LoRA` (`notes.md` §46). Compute-graph card below (`notes.md` §31, §32, §42). Public logit dump: mini-jev-runs | | **GLi\* encoder family** (GLiNER locate / GLiClass categorize / GLiNER2.5 local multi-head / GLiGuard safety schema) | One-pass labels-in-encoder; spans, sequence labels, a safety schema, or both | Spans + types; per-label sigmoid/softmax; optional relations/records | Laptop/local; large or changing label sets; "what's *in* the text" vs "what *is* the text" vs "which safety labels fire" | Affinities are not automatically a gateable P(permit). GLiGuard is not a Jev weight clone. Species map below. Not a Jev how-to and not a GLiNER or GLiGuard install | | **Listwise / pairwise discriminative ranker** | Order of a list (nDCG, softmax-over-list) | Relevance scores, not P(relevant) | Rerank a retrieved shortlist | Translation-invariant listwise losses are **not** calibrated for thresholds ([listwise vs pointwise](https://doi.org/10.48550/arxiv.2208.06164); [RCR](https://arxiv.org/html/2211.01494v2)). Fail **open** (keep retrieval order) | | **Vision scorer** | Image–text affinity or region Choice | Cosine/sigmoid affinity, or a closed region/label pick | Perception as classification over *candidates you extracted* | CLIP softmax = competition in the offered set; SigLIP sigmoid = pairwise affinity, not class-conditional p ([SigLIP](https://huggingface.co/docs/transformers/v4.39.2/en/model_doc/siglip)). Not a VLM captioner | @@ -54,7 +54,7 @@ code owns side effects." They are not aliases. ```text locate GLiNER (span NER) what's *in* the text categorize GLiClass / GLiGuard what the text is; which schema labels fire -decide Jev / Laya / openjev / kev Choice / Score / Noul over a state +decide Jev / Laya / openjev / kev / blackwood Choice / Score / Noul over a state (blackwood: image-in) rank listwise / cross-encoder order a retrieved shortlist perceive CLIP / SigLIP / region Choice score candidates you extracted ``` @@ -90,6 +90,10 @@ below, next to the when-to-use table. *your* ECE — ID numbers are not OOD. A constrained autoregressive decode can emit a label and still not be this species — compute-graph card below. Hume's announced 27B dense drop is Watch. + **blackwood-rlcd** is the open multimodal *decide* head that ships now + (screenshot + marked candidates → Choice; CC BY-NC; Jev still leads + general text; `notes.md` §46) — not that drop, and not a vision-scorer + affinity. - **Categorize (safety schema).** [GLiGuard](https://github.com/fastino-ai/GLiGuard) ([arXiv 2605.07982](https://arxiv.org/abs/2605.07982); Zaratiana, Newhauser, Hurn-Maloney, Lewis, Fastino): a GLiNER2 encoder that @@ -253,17 +257,19 @@ capability shift, independent of vendor: calibration optional. Distilling a hosted teacher is not independent gold. Hume's 27B dense drop is the large-local Watch, not a third how-to. -6. **Cross-modal is still thin.** Discourse, GLiNER/GLiClass, Laya, kev, - and the encoder open-jev are text-first. Vision is a scoring pattern - (above), not a shipped omni decision API. Hume reports that a - multimodal *base* plus text post-training generalizes to images with - little intentional multimodal training — a Watch claim, not a - recipe (`notes.md` §33). Treat "Jev but for images" as a hole to fill - with the vision-scorer family or with that drop *when it ships*. - Locate (spans on a screenshot OCR) is still locate, not perceive. - Pixel-free computer-use (jev-macos-loop, jev-mobile) keeps pixels on - the device and sends text-only decisions. SAM 3.1 or ASR then Jev - is specialist composition, not the Watch drop (`notes.md` §39). +6. **Cross-modal has an open decide head now; Archer is still Watch.** + Discourse, GLiNER/GLiClass, Laya, kev, and the encoder open-jev are + text-first. Vision scorers (CLIP/SigLIP) remain a scoring pattern, + not a decision API. [`BlackwoodAI/blackwood-rlcd`](https://huggingface.co/BlackwoodAI/blackwood-rlcd) + is the shipped omni *decide* model: screenshot or text in, typed + Choice out, Jev-compatible shim (`notes.md` §46). Hume's 27B dense + drop stays **Watch** (no Hub weights this pass; multimodal, no audio). + Prefer specialist composition (SAM / ASR / OCR → schema → System One) + when you already have the producer; prefer a shared multimodal + decision model when the joint of pixels and options matters. Pixel-free + computer-use (jev-macos-loop, jev-mobile) still keeps pixels on the + device and sends text-only decisions. Locate (spans on a screenshot + OCR) is still locate, not perceive. 7. **The agent that only has a generator is incomplete.** The missing organ is a judgment-class model plus policy in code — not another prompt. The agent that only has a ranker is also incomplete: it can @@ -381,7 +387,8 @@ program. | Calibrated p(y\|x) over a closed set | Trained decision-only head (Jev, or an open head you have proper-scored and measured on your labels) | Threshold a generated "90%", an affinity you have not calibrated, TypeAR constrained scores, or a LoRA student's agreement with the teacher | | Laptop-local System One API for development / eval | **kev** — trained decision-only readout; official SDK with a `base_url` change (`notes.md` §45) | Treat 0.5B ID ECE as a knowledge or frontier substitute, or as OOD calibration | | Dependent sequential decisions | Constrained AR that conditions later steps on earlier answers (TypeAR sequential), or code-owned transitions and a new request per stage | Treat sibling questions on one request as if they attend each other | -| Open multimodal self-host / data-residency | Hume's announced **decision-model** drop **when it ships** (Qwen3.8 27B **dense**, 265k, multimodal, no audio; one forward pass locally once AR is removed; MoE next then shrink). Driver: healthcare AU residency, not anti-TypeSafe | Ship on "smarter than Jev." That is his early claim, against his own order-sensitivity and in-distribution calibration warnings. **WATCH** — no Hub weights this pass. Laya remains text-only. kev is text-only. jev-visual is region Choice, not this drop | +| Open multimodal self-host / data-residency *now* | **blackwood-rlcd** — trained decision-only readout with image-in; Jev-compatible shim; CC BY-NC (`notes.md` §46) | Wait for Archer's 27B. Treat screenshot-vs-Jev-text as the same input. Threshold a commercial workflow on a non-commercial license. Skip self-eval because web-element acc is 0.907 | +| Open multimodal self-host / data-residency *when it ships* | Hume's announced **decision-model** drop **when it ships** (Qwen3.8 27B **dense**, 265k, multimodal, no audio; one forward pass locally once AR is removed; MoE next then shrink). Driver: healthcare AU residency, not anti-TypeSafe | Ship on "smarter than Jev." That is his early claim, against his own order-sensitivity and in-distribution calibration warnings. **WATCH** — no Hub weights this pass. Laya remains text-only. kev is text-only. jev-visual is region Choice, not this drop | | Image-in now, different graph | Diffusion structured reads that already accept images on a Jev-shaped interface ([djev-spark](https://github.com/mmastrac/djev-spark)) | Wait on the row above for image-in, or treat this graph as a proof it beats a decision head | ### When to use which decision surface @@ -400,8 +407,9 @@ is the generator, not a sixth surface. | Surface | Calibration | VOI / gather | Latency / $ | Deployment control | Multimodal | Enum size | |---|---|---|---|---|---|---| | **Proprietary Jev** | Decision objective; in-dist ECE 0.0313, OOD collapse (`notes.md` §7). Choice `confidence` is arithmetic on the distribution (§31) | Independent questions cheap; sequential gather is a new request | Cloud envelope; ~$0.042/MTok input (their figure, unreproduced; `/pricing` 404 on 2026-09-18, `notes.md` §1, §42) | No weights. AU health data cannot ride this API if residency forbids it | Text. jev-visual is region Choice | ≤255 Choice | +| **blackwood-rlcd** (open multimodal RLCD; CC BY-NC) | Decision objective; temp-calibrated. Card ECE **0.037** on 300 web steps; letter-shuffle flip **0.133** vs Jev 1.13 **0.587**. Jev still leads general text **0.850** vs **0.786** (`notes.md` §46) | Screenshot/DOM candidates code already marked → Choice. Not gather-as-act | ~200 ms / decision 1×H100 (their figure) | Self-host; non-commercial license | **Image-in now.** Not Archer Watch. Not CLIP | Lettered candidates; Jev-shaped Choice | | **Archer open decision-model** | **Watch.** No Hub weights this pass. "Smarter than Jev" is a claim against *his* calibration/order warnings | Same *hole* as Jev when it ships | 27B dense for one-forward-pass local speed once AR is removed; MoE next, then shrink. Quant-friendly is a claim | Healthcare AU data-residency / deployment control, **not** anti-TypeSafe | Multimodal, no audio. Text post-training reportedly generalizes to images with little intentional multimodal training | Unknown until the drop | -| **TypeAR / pcdServer** (constrained AR) | Next-token constraint ≠ Noul. No abstention primitive. Public logit dump: [`Mikhail/mini-jev-runs`](https://huggingface.co/datasets/Mikhail/mini-jev-runs) (27.9k; scores "deliberately *not* calibrated") | TypeAR sequential conditions later fields; pcdServer batches independent fields after one prefix. Neither is gather-as-act | TypeAR 5.8× is *their* K=16 boolean example. pcdServer: native llama.cpp, Apple+Linux | Self-host the generator / GGUF | Whatever the base model has | TypeAR enums ≤16; pcdServer 2–256 strings, 1–63 fields | +| **TypeAR / pcdServer** (constrained AR) | Next-token constraint ≠ Noul. No abstention primitive. Public logit dump: [`Mikhail/mini-jev-runs`](https://huggingface.co/datasets/Mikhail/mini-jev-runs) (27.9k; scores "deliberately *not* calibrated"). Decision-token QLoRA trains *that* token under parallel constrained decode (`Foodoo1/Qwen3-14B-RLCD-Decision-LoRA`; synthetic fraud receipt, not a financial product; `notes.md` §46) | TypeAR sequential conditions later fields; pcdServer batches independent fields after one prefix. Neither is gather-as-act | TypeAR 5.8× is *their* K=16 boolean example. pcdServer: native llama.cpp, Apple+Linux. Foodoo1: ~234 ms / 4-field broadcast on RTX 3090 4-bit (their figure) | Self-host the generator / GGUF / adapter | Whatever the base model has | TypeAR enums ≤16; pcdServer 2–256 strings, 1–63 fields | | **Encoder open-jev** (DeBERTa-v3-large 434M) | Public gold, CE+Brier, val temperature. In-domain ECE 0.022 / acc 0.854; OOD acc 0.690 / ECE 0.035. **Not** a Jev teacher-copy | One pass over state + all questions; 512 tok | Author: 28 ms / 10 questions H100; 1.8 s / 4q M1 Max CPU | apache-2.0, self-host | Text | Jev-shaped 255 / Score 2–10 / Noul; 512 ctx | | **Tiny LoRA distill** (jev-gate-student-b) | Teacher-copy. P(relevant) from yes/no logits. Held-out n=60 vs vanilla 0.5B; 148,160-row corpus | Memory-gating / context sieve; **fail-open** on errors | Qwen2.5-0.5B LoRA; ~59 ms RTX 3060 | Local, apache-2.0 | Text | Binary relevance | | **Nimble** (open LoRA recipe, not a distill) | Hard synthetic labels. They say temperature was not tuned to correctness rates. 324-row agreement is their receipt, not an ECE (`notes.md` §35) | Not a gather primitive | Their latency table, not re-run | Self-host the adapter. Model card Apache-2.0; repo license absent | Text only | Enum ≤26; 2,048 tokens | @@ -410,11 +418,13 @@ is the generator, not a sixth surface. **Three open paths** (not three species, not extra when-to-use rows): encoder open-jev (DeBERTa, public gold); AR constrained decode (TypeAR -Python/SGLang, pcdServer native GGUF); trained decision-only (Laya / -Nimble / **kev** / Archer **Watch**). **kev** is the cleanest *runnable* -productization of Archer's reconstruction on that third path -(API-compatible; `notes.md` §45). Watch stays Watch. Encoder vs decoder -**replicas** of that third path: DeBERTa is public gold with an OOD +Python/SGLang, pcdServer native GGUF; decision-token LoRA on that graph); +trained decision-only (Laya / Nimble / **kev** / **blackwood-rlcd** / +Archer **Watch**). **kev** is the cleanest *runnable* productization of +Archer's reconstruction on that third path (API-compatible; text-only; +`notes.md` §45). **blackwood-rlcd** is the third path with **image-in +now** (CC BY-NC; Jev-compatible shim; `notes.md` §46). Watch stays Watch. +Encoder vs decoder **replicas** of that third path: DeBERTa is public gold with an OOD drop; openjev-lm / jev-gate LoRAs are teacher-copies with named receipts (overnight 6-vCPU, $0/call — economics, not a new species; `notes.md` §25, §44). kev is public gold on the same 0.5B backbone as openjev-lm, @@ -422,7 +432,10 @@ not a teacher-copy. Pick from the hole. A constrained softmax is still not a Noul. Laya companion packaging this hour: [`laya-typed-decisions`](https://huggingface.co/convaiinnovations/laya-typed-decisions) (same 421.3M; acc 0.766 / Brier 0.066 on `LocalLLaMA/typed-decisions`, -unverified — do not overwrite `notes.md` §18). +unverified — do not overwrite `notes.md` §18). Shared bake-off this +hour: [`pngwn/open-jev-laya-bench`](https://huggingface.co/datasets/pngwn/open-jev-laya-bench) +(26 + 9 tasks, 11,959 items; ECE/NLL/Brier; **not** TypeSafe Jev vs +Laya; LLM-as-judge is not the score — `notes.md` §46, `validation.md`). Reject: TypeAR or pcdServer scores as fail-closed P(permit); a LoRA student's agreement with Jev as independent gold; kev's ID ECE as a license to skip @@ -435,7 +448,8 @@ universal ranking. Diffusion beating a decision head is Hypothesis. Detail: `research/notes.md` §33 (surfaces), §34 (marginals), §35 (Nimble), §36 (diffusion), §38 (entropy allocator, Hypothesis), -§42 (pcdServer serving, meta-VOI, games), §45 (kev). Before +§42 (pcdServer serving, meta-VOI, games), §45 (kev), §46 (blackwood, +laya-bench, decision-token LoRA). Before adopting a surface, the bake-off is a jevals-shaped suite and, for a product loop, a Harbor taskset (`validation.md`, Eval & hill-climb). The stage pipeline into that decision is the same file @@ -456,10 +470,13 @@ the route or the patches. `notes.md` §36. ### "Perception specialist then judgment specialist" vs "Shared multimodal System One" -**Hypothesis.** Basit ask, primary post not retrieved (X and web, -2026-09-18; no tweet id). Useful *application* patterns for omni-ish -products. Not a SAM tutorial, not an ASR tutorial, and not a native -omni System One. +Two placements, not a slogan. Specialist composition stays **Hypothesis** +as a general recipe (Basit ask, primary post not retrieved; `notes.md` +§39). Shared multimodal *decide* now has a named open model: +[`BlackwoodAI/blackwood-rlcd`](https://huggingface.co/BlackwoodAI/blackwood-rlcd) +(**Empirical** as that vendor receipt, not re-run; **Hypothesis** on +*your* labels; `notes.md` §46). Not a SAM tutorial, not an ASR tutorial, +and not Archer's 27B drop (**Watch**). [SAM 3.1](https://huggingface.co/facebook/sam3.1) is Meta Segment Anything 3.1 ([release](https://github.com/facebookresearch/sam3/blob/main/RELEASE_SAM3p1.md), @@ -474,23 +491,28 @@ of transcript-then-Jev, not the source of this ask: (2026-09-17). His latency and price are his receipt, not a class number. -**Information dies at the interface.** The decision call sees the -schema you serialized, not the pixels or the waveform. Prefer a shared -multimodal decision model when that joint signal matters. Archer -Hume's drop stays **Watch** (no Hub weights as of 2026-09-18; -multimodal, no audio; do not promote to Empirical). djev-spark already -accepts images (multipart or JSON; think and sequential reject -images). A future audio-capable shared model is the same hole, not -this stack. +**Information dies at the act.** Specialist composition: the decision +call sees the schema you serialized, not the pixels or the waveform. +Shared multimodal decide (blackwood): the call sees the screenshot +*and* the candidates **code already marked** (letters on the image); +it still does not invent a click. Prefer specialist composition when +you already have the producer; prefer a shared multimodal decision +model when the joint of pixels and options matters — you no longer +have to wait for Archer to place that hole. Archer Hume's drop stays +**Watch** (no Hub weights as of 2026-09-18; multimodal, no audio). +djev-spark already accepts images (multipart or JSON; think and +sequential reject images) on a *different* compute graph. A future +audio-capable shared model is the same hole, not this stack. Same cut as a low/medium-entropy allocator: specialist state, then typed decisions. Buckets are product rhetoric, not a meter, and "review this PR" is still partly generative. One Jev call is still -factorized **marginals** (Meijer); the joint of the raw signal and the -decision lives outside the call. The handoff is a contract surface — -one sentence in `formal-methods.md`. Code owns the schema and the act. -`notes.md` §39. Measure and hill-climb that composition in -`validation.md` (**Hypothesis**, `notes.md` §41). +factorized **marginals** (Meijer); blackwood's one prefill is still +not a joint over the raw waveform. The handoff is a contract surface +— one sentence in `formal-methods.md`. Code owns the schema and the +act. `notes.md` §39, §46. Measure and hill-climb in `validation.md` +(`notes.md` §41). Jev still leads blackwood on general *text* (0.850 +vs 0.786 on their 8,456-item table) — omni is not a text free lunch. ### Open recipe (Bespoke Nimble) — not a distill @@ -574,6 +596,26 @@ tests (isolation, permute, IIA, boundary forgery) mirror Archer probes — they falsify the reconstruction, they do not prove kev = Jev. `notes.md` §45. +### blackwood-rlcd — open multimodal RLCD (not Archer, not CLIP) + +[`BlackwoodAI/blackwood-rlcd`](https://huggingface.co/BlackwoodAI/blackwood-rlcd) +(Hub README this pass; CC BY-NC 4.0; `image-text-to-text`). Trained +decision-only readout with **image-in**. One prefill; option-letter +logits; temperature-calibrated; Jev-compatible `/v1/systemone` shim. +Screenshot + candidates **code already marked** → typed Choice; code +clicks. Not Archer's 27B drop. Not a vision-scorer affinity. Not a +commercial drop-in. Do not copy serve flags. + +**Evidence (card; paired per item; not re-run).** Web element, 300 +held-out: acc **0.907** vs Jev 1.13 text-only **0.480**; letter-shuffle +flip **0.133** vs **0.587**; ECE **0.037** vs **0.091**; ~**200 ms** +1×H100 vs 441 ms OpenRouter. General text, 85 sets / 8,456 items: +**0.786** vs Jev **0.850** — Jev still leads. Screenshot rows compare +screenshot input with Jev's text input on the same steps. Randomized +viewport crop. **Empirical** as that named receipt. **Hypothesis** on +*your* labels. Bake-off candidate beside kev / Laya (`validation.md`). +`notes.md` §46. + ### Constrained-AR surface (not `decide`) [TypeAR](https://github.com/zmtomorrow/TypeAR) ("Type-Safe Decoding for @@ -619,6 +661,17 @@ install. Same author's earlier `parallelConstraintDecoding` is the two-forward-pass cousin already in the ecosystem snapshot. `notes.md` §42. +**Decision-token LoRA (same graph, trained).** If you specialize a +pretrained generator for this surface, train **the single decision +token** under parallel constrained decoding — not generated prose. +[`Foodoo1/Qwen3-14B-RLCD-Decision-LoRA`](https://huggingface.co/Foodoo1/Qwen3-14B-RLCD-Decision-LoRA) +(Apache-2.0 adapter on Qwen3-14B): loss only on that token; one +prefill + KV broadcast across fields. Their held-out 200-case / 4-field +receipt: fraud_risk **64.0% → 95.0%**, overall **85.2% → 98.8%** at +**~234 ms** / broadcast on RTX 3090 4-bit. Synthetic fraud-triage; +do not use for real financial decisions. Softmax over allowed tokens +is still not a Noul. `notes.md` §46. + **Hypothesis, not a stack.** TypeAR's example names the same Qwen3.8-27B family as Hume's announced weights. Running that surface (TypeAR or pcdServer) on those weights versus stock Qwen is a composition to test diff --git a/.agents/skills/augustus/references/mappings.md b/.agents/skills/augustus/references/mappings.md index 330b5cb..9a7be08 100644 --- a/.agents/skills/augustus/references/mappings.md +++ b/.agents/skills/augustus/references/mappings.md @@ -151,9 +151,13 @@ vs **out-of-process CLI** ([`kylemclaren/jevql`](https://github.com/kylemclaren/ scan, not an index; `max_rows` is a spend guard; thresholds stay in SQL. [`ant4g0nist/joxide`](https://github.com/ant4g0nist/joxide): zoxide owns the directory index; Jev scores a shortlist; destinations are existing -local paths only; fail-open. Row contents leave the store (same -residency warning as AU health). Do not copy SQL, env, or CLI flags. -`notes.md` §42, §44. +local paths only; fail-open. Dataframe cousin this hour: +[`yalindogusahin/jevpandas`](https://github.com/yalindogusahin/jevpandas) +— `evaluate` / `filter` / `classify` / `score` / batched `ask` over a +pandas frame; classify example includes `other`; failures never become +negative predictions; LICENSE absent this pass. Row contents leave the +store (same residency warning as AU health). Do not copy SQL, env, or CLI flags. +`notes.md` §42, §44, §46. ## 5. Hierarchy → bounded heuristic search @@ -573,9 +577,13 @@ questions, **Hypothesis**. Links: `mental-models.md` §thresholds; **Method**: code (or a recipe, a law, a text layer) **proves** the easy cases; a System One model judges only what the structure cannot decide. Composition-algebra position 3 *after* a constraint, not instead of one. -**Transfers**: allowlist / refused-in-code / unknown→judge +**Transfers**: allowlist / refused-in-code / unknown→judge. The +allowlist **proves** every verb is a listed read-only tool; the model +judges **only unlisted** leftovers; the gate **cannot block** (fail-open +unless a sandbox sits under) ([jevgate](https://github.com/thevibeworks/jevgate): Proven / Refused / -Unknown; cannot block; Jev alone leaks). Same sandwich as page OCR +Unknown; Jev alone leaks — `/bin/ls` at 0.04 is why it is the third +tier). Same sandwich as page OCR ([doc-router](https://github.com/misbahsy/doc-router): pdf-inspector first, "needs OCR?" Noul on the remainder — 155→87 pages billed, **1.74×** $ on 19 docs / 155 pages). **Does not:** putting the model first so a diff --git a/.agents/skills/augustus/references/mental-models.md b/.agents/skills/augustus/references/mental-models.md index 5a70e2e..3cb5397 100644 --- a/.agents/skills/augustus/references/mental-models.md +++ b/.agents/skills/augustus/references/mental-models.md @@ -116,9 +116,10 @@ LoRA students report agreement with the teacher (`notes.md` §33). Hume prefers the class name **decision models** over "system one" (`notes.md` §33); this file still says System One when quoting TypeSafe. Three open paths, not three species: encoder open-jev, AR constrained -decode (TypeAR + pcdServer), trained decision-only (Laya / Nimble / -kev / Archer Watch). kev is the runnable Archer reconstruction on that -third path; Watch stays Watch. A constrained softmax is still not a Noul (`notes.md` §42, §45). +decode (TypeAR + pcdServer; decision-token LoRA), trained decision-only (Laya / Nimble / +kev / **blackwood-rlcd** / Archer Watch). kev is the runnable Archer reconstruction on that +third path (text-only); blackwood-rlcd is that path with **image-in now** (CC BY-NC); +Watch stays Watch. A constrained softmax is still not a Noul (`notes.md` §42, §45, §46). **Readout versus a token; IIA is a property.** A direct probability and a generated "91%" are different objects; the format calibrates neither @@ -400,7 +401,8 @@ Use these as *existence proofs of a position*. Write your own card. | Hiring | interview / reject / hold | evidence Nouls; veto rules in policy | labor law, scorecards you wrote | | Inbox | reply / snooze / archive | urgency Noul + aboutness Choice | send, calendar | | Document / lab routing | which pages need the expensive observation | Noul on remainder after a text layer / recipe | local extract, merge order (**Empirical** as OCR-router *shape*) | -| Shell / tool allowlist | unlisted remainder after a proof | five Nouls on unknown verbs | Proven/Refused in code (**Empirical**: jevgate) | +| Shell / tool allowlist | unlisted remainder after a **proof** | five Nouls on unknown verbs | Proven/Refused in code; cannot block (**Empirical**: jevgate) | +| Screenshot candidates → act | lettered elements code already marked | Choice over those letters | Click in code (**Empirical** as blackwood-rlcd *shape*; CC BY-NC) | | Moderation | hold before publish | hazard Nouls (**Empirical** as family) | block/review policy | | Phishing / fraud screen | hold vs deliver | SDT criterion on a Noul | blocklist, SPF/DKIM exact (**Hypothesis**) | | Personal ops | cook done / not | "looks done" Noul | thermometer probe | diff --git a/.agents/skills/augustus/references/methods-catalog.md b/.agents/skills/augustus/references/methods-catalog.md index a950f66..f2febe9 100644 --- a/.agents/skills/augustus/references/methods-catalog.md +++ b/.agents/skills/augustus/references/methods-catalog.md @@ -56,7 +56,8 @@ judgment component is new). | Entropy as uncertainty signal | Measuring "how spread is this belief" | Entropy of returned distributions across repeats or options — computed in code from returned probabilities | All arithmetic | **Empirical recipe** (cookbook pattern) | | Detector / Neyman filter (context) | Is this artifact relevant to the current task? | One relevance Noul per block before it enters context; stub + recall key | Cache, recall, safety keeps | **Empirical recipe** (winnow ≤0.22 hide; compaction 2-noul rule; pi-jev-context hide-not-delete) | | Anomaly detection | Does this deviate from expected shape? | Guard nouls + harm Score over {input, output, tool trace} | Baselines, alert thresholds | **Empirical recipe** (guardrails cookbook; pi-jev output judge) | -| Allowlist ∩ remainder (code-then-model) | Unlisted / unstructured leftovers after a proof | Typed questions only on the unknown tier; admit iff every p < τ | Proven/refused in code; cannot block unless a sandbox sits under | **Empirical recipe** (jevgate 0/59 unsafe unasked held-out; doc-router 1.74× $). Domain-general: `mappings.md` §18 | +| Allowlist ∩ remainder (code-then-model) | Unlisted / unstructured leftovers after a **proof** | Typed questions only on the unknown tier; admit iff every p < τ | Proven/refused in code; cannot block unless a sandbox sits under | **Empirical recipe** (jevgate 0/59 unsafe unasked held-out; allowlist *proves* read-only verbs; doc-router 1.74× $). Domain-general: `mappings.md` §18 | +| Decision-token LoRA (constrained-AR) | Specialize a generator for parallel constrained fields | Loss only on the single decision token; KV broadcast across fields | Schema, candidate tokens, policy | **Empirical recipe** as Foodoo1 200-case / 4-field receipt (fraud_risk 64→95%, overall 85.2→98.8%, ~234 ms); **Hypothesis** as a general recipe. Synthetic; not a financial product. Softmax ≠ Noul | | Teacher distill of judgments | Copy a hosted decision API onto a small local head | LoRA / frozen-encoder heads trained on teacher answers | Independent gold labels; ECE on *your* cases | **Empirical recipe** as one 70-row run (openjev-lm 92.9%); **Hypothesis** as a general recipe | ## Verification & logic diff --git a/.agents/skills/augustus/references/mixed-architecture.md b/.agents/skills/augustus/references/mixed-architecture.md index 75cf209..328bb43 100644 --- a/.agents/skills/augustus/references/mixed-architecture.md +++ b/.agents/skills/augustus/references/mixed-architecture.md @@ -31,7 +31,10 @@ If the request is "how do I call Jev?", stop and load `typesafe-ai`. If it is "should this step be a judgment-class model, an LLM, a regex, or a trained classifier — and which family?", stay here. Family table: `judgment-class.md`. Cross-domain frames: `mental-models.md`. Proof vs -judgment: `formal-methods.md`. +judgment: `formal-methods.md`. [wellposed](https://github.com/suraj-phanindra/wellposed) +is an Empirical *recipe* of tenbin's lint hole (missing `other` → +confidence 1.00 wrong); it is not a second Augustus skill +(`notes.md` §46). ## Default architecture @@ -261,6 +264,16 @@ treating a workflow AST as a proof. The outer loop stays with the LLM or with code. **Hypothesis** as "Jev builds the AST"; **Empirical** as named topologies. Not an MCP how-to. +**S1 keeps control; S2 is one-use advice.** Topology B under latency: +the reflex (typed Choice over legal actions) never hands the stick to +the planner. Optional System 2 is *advice* on low confidence, one-use, +asynchronous — the reflex does not pause +([jev-reflex-autonomy-lab](https://github.com/khordoo/jev-reflex-autonomy-lab); +experimental drone viz, not a flight controller; GitHub license null +this pass). Same Kahneman split as the toolbox row (S2 proposes, S1 +discriminates; never the reverse). `notes.md` §46. +`agent-self-assessment.md`. + **Effect-oriented loop (same author, later post).** Topology B inside an effect system ([tweet](https://x.com/JamesWard/status/2100981305009664299)): the host @@ -352,7 +365,12 @@ decision-design card. Do not clone APIs from READMEs. | Finish-line gate | Noul/Score/Choice on evidence | Deterministic shell checks first | hermes-jev-north-star | | Home automation read | Choice/Score/Noul as an entity | Automations, device I/O | `AboveColin/HA-Jev` | | Browser loop without generation | Action Choice over visible elements | Perception, constraints, click | lizard-agent | +| Screenshot / DOM candidates → Choice | Omni decide over letters code marked | Click/act in code; fail-open to specialist OCR | blackwood-rlcd (CC BY-NC; not Archer) | | Android / macOS computer-use | Choice over prevalidated candidates | UI tree / AX / OmniParser; no generated coordinates | jev-mobile, jev-macos-loop | +| S1 reflex + optional S2 advice | Typed action Choice; planner one-use on low p | Collision, legality, the stick stays with S1 | jev-reflex-autonomy-lab (experimental) | +| Decision-as-business-tool | Named judgment; gate is part of the result | Registry, arithmetic, hard guards | jev-decision-layer (unofficial) | +| NL cases → checked e2e | Jev selects observed controls | Playwright expectations; PASS/FAIL/BLOCKED | jev-e2e (alpha) | +| Dataframe semantic index | Noul / Choice / Score per row | pandas, thresholds, never invent negatives | jevpandas | | Model router | Requirement Scores; policy in code | Eligibility, cost/quality/latency objective | routeKit | | Bulk-judgment coprocessor | Choice/Noul off the frontier context | Counts, policy, fail-open gate | jev-mode | | Closed-catalog System One shell | Choice over host tools | Execute, arithmetic, credentials | jot | @@ -373,14 +391,17 @@ Reproduce/open heads (`rongxinzy/LightJev`, openjev family, encoder [`open-jev-deberta-v3-large`](https://huggingface.co/com-kotobalabs/open-jev-deberta-v3-large), LoRA [`jev-gate-student-b`](https://huggingface.co/SargeDev/jev-gate-student-b), companion packaging [`laya-typed-decisions`](https://huggingface.co/convaiinnovations/laya-typed-decisions), -[`jaredpalmer/kev`](https://github.com/jaredpalmer/kev)) +[`jaredpalmer/kev`](https://github.com/jaredpalmer/kev), +[`BlackwoodAI/blackwood-rlcd`](https://huggingface.co/BlackwoodAI/blackwood-rlcd)) are evidence that the *interface* (Choice/Score/Noul, or yes/no logits as P(relevant)) is the transferable part — not a request to implement a backbone or a second API skill. Laya: self-hostable, text-only, 512 tokens/question; vendor benches vs Jev are **claims**. Encoder open-jev: public gold, OOD drop. LoRA student: teacher-copy. **kev**: public gold, pointer readout, System One API drop-in; ID ECE only; not a teacher-copy -(`notes.md` §45). Hume's 27B +(`notes.md` §45). **blackwood-rlcd**: open multimodal RLCD, Jev-compatible +shim, CC BY-NC; Jev still leads general text; not Archer Watch +(`notes.md` §46). Hume's 27B decision-model drop is **Watch**. Closed calibrated API vs open weights is a self-eval tradeoff (`research/notes.md` §18, §33, §45). When-to-use axes: `judgment-class.md`. TypeSafe remains the documented *exemplar*, diff --git a/.agents/skills/augustus/references/optimizer-integration.md b/.agents/skills/augustus/references/optimizer-integration.md index 4185274..367c4fd 100644 --- a/.agents/skills/augustus/references/optimizer-integration.md +++ b/.agents/skills/augustus/references/optimizer-integration.md @@ -13,7 +13,9 @@ signatures, and `typesafe-ai` plus the live docs own Jev's request body. Do not write either from this page. DSPy and Ax tune the LM-program slice only. They are never the primary System One calibration score; that seat is a jevals-shaped labeled suite, and a product loop is a -Harbor taskset (`validation.md`, Eval & hill-climb). +Harbor taskset (`validation.md`, Eval & hill-climb). Shared bake-off +exemplar this hour: [`pngwn/open-jev-laya-bench`](https://huggingface.co/datasets/pngwn/open-jev-laya-bench) +scores ECE/NLL/Brier — not an LLM-as-judge paragraph (`notes.md` §46). ## Judgment: what these optimizers may climb (Hypothesis) diff --git a/.agents/skills/augustus/references/question-design.md b/.agents/skills/augustus/references/question-design.md index 54cf681..84703c3 100644 --- a/.agents/skills/augustus/references/question-design.md +++ b/.agents/skills/augustus/references/question-design.md @@ -44,7 +44,7 @@ request, and treat a stale pin as a prior, never a setting. ## Criteria shape - Criteria are an extension of the instruction and must ask the same thing, in the same direction (a Noul whose `true` side describes "no" performs worse). -- Choice options: contrastive `what` / `not_for` / short concrete `examples` (instances, not descriptions of instances). Add an `other` / `none-of-the-above` when the list may not cover inputs. +- Choice options: contrastive `what` / `not_for` / short concrete `examples` (instances, not descriptions of instances). Add an `other` / `none-of-the-above` when the list may not cover inputs. Skipping that hatch is not a style nit: the model will pick a listed option at confidence 1.00, and no downstream gate will see a problem (`notes.md` §46). - Score levels (2–10): describe **situations**, one dimension each, each standing alone (Jev sees neither the level's number nor its neighbors — "worse than previous" means nothing). No numerals. Levels may be objects `{"summary", "signals"}`. Give a rare extreme its own level when code treats it differently. - Composite scoring: one Score per dimension, normalize by `len(criteria)-1`, weight and combine in code. Change policy by changing weights — never by rewriting questions. - Taxonomy walk: one Choice per tree level, walk in code; each option's value is its subtree (direct children + sample leaves); follow several branches when probabilities are close. @@ -54,6 +54,8 @@ request, and treat a stale pin as a prior, never a setting. | Symptom | Likely cause | Fix | | --- | --- | --- | | Wrong answers, high confidence | Instruction read literally | State exact condition; put boundary cases in criteria | +| Wrong answers, **confidence ~1.00**, no `other` | Forced pick: the offered set does not cover the input; the model *must* choose | Add `other` / none-of-the-above. **Confidence gating cannot catch this** ([wellposed](https://github.com/suraj-phanindra/wellposed) live probe: unsubscribe email → `"support issue"` at 1.00 without `other`, `"other"` at 0.93 with it). Overlapping options collapse confidence (loud). `notes.md` §46. `tenbin` owns the lint skill | +| Question names a state path that does not exist | Dead reference; the API still answers | Lint the request (walk JSON). Structural, not semantic. wellposed recipe; do not copy the CLI | | Low-confidence Choice | Options overlap / none fits | `what`/`not_for`/`examples`; add `other` | | Low-confidence Score | Overlapping levels, two dimensions, thin state | Distinct-situation levels; split question; add state field | | Scores cluster mid-scale | Levels are degrees/numbers | One concrete situation per level; remove numerals | diff --git a/.agents/skills/augustus/references/toolbox-mapping.md b/.agents/skills/augustus/references/toolbox-mapping.md index f998655..1b989fb 100644 --- a/.agents/skills/augustus/references/toolbox-mapping.md +++ b/.agents/skills/augustus/references/toolbox-mapping.md @@ -69,9 +69,9 @@ component; keep the rest of the method in code. | Probabilistic method: priors | Choice distribution as P(s,a) policy prior (MCTS/PUCT); Choice confidences as calibrated gating | **Empirical recipe** (jev-mcts, calibrated vs exact truth) | | Search: value function | Score rubric as leaf value V(s) — only where a simulator validates outcomes; speculative depth hard-capped at 2 | **Empirical recipe** (jev-mcts fidelity split) | | Measurement theory: probe vs estimate | Only post-execution probes concede milestones; model estimates never do — "estimation wearing a measurement costume" is the rejection template | **Empirical recipe** (jev-mcts, pi-warden done-check) | -| Experimental design: perturbation | Behavioral tests as the stats layer: candidate removal, option-order shuffle, distractor injection, boundary cases | **Contract-level** (validation.md) | +| Experimental design: perturbation | Behavioral tests as the stats layer: candidate removal, option-order shuffle, letter-shuffle on screenshot Choice, distractor injection, boundary cases | **Contract-level** (validation.md); letter-shuffle receipt: blackwood-rlcd 0.133 vs Jev 1.13 0.587 on 300 web steps (`notes.md` §46) | | Discrete math: width vs depth | Fan out in width (parallel ≈ free), pay depth linearly; two-stage only when next options depend on an earlier answer | **Empirical recipe** (fan-out: 12.2× cheaper, 10× faster) | -| Psychology: Kahneman | System 2 generates/proposes (LLM), System 1 discriminates (Jev); never the reverse | **Empirical recipe** (mcts-agent role split; 2026-09-18 mixed-architecture discourse) | +| Psychology: Kahneman | System 2 generates/proposes (LLM), System 1 discriminates (Jev); never the reverse. S1 keeps control; optional S2 is one-use advice and does not fly | **Empirical recipe** (mcts-agent role split; 2026-09-18 mixed-architecture discourse; jev-reflex-autonomy-lab as a control-loop shape, `notes.md` §46) | | IR / cascades | Cheap relevance / irrelevance before an expensive ranker or generator | **Empirical recipe** (RAG cookbook; jevprune; git-jev-stage; LlamaIndex Jev rerank) | | Spec / lint | Project-defined semantic rules as predicates over a diff | **Empirical recipe** (jev-pref contract; JevLint file-level Noul; pi-warden; snifftest unsure-band) | | Formal methods / DST / safety | Judgment triages counterexamples, failing seeds, and named-rule conformance; proof/MC/DST stay with their tools. Alloy finder ≠ Apalache BMC ≠ Quint run. DST trio: Antithesis hypervisor / Resonate HQ Lean+oracle+SDK (durable async) / PufferLib env+seed. Noul is a sensor, not a discharged PO. Semi-formal diagrams are vocabularies, not enforcers | **Hypothesis as product**, **Contract** as ownership (matching `mappings.md` §8 and `methods-catalog.md`; worked shape pi-warden — `formal-methods.md`, `formal-semi-formal.md`) | diff --git a/.agents/skills/augustus/references/validation.md b/.agents/skills/augustus/references/validation.md index af12e0a..d757476 100644 --- a/.agents/skills/augustus/references/validation.md +++ b/.agents/skills/augustus/references/validation.md @@ -4,7 +4,12 @@ 1. Exact computation, semantic judgment, or both? Exact parts stay in code. 2. Can the right answer be represented? (candidate present? level exists? - `other` option where coverage is open?) + `other` option where coverage is open?) Missing `other` on an open + coverage set forces a wrong Choice at confidence 1.00 — **confidence + gating cannot catch it**. Lint the *request* (broken state paths, + bundled judgments) before you trust the answer + ([wellposed](https://github.com/suraj-phanindra/wellposed) recipe; + `tenbin` owns the skill; `question-design.md`; `notes.md` §46). 3. Missing / contradictory / malicious / stale evidence — what happens? 4. Which constraints must code enforce regardless of model output? 5. What does each number mean — and which reading would be invalid? @@ -15,7 +20,9 @@ ## Behavioral tests (measure; Jev promises no invariances) Candidate removal (drop the winner — does probability spread sensibly?); -option-order shuffle; **irrelevant-option / IIA** (append an option that +option-order shuffle; **letter-shuffle on screenshot Choice** +(blackwood-rlcd card: flip **0.133** vs Jev 1.13 text-only **0.587** on +300 web steps — vendor receipt, not re-run; `notes.md` §46); **irrelevant-option / IIA** (append an option that should not move odds among the rest — Hume's reconstruction, `research/notes.md` §31, not a new invariance the API promises); public cousin for the read-the-letter graph: @@ -293,9 +300,12 @@ Rules: - **Room / omni products** (same ask): structural gates first; video-as-judge last. Same sandwich as allowlist-then-remainder (`mappings.md` §18). Perception-then-judgment is composition; - information dies at the interface, and a Noul is not over raw pixels - (`notes.md` §39). LLM-as-judge is not the primary score for a - calibrated System One. + information dies at the act. A shared multimodal decide head still + judges marked candidates, not an open click (`notes.md` §39, §46). + LLM-as-judge is not the primary score for a + calibrated System One. Shared bake-off exemplar: + [open-jev-laya-bench](https://huggingface.co/datasets/pngwn/open-jev-laya-bench) + (ECE/NLL/Brier). ### Composition @@ -312,15 +322,30 @@ row is enough. ECE above is wanted, not a Nimble result. ### Bake-off mandate -Before adopting proprietary Jev vs Laya vs TypeAR vs Nimble vs kev vs Archer -vs openjev-lm, run a jevals-shaped labeled suite (or an equivalent +Before adopting proprietary Jev vs Laya vs TypeAR vs Nimble vs kev vs +blackwood-rlcd vs Archer vs openjev-lm, run a jevals-shaped labeled suite (or an equivalent with this hygiene) and, for a product loop, a Harbor taskset. A design card with no eval path is incomplete. +**Shared bake-off exemplar (Empirical as that named receipt, not a +ranking).** [`pngwn/open-jev-laya-bench`](https://huggingface.co/datasets/pngwn/open-jev-laya-bench) +(`RESULTS.md` this pass): System One Qwen3.5-4B scorer vs Laya 421M, +**26 neutral + 9 home**, **11,959** test / **3,269** cal. Scores are +accuracy + **ECE / NLL / Brier** (acc@50% coverage too). Neutral macro +acc Δ **+0.023 [+0.013, +0.032]**; home Δ **+0.229 [+0.198, +0.262]** +(intervals exclude 0). **Not TypeSafe Jev vs Laya.** Neutral prompted- +instruct on the same 4B is statistically tied with the fine-tune +(+0.003, interval includes 0). **LLM-as-judge is not the primary +System One score.** Harbor/jevals practice in the wild: held-out +`test`, temperature on `cal`, leave-one-task-out, prompted arms. +`notes.md` §46. Do not copy the scoring scripts. + Archer weights are still a **Watch** — not on the Hub as of 2026-09-18 (`notes.md` §31–§33). That bake-off is future, not Empirical. kev is the shipped 0.5B reconstruction on the trained decision-only path, not -that drop (`notes.md` §45). "A 9B +that drop (`notes.md` §45). blackwood-rlcd is the open multimodal +decide head on that path **now** (CC BY-NC; Jev still leads general +text; `notes.md` §46). "A 9B LoRA is enough versus Jev" stays **Hypothesis** (`notes.md` §35). openjev-lm is the name of that distill. kev is not a Jev distill. Nimble is not a Jev distill diff --git a/CHANGELOG.md b/CHANGELOG.md index 21142f1..f5d3087 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -124,6 +124,23 @@ Pegged against [`typesafe-ai/skills` v0.5.7](https://github.com/typesafe-ai/skil knowledge/frontier substitute; not a Jev teacher-copy. Contrast vs TypeAR, encoder DeBERTa, proprietary Jev. jevals/Harbor bake-off candidate. No serve how-to. +- Hourly ~14:03 Boise fold (`research/notes.md` §46): Archer still + Watch. Open multimodal RLCD + ([blackwood-rlcd](https://huggingface.co/BlackwoodAI/blackwood-rlcd), + CC BY-NC): screenshot + marked candidates → Choice; web acc 0.907 vs + Jev 1.13 text-only 0.480; letter-shuffle 0.133 vs 0.587; ECE 0.037; + ~200 ms H100; Jev still leads general text 0.850 vs 0.786. Shared + bake-off ([open-jev-laya-bench](https://huggingface.co/datasets/pngwn/open-jev-laya-bench)): + 26+9 tasks, 11959 items; ECE/NLL/Brier; macro acc Δ +0.023 + neutral / +0.229 home; LLM-as-judge is not the score. Decision-token + QLoRA + ([Foodoo1/Qwen3-14B-RLCD-Decision-LoRA](https://huggingface.co/Foodoo1/Qwen3-14B-RLCD-Decision-LoRA)): + fraud_risk 64→95%, overall 85.2→98.8% at ~234 ms/4-field broadcast; + synthetic. jevgate frame: allowlist *proves*, Jev judges only + unlisted, fail-open. wellposed: missing `other` → confidence 1.00 + wrong; gating cannot catch it (`tenbin` owns the lint skill). + S1 reflex keeps control (jev-reflex-autonomy-lab). MED: + jev-decision-layer, jev-e2e, jevpandas. No wrapper. - Effect-oriented loops (`notes.md` §28, `mappings.md` §19): Ward's ZIO client keeps Jev as the outer Choice and the handler as the effect. Not Effect.ts. GLiNER author: GLiNER2 "like jev" is GLiGuard diff --git a/README.md b/README.md index 1bc5b14..57b6aaa 100644 --- a/README.md +++ b/README.md @@ -31,6 +31,7 @@ never launder a Noul as a proof. - `.agents/skills/augustus/references/judgment-class.md` — the class (Jev exemplar, not monopoly): open heads (Laya, kev, encoder DeBERTa, LoRA distill), constrained-AR (TypeAR, pcdServer), announced decision-model (Watch), + open multimodal RLCD (blackwood-rlcd; not Archer), GLiNER/GLiClass species (locate vs categorize vs local multi-head), listwise vs decision objectives, vision scoring, when-to-use axes, agent-architecture portents @@ -49,15 +50,16 @@ never launder a Noul as a proof. - `.agents/skills/augustus/references/applied-mappings.md` — context sieve, exact-text keep/drop, env triage, moderation/ranking, skill routing - `.agents/skills/augustus/references/faq.md` — "just classification", - stack replacement, Jev vs open head vs encoder vs LoRA vs constrained AR vs kev, - GLiNER vs GLiClass vs CLIP, LLM-as-judge, in-engine vs CLI store, + stack replacement, Jev vs open head vs encoder vs LoRA vs constrained AR vs kev vs blackwood, + wait-for-Archer, missing-other confident-wrong, GLiNER vs GLiClass vs CLIP, LLM-as-judge, in-engine vs CLI store, hard envelope (bitrate / planner), not-another-how-to - `.agents/skills/augustus/references/mappings.md` — classical-method mappings with boundaries, counterexamples, acceptance tests (including Hypothesis cards §6–§19 — promote only with a test that ran) - `.agents/skills/augustus/references/validation.md` — design gate, eval recipes, Jev-for-skills (routing, self-monitoring, testing, modularity, - frontmatter), and Eval & hill-climb (jevals hygiene + Harbor taskset) + frontmatter), and Eval & hill-climb (jevals hygiene + Harbor taskset; + open-jev-laya-bench as ECE/NLL/Brier bake-off exemplar) - `.agents/skills/augustus/references/boundary-audit.md` — existing-system insertion: fit test, opportunity map, smallest boundary, red flags - `.agents/skills/augustus/scripts/evaluate_decisions.py` — offline evaluator diff --git a/docs/ecosystem.md b/docs/ecosystem.md index f004a0d..12be887 100644 --- a/docs/ecosystem.md +++ b/docs/ecosystem.md @@ -26,6 +26,10 @@ weekdays. Jev is the densest public corpus, not the class monopoly. ### Agent harnesses & self-supervision - **Kevthetech143/super-jev** — domain-independent loop: observe → questions → decide → **permit (independent of confidence)** → execute (idempotency key) → verify → JSONL replay. - **AntonioCoppe/jev-harness** — policy + confidence gate + shadow mode + offline eval CLI asserting on the **action**; 24-row filter 48.9s (Claude CLI) vs 1.3s Jev. Harbor/jevals-adjacent practice. `notes.md` §33, §44. +- **khordoo/jev-reflex-autonomy-lab** — S1 Jev reflex keeps control; optional S2 planner is one-use advice on low confidence. Experimental viz, not a flight controller. `notes.md` §46. +- **perixtar/jev-e2e** — NL cases; Jev selects observed controls; Playwright independently checks. PASS/FAIL/BLOCKED. Alpha. `notes.md` §46. +- **Wany-i/jev-decision-layer** — business decision tool; caller names the judgment; `gate` is part of the result. Unofficial. `notes.md` §46. +- **yalindogusahin/jevpandas** — pandas semantic index; noul/choice/score; LICENSE absent this pass. `notes.md` §46. - **Friedjof/jev-mobile** — durable Android worker + Mobile MCP; Jev sees prevalidated candidates only. `notes.md` §33. - **jcpsimmons/jev-macos-loop** — Apple-silicon computer-use; local OmniParser/OCR/AX; text-only Jev. Finder demo independently verified. - **rajdhakad9826/routeKit** — Jev estimates task requirements; policy engine selects the LLM. Jev does not pick the model. @@ -42,15 +46,17 @@ Per-keystroke launchers (104ms median, sequence-tagged staleness), firehose mode - **arnabgho/rlcd-lite** — GRPO + Brier proper-scoring-rule reward → calibrated decisions; binary reward doesn't calibrate. - **stephanj/parallelConstraintDecoding** — whole JSON schema of booleans/enums in two forward passes (prefill → parallel masked fields). - **Foadsf/jev-for-engineers**, **AbdelStark/jev-benchmarks**, **BrendanH18/jev-lab** — measurement discipline and cost/latency visibility. -- **dayhaysoos/jevals** — local MIT workbench: labeled cases (Noul / Choice / Score), compare runs, WebMCP + agent skill. Empirical acceptance-test surface for Hypothesis mapping cards; complements `evaluate_decisions.py`. Not affiliated with TypeSafe. Pointer: `research/notes.md` §24. Hygiene and the Harbor substrate: `validation.md` Eval & hill-climb (`notes.md` §40). +- **dayhaysoos/jevals** — local MIT workbench: labeled cases (Noul / Choice / Score), compare runs, WebMCP + agent skill. Empirical acceptance-test surface for Hypothesis mapping cards; complements `evaluate_decisions.py`. Not affiliated with TypeSafe. Pointer: `research/notes.md` §24. Hygiene and the Harbor substrate: `validation.md` Eval & hill-climb (`notes.md` §40). Shared bake-off exemplar this hour: [`pngwn/open-jev-laya-bench`](https://huggingface.co/datasets/pngwn/open-jev-laya-bench) (ECE/NLL/Brier; LLM-as-judge is not the score; `notes.md` §46). - **jeiel85/jevscope** — local-first visual debugger + JSONL regression for Choice/Score/Noul; compare two definitions; policy buckets are JevScope-derived. Sits next to jevals. Pointer: `research/notes.md` §25. ### Local / open heads & GLi\* species - **GLiNER / GLiNER2.5 / GLiClass** — species map: locate spans vs categorize the sequence vs local multi-head (fastino-ai GLiNER2.5 CPU-first). Peer of Jev, not a footnote. `references/judgment-class.md`. Author primary source: GLiNER2 "like jev" is schema-conditioned categorize (GLiGuard), not a Noul (`notes.md` §28). 36× Browser Use claim is a tweet (`notes.md` §25). - **GLiGuard** (fastino-ai) — 0.3B GLiNER2 encoder, checkpoint `fastino/gliguard-LLMGuardrails-300M`. One bidirectional pass over a safety schema. Same interface shape as batched questions; different objective. Not a Jev weight clone. `judgment-class.md`; `notes.md` §30. - **DECRUX9812/openjev-lm** — Qwen2.5-0.5B+LoRA distilled from hosted Jev answers; 65/70 = 92.9% on 70 hand-labelled rows (one annotator, one domain, one seed) overnight on 6 vCPU, $0/call. Its 98.1% on fresh rows is teacher *agreement*, not gold. Receipts pattern: `notes.md` §25, §44. -- **convaiinnovations/laya** — open Choice/Score/Noul head, text-only, 512 tok. Companion packaging this hour: [`laya-typed-decisions`](https://huggingface.co/convaiinnovations/laya-typed-decisions) (421.3M, acc 0.766 / Brier 0.066 unverified). `notes.md` §18, §42. +- **convaiinnovations/laya** — open Choice/Score/Noul head, text-only, 512 tok. Companion packaging this hour: [`laya-typed-decisions`](https://huggingface.co/convaiinnovations/laya-typed-decisions) (421.3M, acc 0.766 / Brier 0.066 unverified). Shared bake-off: [`pngwn/open-jev-laya-bench`](https://huggingface.co/datasets/pngwn/open-jev-laya-bench) (26+9 tasks, 11959 items; ECE/NLL/Brier; not TypeSafe Jev vs Laya). `notes.md` §18, §42, §46. - **jaredpalmer/kev** — Qwen2.5-0.5B LoRA + pointer readout; Apache-2.0; `kev-0.5b` 38 MB on GitHub release v0.1.0. Runnable Archer reconstruction (`POST /v1/systemone`). Public gold, not a Jev teacher. Isolation exact; ID ECE 0.065 (0.031 after T); acc 0.799 / 1,350. Not a knowledge/frontier substitute. `notes.md` §45. +- **BlackwoodAI/blackwood-rlcd** — open multimodal RLCD (image-text-to-text), Jev-compatible shim, CC BY-NC 4.0. Screenshot + marked candidates → Choice. Card: web acc 0.907 vs Jev 1.13 text-only 0.480; letter-shuffle 0.133 vs 0.587; ECE 0.037; ~200 ms H100. Jev still leads general text 0.850 vs 0.786. Not Archer Watch. `notes.md` §46. +- **Foodoo1/Qwen3-14B-RLCD-Decision-LoRA** — decision-token QLoRA on Qwen3-14B under parallel constrained decoding. Held-out 200-case / 4-field: fraud_risk 64→95%, overall 85.2→98.8% at ~234 ms. Synthetic; not a financial product. `notes.md` §46. - **zmtomorrow/TypeAR** — constrained autoregressive decoding surface: typed fields on a pretrained open model, no retraining. Not a proper-scoring decision head. `research/notes.md` §32. - **stephanj/pcdServer** — native Parallel Constrained Decoder (C++20, llama.cpp GGUF, Apple+Linux). TypeAR-class serving: 2–256 enums, 1–63 parallel fields; softmax over allowed values is not a Noul. `notes.md` §42. - **com-kotobalabs/open-jev-deberta-v3-large** — encoder open-jev, DeBERTa-v3-large 434M, apache-2.0, public gold (not a Jev teacher). In-domain ECE 0.022; OOD acc 0.854→0.690. `notes.md` §33. @@ -59,10 +65,11 @@ Per-keystroke launchers (104ms median, sequence-tagged staleness), firehose mode - **Archer Hume open decision-model** — **Watch.** Qwen3.8 27B dense, 265k, multimodal no audio; one forward pass locally once AR is removed. Driver: AU healthcare data-residency. No Hub weights this pass. `notes.md` §31–§33. Runnable *architecture* productization (0.5B, not that drop): **jaredpalmer/kev**. `notes.md` §45. - **bespokelabsai/nimble** — open recipe: contrastive hard labels, not a Jev distill. Model card Apache-2.0 LoRA `bespokelabs/Bespoke-Nimble-9B` on Qwen3.5-9B (repo license absent). Their 324-row holdout is a named receipt, not a ranking. `research/notes.md` §35. - **mmastrac/djev-spark** — DiffusionGemma 26B-A4B NVFP4, Jev-shaped decisions, images as an extension. Third compute graph. Interface claim, not a win over a decision head. `research/notes.md` §36. -- **Perception then judgment** — SAM 3.1 (masks and tracks) or ASR (a transcript) are upstream producers, not the perceive species. System One on that state is decide. Composition, not native omni. Information dies at the interface. `research/notes.md` §39. +- **Perception then judgment** — SAM 3.1 (masks and tracks) or ASR (a transcript) are upstream producers, not the perceive species. System One on that state is decide. Composition, not native omni. Information dies at the interface. Open multimodal *decide* that ships now: blackwood-rlcd (not Archer). `research/notes.md` §39, §46. ### Structural prove ∩ remainder -- **thevibeworks/jevgate** — Proven / Refused / Unknown; cannot block; 0/59 unsafe unasked held-out. Allowlist ∩ System One. +- **thevibeworks/jevgate** — Proven / Refused / Unknown; cannot block; 0/59 unsafe unasked held-out. Allowlist **proves** read-only verbs; Jev judges only unlisted. Allowlist ∩ System One. +- **suraj-phanindra/wellposed** — lint the Jev request before it comes back confidently wrong. Missing `other` → confidence 1.00 on a wrong Choice; gating cannot catch it. `tenbin` owns the lint skill. `notes.md` §46. - **misbahsy/doc-router** — page OCR router: 155→87 billed, 1.74× $ on 19 docs / 155 pages. Same sandwich. ### Agent harnesses extras (this hour) diff --git a/research/archive/findings.md b/research/archive/findings.md index d154caf..7b0e3d8 100644 --- a/research/archive/findings.md +++ b/research/archive/findings.md @@ -705,6 +705,43 @@ Cross-repo addition: (av) the trained decision-only path now has a shipped API-compatible reconstruction (kev); Watch remains the 27B announcement. +## Batch #29 (2026-09-18, ~14:03 Boise hourly) + +Note: `research/notes.md` §46. Docs-only. Folded into PR #2. Archer +27B drop still Watch (Hub empty). No invented metrics. + +- **blackwood-rlcd (Empirical as named vendor receipt; Hypothesis on + your labels).** Open multimodal RLCD, CC BY-NC, Jev-compatible shim. + Web 0.907 vs Jev 1.13 text-only 0.480; letter-shuffle 0.133 vs 0.587; + ECE 0.037; ~200 ms H100. Jev still leads general text 0.850 vs 0.786. + Screenshot vs Jev-text is not the same input. Omni decide can ship + without waiting for Archer. Soft judgment over marked pixel + candidates; code clicks. +- **open-jev-laya-bench (Empirical as that named receipt; not a Jev + ranking).** 26+9 tasks, 11959 test / 3269 cal. Macro acc Δ +0.023 + [+0.013,+0.032] neutral, +0.229 [+0.198,+0.262] home. ECE/NLL/Brier. + LLM-as-judge is not the score. Harbor/jevals practice in the wild. +- **Foodoo1 decision-token QLoRA (Empirical as 200-case receipt).** + Train the single decision token under parallel constrained decode. + fraud_risk 64→95%, overall 85.2→98.8% at ~234 ms/4-field. Synthetic; + not a financial product. Softmax ≠ Noul. +- **jevgate frame (already §25):** allowlist *proves*; Jev judges only + unlisted; fail-open (cannot block). +- **wellposed (Empirical as request-lint recipe):** missing `other` → + confidence 1.00 wrong; gating cannot catch it. Broken state paths. + `tenbin` owns the lint skill. +- **jev-reflex-autonomy-lab:** S1 keeps control; optional S2 one-use + advice. No metrics. License null this pass. +- **MED:** jev-decision-layer (gate is part of the result); jev-e2e + (Playwright checks; confident model cannot substitute); jevpandas + (dataframe semantic index; LICENSE 404). + +Cross-repo addition: (aw) omni decide is a shipped open head, not a +Watch-only hole; (ax) bake-off substrate with ECE/NLL/Brier in the +wild; (ay) decision-token LoRA is how you train constrained-AR, not a +new species; (az) confidence gating cannot catch a forced Choice. + + diff --git a/research/notes.md b/research/notes.md index 519f0ff..352d77e 100644 --- a/research/notes.md +++ b/research/notes.md @@ -2397,3 +2397,224 @@ labels. Cards: `judgment-class.md`; FAQ; `validation.md`; `mental-models.md`; `mixed-architecture.md`; `formal-methods.md` (pointer-softmax is still a sensor); `optimizer-integration.md`. No wrapper. + +## 46. 14:03 Boise hourly — open multimodal RLCD, bake-off substrate, decision-token LoRA (2026-09-18) + +America/Boise 14:03 = 20:03 UTC. Docs-only fold into PR #2 +(`cursor/augustus-store-envelope-00b4`). Archer Hume 27B drop still +**WATCH** (Hub authors `archerhume` / `4rcherhume` empty this pass; +user watch still ~2026-09-19). No invented metrics. Frames first, +not a hit list. Identity lock vs `typesafe-ai` / `tenbin` / +`decision-first` holds. No wrapper, no serve how-to. + +HTTP 200 / Hub fetch this pass: +[`BlackwoodAI/blackwood-rlcd`](https://huggingface.co/BlackwoodAI/blackwood-rlcd) +README; [`pngwn/open-jev-laya-bench`](https://huggingface.co/datasets/pngwn/open-jev-laya-bench) +`RESULTS.md` (no dataset README); [`Foodoo1/Qwen3-14B-RLCD-Decision-LoRA`](https://huggingface.co/Foodoo1/Qwen3-14B-RLCD-Decision-LoRA) +README; GitHub READMEs for +[`thevibeworks/jevgate`](https://github.com/thevibeworks/jevgate), +[`suraj-phanindra/wellposed`](https://github.com/suraj-phanindra/wellposed), +[`khordoo/jev-reflex-autonomy-lab`](https://github.com/khordoo/jev-reflex-autonomy-lab), +[`Wany-i/jev-decision-layer`](https://github.com/Wany-i/jev-decision-layer), +[`perixtar/jev-e2e`](https://github.com/perixtar/jev-e2e), +[`yalindogusahin/jevpandas`](https://github.com/yalindogusahin/jevpandas). + +### HIGH + +1. **[`BlackwoodAI/blackwood-rlcd`](https://huggingface.co/BlackwoodAI/blackwood-rlcd)** + (Hub, 2026-09-18 v1; `pipeline_tag: image-text-to-text`; CC BY-NC 4.0; + likes=1 this pass). **Open multimodal RLCD decision model.** + Screenshot or text in; Choice / Score / Boolean out in one prefill; + Jev-compatible `/v1/systemone` shim. Options listed as letters; + answer from option-letter logits at the last position; temperature- + calibrated; no sampling; `output_tokens` always 0. Not Archer's 27B + drop. Not CLIP/SigLIP (those are vision *scorers*). Same *decide* + species as Jev / Laya / kev, with image-in. + + **Architecture (load-bearing).** Omni perception→decision can ship + **without waiting for Archer**. Soft judgment over **pixel + candidates code already marked** (letters drawn on the screenshot; + criteria keyed by those letters) inside a deterministic click/act. + Specialist composition (SAM / OCR / AX → text → Jev) still exists + (`§39`); this is the shared multimodal System One that card was + waiting for. Information still dies at the *act*: the model picks a + letter; code clicks. A Noul is still not a proof. CC BY-NC: research + / personal, not a commercial drop-in. Do not copy vLLM flags, the + shim, or curl bodies into skill cards. + + **Evidence (model card; paired per item; not re-run; Empirical as + their named receipt).** Web element choice, 300 held-out steps: + acc **0.907** vs Jev 1.13 text-only **0.480**; letter-shuffle flip + **0.133** vs **0.587** (lower better); ECE **0.037** vs **0.091**; + latency **~200 ms** (1×H100) vs 441 ms (Jev via OpenRouter). + 4-lettering (four orderings averaged, 4× compute) acc **0.953**. + Desktop held-out acc **0.76** vs **0.654**. NL predicates over + records **0.973** vs **0.965**. Tetris lines / 60 pieces **14.2** vs + **13.6**. General text, 85 public sets / 8,456 items: blackwood + **0.786**, Jev **0.850** — Jev still leads. Domain sets 27 / 12,746: + **0.710** vs **0.703**. Pixel rows use a randomized viewport crop. + Jev is text-only by design: screenshot rows compare screenshot + input with Jev's *text* input on the same steps. **Hypothesis** that + it substitutes for Jev on *your* labels. Boolean questions on skewed + sets can be over-confident (their limit). Cards: `judgment-class.md` + (family, holes, when-to-use, dedicated card, perception rewrite); + FAQ wait-for-Archer; `validation.md` letter-shuffle; `formal-methods.md` + one sentence; `mixed-architecture.md` gallery; `applied-mappings.md` §2. + +2. **[`pngwn/open-jev-laya-bench`](https://huggingface.co/datasets/pngwn/open-jev-laya-bench)** + (`RESULTS.md` this pass; no dataset README). Shared bake-off: + [`pngwn/system-one-qwen3.5-4b-scorer`](https://huggingface.co/pngwn/system-one-qwen3.5-4b-scorer) + `@ e6464dce` vs [`convaiinnovations/laya`](https://huggingface.co/convaiinnovations/laya) + `@ 7c76b622`. **26 neutral + 9 home** tasks, **11,959** test items, + **3,269** cal items. Hardware: NVIDIA A100-SXM4-80GB. Config chosen + on `cal` (scorer vs `laya_text`). **Not TypeSafe Jev vs Laya** — do + not launder this table into a proprietary-Jev ranking. Laya's + published in-task acc/ECE sit on unpublished sets and are context, + not like-for-like. + + **Harbor / jevals practice in the wild.** Measurement substrate + exemplar: accuracy + **ECE / NLL / Brier** (and acc@50% coverage), + held-out `test`, temperature chosen on `cal`, leave-one-task-out, + prompted-base / prompted-instruct arms on the same 4B, per-task T + vs shipped T. **LLM-as-judge is not the primary System One score.** + That is the eval path the skill already wanted (`validation.md` + Eval & hill-climb; `notes.md` §40). Do not copy the scoring scripts. + + **Evidence (`RESULTS.md`; not re-run; Empirical as that named + receipt).** Neutral, shipped T: System One 4B macro acc **0.760**, + Laya **0.737**, Δ **+0.023 [+0.013, +0.032] \*** (interval excludes + 0). Neutral ECE-10 **0.074** vs **0.122**; macro NLL **0.569** vs + **1.272**; macro Brier **0.324** vs **0.392**. Home (the 4B's own + held-out split): **0.715** vs **0.486**, Δ **+0.229 [+0.198, +0.262] + \***. Home ECE-10 **0.072** vs **0.272**; NLL **0.697** vs **3.836**; + Brier **0.363** vs **0.759**. Ahead on 13/26 neutral and 8/9 home. + Neutral prompted-instruct (same 4B, ≤26 options) macro acc **0.763** + vs fine-tune **0.760**, Δ **+0.003 [-0.004, +0.011]** — interval + includes 0; do not slogan "fine-tune always wins zero-shot." Home + prompted arms drop banking77 / ticket-queue (enum >26). Latency + one-question p50 on the same GPU: tweet_emotion Laya 27.9 ms vs + scorer 76.7 ms; home_banking77 (K=77) 30.0 vs 382.8 ms. **Hypothesis** + as a ranking of *your* head. Cards: `validation.md` bake-off; FAQ + LLM-as-judge; `judgment-class.md` Laya companion. + +3. **[`Foodoo1/Qwen3-14B-RLCD-Decision-LoRA`](https://huggingface.co/Foodoo1/Qwen3-14B-RLCD-Decision-LoRA)** + (Apache-2.0 adapter; base `Qwen/Qwen3-14B`; PEFT). **Decision-token + QLoRA** under **parallel constrained decoding**: one prefill, KV + broadcast across fields, logit slice over candidate first tokens — + the TypeAR / pcdServer / `stephanj/parallelConstraintDecoding` + inference pattern (`findings.md` batch #1). Loss computed **only on + the single decision token** of each field. The base already hits + easy fields (language/sentiment 100%); reasoning-heavy fields fail + because the decision happens at one token with no room to think. + Pattern, not a fraud product: train the decision token, do not + fine-tune generated prose. Evaluated with + [`harshatheg/Qwen-2.5-1B-RLCD`](https://huggingface.co/harshatheg/Qwen-2.5-1B-RLCD) + inference code (that Hub repo still has no weights — stock Qwen + + custom code; already archived). Synthetic fraud-triage schema; + **do not use for real financial decisions** (their limit). Do not + copy the prompt or LoRA flags into skill cards. + + **Evidence (card; held-out 200-case, 4 fields, 4-bit NF4, RTX 3090; + not re-run).** fraud_risk (4-way) **64.0% → 95.0%**; block_account + **77.0% → 100%**; language and sentiment stay **100%**; overall + (800 decisions) **85.2% → 98.8%**. Mean latency **~234 ms** per + case (all 4 fields in one broadcast). Remaining errors all + LOW→ELEVATED (conservative); zero high-risk judged low. Train: + 16,608 single-token examples from 4,152 synthetic cases (648 + templates × 72 messages EN/ZH/ES/JA), 1 epoch, ~97 min on one 3090. + **Empirical** as that named receipt. **Hypothesis** as a general + recipe on *your* schema. Cards: `judgment-class.md` constrained-AR; + `methods-catalog.md`; FAQ open-weights vs constrained decode. + +4. **[`thevibeworks/jevgate`](https://github.com/thevibeworks/jevgate)** + — already §25 / `mappings.md` §18. This hour the *mental model*, + matching the GitHub description: **an allowlist proves** every verb + is a listed read-only tool; **Jev judges only unlisted** leftovers; + **fail-open (cannot block)**. Hard envelope in code + soft judgment + only on the residual. Proven runs in 4 µs and never leaves the box; + Refused never asks the model (so a comment cannot talk a writer + through); Unknown is five Nouls, admit iff every p < 0.2. Jev on + `/bin/ls` alone leaks (0.04) — that is why it is the third tier. + Real-traffic counts already on the README (116,979 Bash calls: + 26.4% proven; 16.2% reach Jev; 47 admitted, all read-only by hand) + stay author-reported; do not re-promote 0/59 as if new. MIT. 0★ + this pass. No rewrite of the 0.2 threshold as a constant. Cards: + `mappings.md` §18 "proves"; FAQ allowlist-then-judge; methods-catalog. + +5. **[`suraj-phanindra/wellposed`](https://github.com/suraj-phanindra/wellposed)** + (MIT, JS, 0★, created 2026-09-17). **Lint the request before it + comes back confidently wrong.** Neighbor-skill identity: `tenbin` + still owns design-time lint/measure; this is an Empirical *recipe* + of that hole, not a second Augustus skill and not a wellposed + how-to. Standing Choice contract (`SKILL.md`): probabilities are + conditional on the offered set; absent candidates can never be + chosen; add `other` where coverage is open. wellposed is the + receipt that **violating that contract is silent**. + + **Evidence (README; not re-run).** 40 generated requests: **0/40** + syntax errors (API validation already covers that); **16/40 (40%)** + asked something that did not make sense; **0/11** list-questions + included a none-of-the-above. Live probe: unsubscribe email, four + department options, no `other` → `"support issue"` **confidence + 1.00**; with `other` → `"other"` confidence 0.93. Overlapping + `angry`/`furious` collapsed confidence to **0.19** — loud; ordinary + gates catch it. Forced wrong Choice is quiet. Structural lint (35 + rules, offline) then optional jev-on-jev semantic layer. Labeled + corpus: recall **22/26 = 85%**, precision **22/24 = 92%** (computed + live from the corpus). Honest limits: one-model labels, so 40% is a + floor; 2026-09-18 adversarial audit added 36 items because the + metric could not see ordinary-English false positives. Broken state + paths (`ticket.assigned_agent.name` with no such path) are + *provably* wrong. Confidence gating **cannot** catch a forced + Choice. Cards: `question-design.md` diagnosis; validation gate #2; + FAQ; tenbin identity lock. + +6. **[`khordoo/jev-reflex-autonomy-lab`](https://github.com/khordoo/jev-reflex-autonomy-lab)** + (TypeScript, 0★, created 2026-09-18T19:46Z; GitHub license null this + pass). Interactive multi-drone lab. **S1 Jev reflex keeps control**; + optional S2 planner (OpenRouter / GLM 5.3) is **one-use advice** on + low confidence. Jev does not pause while the planner responds. S2 + does not fly the drone. Kahneman row already taught the split + (`toolbox-mapping.md`: S2 proposes, S1 discriminates; never the + reverse) — this is that split as a control loop, not a flight + controller. Experimental visualization; mock mode without keys; + live fleet success varies. No metrics to promote. Do not copy the + adapter, `.dev.vars`, or ports. Cards: `mixed-architecture.md` dual + orchestration; `agent-self-assessment.md`; toolbox Kahneman row. + +### MED (pointers, not cards of their own) + +7. **[`Wany-i/jev-decision-layer`](https://github.com/Wany-i/jev-decision-layer)** + (MIT, Python stdlib, 0★). Wrap the decision model as a **business + decision tool**: caller names the *judgment*, not the model. + `decide(name, fields)` → outcome + confidence + **`gate` (part of + the result)**. Registry JSON; hard guards in code; `other` required + on Choice. OpenRouter `POST /api/alpha/decisions` (not + `chat/completions` — that 400s). Text-only: screenshots must be + textualized — contrast with blackwood. 28 offline tests, no key. + Unofficial. Do not copy the registry or MCP install. + +8. **[`perixtar/jev-e2e`](https://github.com/perixtar/jev-e2e)** + (MIT, TypeScript, alpha, 0★). Natural-language cases; Jev selects + observed controls; **Playwright executes and independently checks + expectations**. Verdicts PASS / FAIL / BLOCKED. A completed + navigation or a confident model response **cannot substitute for + checked expectations**. Harbor/jevals practice on a browser taskset + (score on the task; harness rolls out). Alpha: controlled demo does + not establish reliability across arbitrary sites. Do not copy CLI + flags or ports. + +9. **[`yalindogusahin/jevpandas`](https://github.com/yalindogusahin/jevpandas)** + (Python; GitHub LICENSE 404 this pass). pandas frame as the store: + `evaluate` / `filter` / `classify` / `score` / batched `ask`. + Classify example includes `other`. Failures never become negative + predictions. No generative chat, no joins, no training on review + labels. Store-as-semantic-index cousin of jevql / sqlite-jev + (`mappings.md` §4) over a dataframe instead of SQL. Samples in + `data/` are synthetic. Do not copy the client. + +Cards: `judgment-class.md`; `validation.md`; `question-design.md`; +`mappings.md` §18 / §4; `mixed-architecture.md`; `faq.md`; +`agent-self-assessment.md`; `mental-models.md`; `formal-methods.md`; +`applied-mappings.md` §2; `methods-catalog.md`; `toolbox-mapping.md`. +No wrapper. diff --git a/research/refresh-log.md b/research/refresh-log.md index 77b1a3e..d728ad4 100644 --- a/research/refresh-log.md +++ b/research/refresh-log.md @@ -391,5 +391,25 @@ `optimizer-integration.md`; SKILL.md path + identity lock. - notes.md §45; sources.json; findings.md batch #28. No wrapper. +## 2026-09-18 20:03 UTC — ~14:03 Boise hourly fold + +- Folded into open PR #2 (`cursor/augustus-store-envelope-00b4`). + Docs-only. Archer 27B drop still **WATCH** (Hub empty; user watch + ~2026-09-19). +- HIGH: blackwood-rlcd (open multimodal RLCD, CC BY-NC, Jev-compatible + shim; web 0.907 vs Jev text-only 0.480; letter-shuffle 0.133 vs + 0.587; ECE 0.037; ~200 ms H100; Jev still leads general text 0.850 + vs 0.786). open-jev-laya-bench (26+9, 11959 items; ECE/NLL/Brier; + Δ +0.023 / +0.229; LLM-as-judge is not the score). Foodoo1 + decision-token QLoRA (64→95% / 85.2→98.8% at ~234 ms; synthetic). + jevgate frame (allowlist proves; fail-open). wellposed (missing + other → confident wrong). jev-reflex-autonomy-lab (S1 keeps control). +- MED: jev-decision-layer, jev-e2e, jevpandas. +- Cards: judgment-class, validation, question-design, mappings §18/§4, + mixed-architecture, faq, agent-self-assessment, mental-models, + formal-methods, applied-mappings, methods-catalog, toolbox. +- notes.md §46; sources.json; findings.md batch #29. No wrapper. + + diff --git a/research/sources.json b/research/sources.json index 5b7397f..9729a0d 100644 --- a/research/sources.json +++ b/research/sources.json @@ -1,6 +1,6 @@ { "refresh_cadence": "hourly", - "retrieved": "2026-09-18T19:54Z", + "retrieved": "2026-09-18T20:03Z", "sources": [ { "kind": "docs", @@ -1249,6 +1249,60 @@ "title": "jaredpalmer/kev MODEL_CARD.md", "url": "https://github.com/jaredpalmer/kev/blob/main/MODEL_CARD.md", "note": "Research prototype; ID ECE 0.065 / 0.031 after T=1.47; isolation probes. notes.md \u00a745." + }, + { + "kind": "huggingface", + "title": "BlackwoodAI/blackwood-rlcd", + "url": "https://huggingface.co/BlackwoodAI/blackwood-rlcd", + "note": "Open multimodal RLCD; CC BY-NC; Jev-compatible shim. Web acc 0.907 vs Jev 1.13 text-only 0.480; letter-shuffle 0.133 vs 0.587; ECE 0.037; ~200ms H100. Jev still leads general text 0.850 vs 0.786. Not Archer. notes.md \u00a746." + }, + { + "kind": "huggingface", + "title": "pngwn/open-jev-laya-bench", + "url": "https://huggingface.co/datasets/pngwn/open-jev-laya-bench", + "note": "Shared bake-off Qwen3.5-4B scorer vs Laya 421M. 26+9 tasks, 11959 test items. ECE/NLL/Brier. Macro acc \u0394 +0.023 neutral, +0.229 home. Not TypeSafe Jev vs Laya. notes.md \u00a746." + }, + { + "kind": "huggingface", + "title": "Foodoo1/Qwen3-14B-RLCD-Decision-LoRA", + "url": "https://huggingface.co/Foodoo1/Qwen3-14B-RLCD-Decision-LoRA", + "note": "Decision-token QLoRA under parallel constrained decoding. fraud_risk 64\u219295%, overall 85.2\u219298.8% at ~234ms/4-field. Synthetic. Apache-2.0 adapter. notes.md \u00a746." + }, + { + "kind": "github", + "title": "thevibeworks/jevgate (14:03 Boise frame)", + "url": "https://github.com/thevibeworks/jevgate", + "note": "Allowlist proves read-only verbs; Jev judges only unlisted; fail-open cannot block. Already \u00a725; frame this hour. MIT. notes.md \u00a746." + }, + { + "kind": "github", + "title": "suraj-phanindra/wellposed", + "url": "https://github.com/suraj-phanindra/wellposed", + "note": "Lint Jev requests. Missing other \u2192 confidence 1.00 wrong; gating cannot catch. MIT. tenbin owns the lint skill. notes.md \u00a746." + }, + { + "kind": "github", + "title": "khordoo/jev-reflex-autonomy-lab", + "url": "https://github.com/khordoo/jev-reflex-autonomy-lab", + "note": "S1 Jev reflex keeps control; optional S2 one-use advice. Experimental. License null this pass. notes.md \u00a746." + }, + { + "kind": "github", + "title": "Wany-i/jev-decision-layer", + "url": "https://github.com/Wany-i/jev-decision-layer", + "note": "Business decision tool; gate is part of the result; other required. MIT. Unofficial. notes.md \u00a746." + }, + { + "kind": "github", + "title": "perixtar/jev-e2e", + "url": "https://github.com/perixtar/jev-e2e", + "note": "Jev selects observed controls; Playwright independently checks. PASS/FAIL/BLOCKED. MIT. Alpha. notes.md \u00a746." + }, + { + "kind": "github", + "title": "yalindogusahin/jevpandas", + "url": "https://github.com/yalindogusahin/jevpandas", + "note": "pandas semantic index; noul/choice/score; LICENSE 404 this pass. notes.md \u00a746." } ] } From 8d61cd6ac781f8642787c52549f7f862f057438e Mon Sep 17 00:00:00 2001 From: Cursor Agent Date: Fri, 18 Sep 2026 20:42:32 +0000 Subject: [PATCH 04/43] Fold Abide as productized soft-rule preference lint Cite README + replay receipt only. Mental models: soft rules vs linter, edit vs turn observation window, banded fail-open, rubric as artifact. jev-pref stays the contract; rh-guard stays eval-integrity. Not a hit list, not multimodal, no hook how-to. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com> --- .agents/skills/augustus/SKILL.md | 6 +- .../references/agent-self-assessment.md | 9 +- .../references/composition-algebra.md | 7 +- .agents/skills/augustus/references/faq.md | 21 ++- .../augustus/references/formal-methods.md | 3 +- .../augustus/references/judgment-class.md | 7 +- .../skills/augustus/references/mappings.md | 19 ++- .../augustus/references/mental-models.md | 1 + .../augustus/references/methods-catalog.md | 2 +- .../augustus/references/mixed-architecture.md | 27 +++- .../augustus/references/question-design.md | 1 + .../augustus/references/toolbox-mapping.md | 2 +- .../skills/augustus/references/validation.md | 15 ++- CHANGELOG.md | 13 ++ README.md | 5 +- docs/ecosystem.md | 4 +- research/archive/findings.md | 26 ++++ research/notes.md | 125 ++++++++++++++++++ research/refresh-log.md | 19 +++ research/sources.json | 18 ++- 20 files changed, 300 insertions(+), 30 deletions(-) diff --git a/.agents/skills/augustus/SKILL.md b/.agents/skills/augustus/SKILL.md index 1c059f1..b9dd905 100644 --- a/.agents/skills/augustus/SKILL.md +++ b/.agents/skills/augustus/SKILL.md @@ -1,6 +1,6 @@ --- name: augustus -description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill, open multimodal RLCD / blackwood), announced open decision-model (Watch — still not landed), constrained-AR (TypeAR, pcdServer, decision-token LoRA), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"wait for Archer vs open multimodal RLCD\", \"screenshot/DOM candidates → typed Choice\", \"eval path\", \"jevals\", \"Harbor taskset\", \"shared bake-off ECE/NLL/Brier\", \"LLM-as-judge is not the System One score\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge (allowlist *proves*; fail-open cannot block), \"missing other → confident wrong Choice\", \"lint the Jev request\", \"S1 reflex keeps control / optional S2 one-use advice\", TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, or \"formally verify with Jev\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." +description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill, open multimodal RLCD / blackwood), announced open decision-model (Watch — still not landed), constrained-AR (TypeAR, pcdServer, decision-token LoRA), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"wait for Archer vs open multimodal RLCD\", \"screenshot/DOM candidates → typed Choice\", \"eval path\", \"jevals\", \"Harbor taskset\", \"shared bake-off ECE/NLL/Brier\", \"LLM-as-judge is not the System One score\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge (allowlist *proves*; fail-open cannot block), \"missing other → confident wrong Choice\", \"lint the Jev request\", \"S1 reflex keeps control / optional S2 one-use advice\", \"soft AGENTS.md rules vs linter\" (Abide / jev-pref), \"edit-phase vs turn-phase observation window\", banded confidence fail-open preference lint, TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, or \"formally verify with Jev\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." license: MIT metadata: version: 0.3.0 @@ -54,7 +54,7 @@ classical method you already trust, substitute it, classify the win "paraphrase brittleness", "allowlist then judge", "TOCTOU-of-Noul", "Jev inside the database / sqlite-jev", "Jev picks bitrate / join order / the model", "wait for Archer", "lint the request / missing - other", or "screenshot Choice / omni System One": + other", "soft AGENTS.md rules vs the linter", or "screenshot Choice / omni System One": read `references/faq.md`, then `references/mental-models.md`, then `references/mixed-architecture.md`, then @@ -125,7 +125,7 @@ classical method you already trust, substitute it, classify the win | Moderation and ranking | Hold-before-publish vs graded rerank; fail policy per action | `references/applied-mappings.md#4-moderation-and-ranking` | | Skill / tool routing | Choice over a closed catalog + whether-anything-fits; code dispatches | `references/applied-mappings.md#5-skill--tool-routing` | | Expensive observation router | Structural prove (text layer) ∩ remainder Noul (needs OCR?) | `references/applied-mappings.md#6-expensive-observation-router` | -| Agent preference lint / semantic gates | Project-defined rules as criteria; provider classifies evidence; code maps outcome | `references/mixed-architecture.md#preference-lint-and-gates` | +| Agent preference lint / semantic gates | Soft project rules as criteria; linter owns hard rules; one Score per rule on the diff; bands + fail-open; name the observation window (edit vs turn) | `references/mixed-architecture.md#preference-lint-and-gates` | | Dual orchestration (Jev ∩ LLM ∩ MCP) | Jev-as-tool vs Jev-as-outer-loop; schemas are exact state | `references/mixed-architecture.md#dual-orchestration-jev--llm--mcp` | | "It's just classification" / "not probabilistic programming" / stack-replacement FAQ | Typed judgment is a software primitive, not a new task; marginals are not a joint; Jev is not the only model | `references/faq.md` | | Feature engineering / multi-criteria analysis | Nouls + Score distributions as named features, weights in code | `references/mappings.md#1-semantic-judgments--features-and-explicit-utility` | diff --git a/.agents/skills/augustus/references/agent-self-assessment.md b/.agents/skills/augustus/references/agent-self-assessment.md index 05f055d..3257d34 100644 --- a/.agents/skills/augustus/references/agent-self-assessment.md +++ b/.agents/skills/augustus/references/agent-self-assessment.md @@ -79,8 +79,13 @@ When the "judge" is really "does this change violate a rule we already wrote?", do not ask Jev whether the code is good. Load `references/mixed-architecture.md#preference-lint-and-gates`. The transferable contract (`doeixd/jev-pref`): the project defines the rule, Jev classifies -visible evidence, code maps the outcome, the agent acts. Shadow-mode the gate -first; permit remains a separate axis from confidence. +visible evidence, code maps the outcome, the agent acts. +[`coldteadotai/abide`](https://github.com/coldteadotai/abide) is the +fuller productized path of that contract (compile / calibrate / tune / +replay; one Score per rule on the diff; bands + fail-open). Soft +rules → soft judgment; the linter owns hard rules. Shadow-mode the +gate first; permit remains a separate axis from confidence. +`notes.md` §47. ## Using Jev to test and optimize the skill suite itself diff --git a/.agents/skills/augustus/references/composition-algebra.md b/.agents/skills/augustus/references/composition-algebra.md index ebefb8b..34425bb 100644 --- a/.agents/skills/augustus/references/composition-algebra.md +++ b/.agents/skills/augustus/references/composition-algebra.md @@ -155,10 +155,11 @@ datasets — every gate is a per-dataset measurement (see validation.md). Jev as gate/selector/verifier *around* a generator, never instead of one. Cost-sensitive prefilter (drop chunks/lines/hunks before the LLM); tool/skill routing (Choice + fits-Noul, code dispatches); preference lint - (project-defined rules as criteria). Fail-open vs fail-closed is per + (project-defined *soft* rules as criteria; linter owns hard rules; + Abide is the productized path, `notes.md` §47). Fail-open vs fail-closed is per action — LlamaIndex Jev rerank fails open (keep retrieval order), select fails closed. Full card: `references/mixed-architecture.md`. -11. **Structural prove ∩ remainder judge** (jevgate, doc-router): code - (allowlist, text layer) decides the easy cases; typed questions only +11. **Structural prove ∩ remainder judge** (jevgate, doc-router, Abide): code + (allowlist, text layer, linter) decides the easy cases; typed questions only on leftovers; fail-open unless a real sandbox sits under. Full card: `mappings.md` §18. diff --git a/.agents/skills/augustus/references/faq.md b/.agents/skills/augustus/references/faq.md index f300df2..5c6bca7 100644 --- a/.agents/skills/augustus/references/faq.md +++ b/.agents/skills/augustus/references/faq.md @@ -179,7 +179,7 @@ labels on a GLiNER2 encoder, not Choice / Score / Noul). Empirical open encoder next to GLiClass; not a weight clone. "like jev" is discourse. A GLiGuard score is not a proof. LLM I/O safety is not a coding-agent tool gate (rh-guard for reward-hacking; jevgate shape for allowlist -∩ remainder). `judgment-class.md`. +∩ remainder; Abide for project soft rules on diffs). `judgment-class.md`. ## Can I threshold CLIP / SigLIP as a safety gate? @@ -295,7 +295,24 @@ database query already answers, **do not call a model** (`wotai-dev/typesafe-jev-tools`, `notes.md` §42). That is meta-VOI, not a hook tutorial. This hour's wording of the same sandwich: the allowlist **proves** read-only verbs; Jev judges only unlisted leftovers; -fail-open (cannot block) (`notes.md` §46). +fail-open (cannot block) (`notes.md` §46). Same family, different +remainder: a **linter proves** lintable rules; [Abide](https://github.com/coldteadotai/abide) +Scores residual soft AGENTS.md / CLAUDE.md rules; fail-open, banded +(`notes.md` §47). Soft judgment is never the sole hard veto. + +## Soft project rules — Jev or the linter? + +The linter owns what it can prove. Soft instruction-file rules +("no helper with one caller", "don't add what wasn't asked") are +residual judgment. Same layering as jevgate: structure first, typed +Score only on the remainder; **fail-open**. Name the observation +window (edit vs turn). Fix false positives in the rubric, not the +model. Productized path: [Abide](https://github.com/coldteadotai/abide); +earlier contract pointer: jev-pref. Complementary, not the same +product: [rh-guard](https://github.com/24601/rh-guard) (reward-hacking / +eval integrity). Request-shape lint still sits upstream (wellposed / +`tenbin`). `mixed-architecture.md`; `question-design.md`; `notes.md` +§47. ## Can confidence gating catch a forced wrong Choice? diff --git a/.agents/skills/augustus/references/formal-methods.md b/.agents/skills/augustus/references/formal-methods.md index f520d28..a6af3b7 100644 --- a/.agents/skills/augustus/references/formal-methods.md +++ b/.agents/skills/augustus/references/formal-methods.md @@ -71,7 +71,8 @@ TOCTOU-of-Noul (§5), not a discharged obligation. **What transfers** into a mixed stack: triage which counterexample, property, or failing seed a human looks at first; score whether a production trace resembles a spec behavior; lint an artifact against a -*named, project-written* rule (`mixed-architecture.md` preference lint). +*named, project-written* rule (`mixed-architecture.md` preference lint; +Abide is the productized path of that hole, `notes.md` §47). **What does not:** closing a proof obligation, replacing TLC/Apalache/ GNATprove, or treating "DST hasn't failed this week" as a safety case. diff --git a/.agents/skills/augustus/references/judgment-class.md b/.agents/skills/augustus/references/judgment-class.md index 625c73c..66938d0 100644 --- a/.agents/skills/augustus/references/judgment-class.md +++ b/.agents/skills/augustus/references/judgment-class.md @@ -119,7 +119,9 @@ below, next to the when-to-use table. **Surfaces.** GLiGuard is for LLM input/output safety. [rh-guard](https://github.com/24601/rh-guard) is a coding-agent reward-hack gate (README fetched this pass); jevgate is the - allowlist-then-judge shape (`mappings.md` §18). Different holes. + allowlist-then-judge shape (`mappings.md` §18); + [Abide](https://github.com/coldteadotai/abide) is project-instruction + soft rules on diffs (`mixed-architecture.md`). Different holes. Do not point one model at both, and do not copy a hook install here. **Aggregation is policy-in-code, already taught.** The README's @@ -644,7 +646,8 @@ the isolation pattern. Enum width and no abstention primitive are brittleness — compose with cost-sensitive abstention (`mappings.md` §2), paraphrase abstain (`mappings.md` §17), and a jevgate-shaped gate (`mappings.md` §18). rh-guard's README (HTTP 200) is a coding-agent -reward-hack gate, a different surface from this one and from GLiGuard. +reward-hack gate, a different surface from this one, from GLiGuard, +and from Abide's project soft-rule Scores (`notes.md` §47). Do not copy the hook install. [pcdServer](https://github.com/stephanj/pcdServer) (MIT, C++20, created diff --git a/.agents/skills/augustus/references/mappings.md b/.agents/skills/augustus/references/mappings.md index 9a7be08..d2dd507 100644 --- a/.agents/skills/augustus/references/mappings.md +++ b/.agents/skills/augustus/references/mappings.md @@ -60,7 +60,9 @@ Links: Score docs, composite-scoring pattern, autoresearch cookbook. choosing among act / decline / gather-evidence / escalate from the distribution, with thresholds owned by each action's consequences. For a calibrated binary probability with FP/FN costs: `t = C_FP/(C_FP+C_FN)`. -**Does not transfer**: universal thresholds (no magic 0.8); model probability +**Does not transfer**: universal thresholds (no magic 0.8 — a product +band such as Abide's ≥0.8 repair is *their* operating point, still +re-measured on your labels); model probability is not auto-calibrated on YOUR population — plot confidence vs accuracy on your data (**Contract**: confidence summarizes distribution shape, nothing more); Noul 0.5 ≠ medium-anything; top-Choice probability ≠ probability the @@ -73,7 +75,12 @@ else: act only if confidence > high bar, else confirm ``` **Example**: trading bot acts on high-confidence reads, stands down when the -book state is ambiguous (jev-trader `late → hold`). **Beyond SWE +book state is ambiguous (jev-trader `late → hold`). **Banded fail-open +(Empirical as a named product receipt):** +[Abide](https://github.com/coldteadotai/abide) on project soft rules: +≥0.8 repair in-session, 0.5–0.8 human note, <0.5 silence; hooks exit 0; +no key → the edit proceeds (`notes.md` §47). Soft judgment is never the +sole hard veto. **Beyond SWE (Hypothesis until labeled):** inbox reply/snooze/archive; "is this paper on-question?"; "call this lead / nurture / drop" — same act/abstain/ gather table, costs written in hours or dollars, threshold per *action*. @@ -606,7 +613,13 @@ one Jev call on the remainder. Empty state was self-contradictory — that is why the refuse-empty rule exists. [`affirmitv/bitrate-advisor`](https://github.com/affirmitv/bitrate-advisor) is the same sandwich on a live encoder: policy proves the cap; Jev -judges only inside it (`notes.md` §44). **Beyond SWE (Hypothesis):** +judges only inside it (`notes.md` §44). +[`coldteadotai/abide`](https://github.com/coldteadotai/abide) is the +same *family* on project instructions: the **linter proves** lintable +rules; Jev Scores only residual soft AGENTS.md rules; fail-open, banded +(`notes.md` §47). Different remainder from jevgate's unlisted verbs +and from rh-guard's eval-integrity hole — do not merge products. +**Beyond SWE (Hypothesis):** recipe book ∩ "does this leftover look done?"; labor-law allowlist ∩ hiring-fit Noul; SPF/DKIM pass ∩ phishing Noul on the body. **Counterexample:** Jev on `/bin/ls` as the first tier. **Test:** planted writers never diff --git a/.agents/skills/augustus/references/mental-models.md b/.agents/skills/augustus/references/mental-models.md index 3cb5397..8f0e34f 100644 --- a/.agents/skills/augustus/references/mental-models.md +++ b/.agents/skills/augustus/references/mental-models.md @@ -402,6 +402,7 @@ Use these as *existence proofs of a position*. Write your own card. | Inbox | reply / snooze / archive | urgency Noul + aboutness Choice | send, calendar | | Document / lab routing | which pages need the expensive observation | Noul on remainder after a text layer / recipe | local extract, merge order (**Empirical** as OCR-router *shape*) | | Shell / tool allowlist | unlisted remainder after a **proof** | five Nouls on unknown verbs | Proven/Refused in code; cannot block (**Empirical**: jevgate) | +| SWE | residual AGENTS.md / CLAUDE.md rules | one Score per named instruction-file rule | linter owns hard rules; bands + fail-open (**Empirical**: Abide replay, `notes.md` §47) | | Screenshot candidates → act | lettered elements code already marked | Choice over those letters | Click in code (**Empirical** as blackwood-rlcd *shape*; CC BY-NC) | | Moderation | hold before publish | hazard Nouls (**Empirical** as family) | block/review policy | | Phishing / fraud screen | hold vs deliver | SDT criterion on a Noul | blocklist, SPF/DKIM exact (**Hypothesis**) | diff --git a/.agents/skills/augustus/references/methods-catalog.md b/.agents/skills/augustus/references/methods-catalog.md index f2febe9..f88e33d 100644 --- a/.agents/skills/augustus/references/methods-catalog.md +++ b/.agents/skills/augustus/references/methods-catalog.md @@ -65,7 +65,7 @@ judgment component is new). | Method | Judgment-shaped component | Jev substitution | Stays in code | Status | |---|---|---|---|---| | Claim–evidence entailment (NLI) | supports / contradicts / not-established per claim–source pair | One Choice per pair + review flag; judge against the cited source text only | Quote extraction, citation graph, audit log | **Empirical recipe** (citation_check cookbook) | -| Spec vs artifact conformance (model checking *mindset*) | Property holds / violated / unverifiable for a named requirement | One Noul per requirement, batched; violated → named rule back into context (pi-warden shape). This is **not** TLC/Apalache/GNATprove | Requirement enumeration, enforcement, logging; the real checker if you have one | **Empirical recipe** (pi-warden: 6→0 rule breaks, 150 paired runs; jev-pref: YOU define the rule). Ownership split: `formal-methods.md` | +| Spec vs artifact conformance (model checking *mindset*) | Property holds / violated / unverifiable for a named requirement | One Noul/Score per requirement, batched; violated → named rule back into context (pi-warden / Abide shape). This is **not** TLC/Apalache/GNATprove | Requirement enumeration, enforcement, logging; the **linter** if the rule is lintable; the real checker if you have one | **Empirical recipe** (pi-warden: 6→0 rule breaks, 150 paired runs; jev-pref: YOU define the rule; Abide: productized compile/calibrate/tune/replay, `notes.md` §47). Ownership split: `formal-methods.md` | | Alloy finder vs Apalache / TLC | Which bound, which counterexample, is the property tautological? | Triage instances/CEs; never "this spec looks right" | Analyzer / SMT / explicit-state engine | **Hypothesis** as product; **Contract** as ownership (`formal-methods.md` §2) | | Type-checking analog | Does this planned call match the schema/operation/target? | Decomposed nouls over {request, schema, trace}; never trust a Jev pass as authorization | Real validation of operation+target in code | **Empirical recipe** (validation.md self-monitoring) | diff --git a/.agents/skills/augustus/references/mixed-architecture.md b/.agents/skills/augustus/references/mixed-architecture.md index 328bb43..61bb41d 100644 --- a/.agents/skills/augustus/references/mixed-architecture.md +++ b/.agents/skills/augustus/references/mixed-architecture.md @@ -307,10 +307,31 @@ ask Jev to own the standard of quality. Good checks: "does this diff introduce new mutable module-level state?", "API impact: none / additive / behavioral / breaking?" +[`coldteadotai/abide`](https://github.com/coldteadotai/abide) is the +fuller **productized** path of that contract (compile / calibrate / +tune / replay / audit; Claude Code / Codex / OpenCode). Soft +AGENTS.md / CLAUDE.md rules → one Score per rule on the **diff**, +never the conversation; hard, linter-checkable rules stay with the +linter — same layering family as jevgate's hard envelope (`mappings.md` +§18). Edit-phase vs turn-phase is an observation-window question +("added more than asked" has no answer after edit 1 of 12; +`question-design.md`). Product bands: ≥0.8 repair, 0.5–0.8 note, <0.5 +silence — **their** operating point, not a universal 0.8 (`mappings.md` +§2). Fail-open: no key / no network → the edit proceeds; hooks exit 0. +`.abide/rubric.json` quotes source lines; false positives are rule +rewrites, not a model swap. Replay of 93 sessions with an independent +reviewer: edit precision ~26%, turn ~73% (`notes.md` §47; +`validation.md`). Text/diff only — not multimodal. Complementary to +[`24601/rh-guard`](https://github.com/24601/rh-guard) (eval-integrity / +reward-hacking on tool use vs project soft rules on diffs): same +hook-host lessons, different judgment class; do not merge products. +Do not copy hooks. + Related placements: -- **AGENTS.md / project prefs as criteria** — jev-pref; pi-warden rule - breaks 6→0 on 150 paired runs (`agent-self-assessment.md`). +- **AGENTS.md / project prefs as criteria** — jev-pref states the + contract; Abide productizes it; pi-warden rule breaks 6→0 on 150 + paired runs (`agent-self-assessment.md`; `notes.md` §47). - **Confidence gates + shadow mode** — `AntonioCoppe/jev-harness` (48.9s Claude CLI vs 1.3s Jev on a 24-row filter). Log would-do until evals pass. Selective abstention (`mappings.md` §2): low confidence is @@ -353,7 +374,7 @@ decision-design card. Do not clone APIs from READMEs. |---|---|---|---| | Hold-before-publish moderation | Hazard Nouls + harm Score | Block/review/pass policy | Near Here / firehose family | | Tool / engine / skill select | Choice + fits-Noul | Dispatch, auth, reject-all | skillranker, LlamaIndex selectors, Toolrouter | -| Preference lint | Per-rule Noul/Choice on a diff | Rule text, outcome map | jev-pref, JevLint | +| Preference lint | Per-rule Score/Noul on a diff | Rule text, linter for hard rules, bands + fail-open | jev-pref (contract), Abide (productized), JevLint | | Context / log prune | Per-line or per-block relevance | Always-keep set, recall keys | jevprune, winnow | | Exact hunk staging | Per-hunk include/exclude/mixed | `git diff`, atomic apply | git-jev-stage | | Semantic `WHERE` | Noul/`jev_prob` over a row | SQL, indexes, LIMIT | jevql (CLI; DB sees ordinary SQL); sqlite-jev (in-engine extension) | diff --git a/.agents/skills/augustus/references/question-design.md b/.agents/skills/augustus/references/question-design.md index 84703c3..4983ffb 100644 --- a/.agents/skills/augustus/references/question-design.md +++ b/.agents/skills/augustus/references/question-design.md @@ -67,6 +67,7 @@ request, and treat a stale pin as a prior, never a setting. | Answer follows state text | Content steers the model | Tighten criteria; adversarial tests; confidence-gate the action | | Rewording trades one error for another | One question, several properties | Split into atomic questions | | Synonymous wording swings p / the act | Stimulus includes question text; no invariance promised | Paraphrase-pair eval; abstain or raise t; rewrite (`mappings.md` §17) | +| Question has no answer yet (edit 1 of 12) | Observation window is wrong: a turn-level property asked at edit time | Name when the evidence exists. Edit-phase vs turn-phase is a question-design cut, not a hook detail ([Abide](https://github.com/coldteadotai/abide): "added more than asked" is a turn rule). `notes.md` §47 | | Each answer right, decision wrong | Policy wrong | Change weights/thresholds in code, leave questions alone | ## Revision discipline diff --git a/.agents/skills/augustus/references/toolbox-mapping.md b/.agents/skills/augustus/references/toolbox-mapping.md index 1b989fb..f5ab5c0 100644 --- a/.agents/skills/augustus/references/toolbox-mapping.md +++ b/.agents/skills/augustus/references/toolbox-mapping.md @@ -73,7 +73,7 @@ component; keep the rest of the method in code. | Discrete math: width vs depth | Fan out in width (parallel ≈ free), pay depth linearly; two-stage only when next options depend on an earlier answer | **Empirical recipe** (fan-out: 12.2× cheaper, 10× faster) | | Psychology: Kahneman | System 2 generates/proposes (LLM), System 1 discriminates (Jev); never the reverse. S1 keeps control; optional S2 is one-use advice and does not fly | **Empirical recipe** (mcts-agent role split; 2026-09-18 mixed-architecture discourse; jev-reflex-autonomy-lab as a control-loop shape, `notes.md` §46) | | IR / cascades | Cheap relevance / irrelevance before an expensive ranker or generator | **Empirical recipe** (RAG cookbook; jevprune; git-jev-stage; LlamaIndex Jev rerank) | -| Spec / lint | Project-defined semantic rules as predicates over a diff | **Empirical recipe** (jev-pref contract; JevLint file-level Noul; pi-warden; snifftest unsure-band) | +| Spec / lint | Project-defined semantic rules as predicates over a diff; linter owns hard rules | **Empirical recipe** (jev-pref contract; Abide productized path — replay 93 sessions, edit precision ~26% / turn ~73% before tune, `notes.md` §47; JevLint file-level Noul; pi-warden; snifftest unsure-band) | | Formal methods / DST / safety | Judgment triages counterexamples, failing seeds, and named-rule conformance; proof/MC/DST stay with their tools. Alloy finder ≠ Apalache BMC ≠ Quint run. DST trio: Antithesis hypervisor / Resonate HQ Lean+oracle+SDK (durable async) / PufferLib env+seed. Noul is a sensor, not a discharged PO. Semi-formal diagrams are vocabularies, not enforcers | **Hypothesis as product**, **Contract** as ownership (matching `mappings.md` §8 and `methods-catalog.md`; worked shape pi-warden — `formal-methods.md`, `formal-semi-formal.md`) | | Decision analysis: VOI | Gather as an enumerated act; pay iff expected decision-loss drop > cost | **Hypothesis** as calculator (`mappings.md` §6; `mental-models.md`) | | Signal detection | Noul as evidence variable; criterion from costs and base rate; ROC/PR on your labels | **Hypothesis** for non-SWE plots (`mappings.md` §7) | diff --git a/.agents/skills/augustus/references/validation.md b/.agents/skills/augustus/references/validation.md index d757476..6b4a8a6 100644 --- a/.agents/skills/augustus/references/validation.md +++ b/.agents/skills/augustus/references/validation.md @@ -316,9 +316,18 @@ Rules: | End-to-end product / agent loop | Harbor taskset | behavioral assertions, cost/perf bounds | | LM-program knobs only | DSPy/Ax (narrow) | never primary System One calibration score | | Reward-hack / eval gaming | [rh-guard](https://github.com/24601/rh-guard) | structural deny + System One sidecar | - -rh-guard is a reward-hack hook, a different surface from jevgate. One -row is enough. ECE above is wanted, not a Nimble result. +| Project soft-rule lint | [Abide](https://github.com/coldteadotai/abide) | Score per rule on the diff; bands; fail-open; replay + independent review | + +rh-guard is a reward-hack hook, a different surface from jevgate and +from Abide (eval-integrity vs allowlist-remainder vs project soft +rules). One row each. ECE above is wanted, not a Nimble result. +Abide replay (author-reported, not re-run; `notes.md` §47): 93 +sessions, 1,256 edits / 147 turns; independent-reviewer precision +**edit ~26% / turn ~73%** before calibrate/tune. Harbor-adjacent +measurement (frozen transcripts, phase split, independent +confirmation), not a Harbor taskset and not a jevals substitute. +Turn-phase soft rules held up better; false positives mostly fixable +in the rubric. Text/diff only. ### Bake-off mandate diff --git a/CHANGELOG.md b/CHANGELOG.md index f5d3087..3b6be1a 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -141,6 +141,19 @@ Pegged against [`typesafe-ai/skills` v0.5.7](https://github.com/typesafe-ai/skil wrong; gating cannot catch it (`tenbin` owns the lint skill). S1 reflex keeps control (jev-reflex-autonomy-lab). MED: jev-decision-layer, jev-e2e, jevpandas. No wrapper. +- Abide (`coldteadotai/abide`, `research/notes.md` §47): productized + Jev preference lint for Claude Code / Codex / OpenCode. Soft + AGENTS.md / CLAUDE.md rules → one Score per rule on the diff (never + the conversation); hard rules stay with the linter (same layering + family as jevgate). Edit- vs turn-phase observation window; banded + confidence (≥0.8 repair / 0.5–0.8 note / <0.5 silence — their + operating point) + fail-open hooks; rubric.json quotes source + lines; calibrate/tune fix false positives in the question. Replay + of 93 sessions (1,256 edits / 147 turns) with independent review: + edit precision ~26%, turn ~73% (author-reported, before tune). + Fuller productized path of the jev-pref contract. Complementary to + rh-guard (eval-integrity vs project soft rules). Text/diff only — + not multimodal. No hook how-to. - Effect-oriented loops (`notes.md` §28, `mappings.md` §19): Ward's ZIO client keeps Jev as the outer Choice and the handler as the effect. Not Effect.ts. GLiNER author: GLiNER2 "like jev" is GLiGuard diff --git a/README.md b/README.md index 57b6aaa..245671b 100644 --- a/README.md +++ b/README.md @@ -51,7 +51,7 @@ never launder a Noul as a proof. exact-text keep/drop, env triage, moderation/ranking, skill routing - `.agents/skills/augustus/references/faq.md` — "just classification", stack replacement, Jev vs open head vs encoder vs LoRA vs constrained AR vs kev vs blackwood, - wait-for-Archer, missing-other confident-wrong, GLiNER vs GLiClass vs CLIP, LLM-as-judge, in-engine vs CLI store, + wait-for-Archer, missing-other confident-wrong, soft project rules vs linter (Abide), GLiNER vs GLiClass vs CLIP, LLM-as-judge, in-engine vs CLI store, hard envelope (bitrate / planner), not-another-how-to - `.agents/skills/augustus/references/mappings.md` — classical-method mappings with boundaries, counterexamples, acceptance tests (including @@ -59,7 +59,8 @@ never launder a Noul as a proof. - `.agents/skills/augustus/references/validation.md` — design gate, eval recipes, Jev-for-skills (routing, self-monitoring, testing, modularity, frontmatter), and Eval & hill-climb (jevals hygiene + Harbor taskset; - open-jev-laya-bench as ECE/NLL/Brier bake-off exemplar) + open-jev-laya-bench as ECE/NLL/Brier bake-off exemplar; Abide replay as + Harbor-adjacent soft-rule measurement) - `.agents/skills/augustus/references/boundary-audit.md` — existing-system insertion: fit test, opportunity map, smallest boundary, red flags - `.agents/skills/augustus/scripts/evaluate_decisions.py` — offline evaluator diff --git a/docs/ecosystem.md b/docs/ecosystem.md index 12be887..8bc71dd 100644 --- a/docs/ecosystem.md +++ b/docs/ecosystem.md @@ -69,6 +69,7 @@ Per-keystroke launchers (104ms median, sequence-tagged staleness), firehose mode ### Structural prove ∩ remainder - **thevibeworks/jevgate** — Proven / Refused / Unknown; cannot block; 0/59 unsafe unasked held-out. Allowlist **proves** read-only verbs; Jev judges only unlisted. Allowlist ∩ System One. +- **coldteadotai/abide** — same family, different remainder: linter proves lintable rules; Jev Scores residual soft AGENTS.md rules; fail-open, banded. `notes.md` §47. - **suraj-phanindra/wellposed** — lint the Jev request before it comes back confidently wrong. Missing `other` → confidence 1.00 on a wrong Choice; gating cannot catch it. `tenbin` owns the lint skill. `notes.md` §46. - **misbahsy/doc-router** — page OCR router: 155→87 billed, 1.74× $ on 19 docs / 155 pages. Same sandwich. @@ -98,6 +99,7 @@ Per-keystroke launchers (104ms median, sequence-tagged staleness), firehose mode - **ax-llm/ax** — native `typesafe` provider; **typesafeainate/dspy-typesafeify** — DSPy decorator PoC. - **riff (scale-venture-partners)** — hybrid static+semantic linter with ruff-style JEV codes and per-finding calibrated p; 14 calls ≈ $0.0004. - **huntedman/JevLint** — file-level convention Nouls (magic-strings, descriptive-names); write→check→fix; no line-level/auto-fix. Sibling of jev-pref. Independent. Pointer: `notes.md` §26. +- **coldteadotai/abide** — productized Jev preference lint (Claude Code / Codex / OpenCode). Soft instruction-file rules as one Score per rule on the diff; linter owns hard rules; bands + fail-open; compile/calibrate/tune/replay. Replay 93 sessions: edit precision ~26% / turn ~73% (independent review, before tune). Fuller path of jev-pref; complementary to rh-guard. Text/diff only. `notes.md` §47. - **super-jev / probably / jev-search** above — see their cards in `archive/findings.md`. ### Mixed architecture (2026-09-18T14 discourse + topic:jev) @@ -106,7 +108,7 @@ Default placement, not a new product class: Jev judges, an LLM writes, code owns control. Movers that sharpened the card: `git-jev-stage` (exact hunk Choice), `jevprune` (per-line relevance with an always-keep set), `llama-index-jev` (rerank fails open / select fails closed), `jev-pref` -(AGENTS.md as criteria), `lizard-agent` (no LLM when nothing needs writing), +(AGENTS.md as criteria; Abide is the productized path, `notes.md` §47), `lizard-agent` (no LLM when nothing needs writing), `jevql` (judgment as SQL `WHERE`; CLI so Postgres never sees `jev()`), `sqlite-jev` (in-engine SQLite extension; same hole), OpenSmoke (Jev over every step, LLM only on flags). Neighbor skills `tenbin` and `decision-first` are *not* Augustus diff --git a/research/archive/findings.md b/research/archive/findings.md index 7b0e3d8..268ff1b 100644 --- a/research/archive/findings.md +++ b/research/archive/findings.md @@ -741,6 +741,32 @@ Watch-only hole; (ax) bake-off substrate with ECE/NLL/Brier in the wild; (ay) decision-token LoRA is how you train constrained-AR, not a new species; (az) confidence gating cannot catch a forced Choice. +## Batch #30 (2026-09-18) — Abide productized preference lint + +Note: `research/notes.md` §47. Docs-only. Folded into PR #2. Not a +second PR. Archer still Watch. No invented metrics. Text/diff only. + +- **coldteadotai/abide (Empirical as README + dated replay; Hypothesis + on your AGENTS.md).** MIT, created 2026-09-18, TypeScript, npm + `@coldtea/abide`. Productized Jev hooks: one Score per soft project + rule on the diff, never the conversation. Linter owns hard rules + (jevgate-family sandwich, different remainder). Edit vs turn is an + observation window. Bands ≥0.8 / 0.5–0.8 / <0.5 are their operating + point, not a universal 0.8. Fail-open hooks. Rubric quotes source + lines; calibrate/tune rewrite dead rules. Replay 93 sessions, 1,256 + edits / 147 turns, $0.22: independent-reviewer precision edit 26% / + turn 73% before tune. Turn-phase soft rules held up better. No + turn-number drift. Replay does not measure in-session repair. +- **Siblings:** jev-pref (contract Abide productizes); rh-guard + (eval-integrity, not project soft rules); wellposed (request lint + upstream); jevgate (hard envelope); JevLint (file-level conventions). + Do not merge products. Do not copy hooks. + +Cross-repo addition: (ba) preference lint has a productized compile / +calibrate / tune / replay path; (bb) observation window (edit vs turn) +is question design; (bc) banded fail-open means soft judgment is never +the sole hard veto; (bd) false positives in the rubric, not the model. + diff --git a/research/notes.md b/research/notes.md index 352d77e..c28b8b8 100644 --- a/research/notes.md +++ b/research/notes.md @@ -2618,3 +2618,128 @@ Cards: `judgment-class.md`; `validation.md`; `question-design.md`; `agent-self-assessment.md`; `mental-models.md`; `formal-methods.md`; `applied-mappings.md` §2; `methods-catalog.md`; `toolbox-mapping.md`. No wrapper. + +## 47. Abide — productized soft-rule preference lint (2026-09-18) + +America/Boise ~14:34 = 20:34 UTC. Docs-only fold into open PR #2 +(`cursor/augustus-store-envelope-00b4`). Not a competing PR. Archer +27B drop still **WATCH**. Identity lock vs `typesafe-ai` / `tenbin` / +`decision-first` holds. No wrapper, no hook how-to, no copied `npx` / +init / ports. Text/diff only — **not multimodal**. Do not invent +metrics. + +HTTP 200 this pass: GitHub README for +[`coldteadotai/abide`](https://github.com/coldteadotai/abide) +(MIT, TypeScript, created 2026-09-18, npm `@coldtea/abide`, 4★ this +pass); replay method in-repo at `benchmarks/replay/README.md`. +Attached SIGNAL + replay README + repo.json as primary placement +guidance. + +### HIGH + +1. **[`coldteadotai/abide`](https://github.com/coldteadotai/abide)** + — productized Jev hooks for Claude Code / Codex / OpenCode that + enforce **soft project rules** (AGENTS.md / CLAUDE.md / instruction + files) no linter can check. On every edit (or turn) it asks Jev + **one Score / probability per rule against the diff only**, never + the conversation. A break names the rule and asks the agent to + repair in-session. jev-pref already stated the contract (YOU define + the rule / Jev classifies / code maps outcome). Abide is the + fuller productized path: compile / calibrate / tune / replay / + audit + multi-host. Do not clone the CLI. + + **Architecture (load-bearing; not a hit list).** + + 1. **Soft rules → soft judgment; hard rules → linter.** Rules a + linter can check are handed to the linter. Same layering family + as jevgate's hard envelope (`mappings.md` §18): structure + **proves** what it can; typed judgment only on the residual. + jevgate's remainder is unlisted shell verbs; Abide's remainder + is project-instruction soft rules. Different holes; same + sandwich. Soft judgment is never the sole hard veto. + 2. **Edit-phase vs turn-phase is an observation-window question.** + Each compiled rule runs at `edit` (this hunk) or `turn` (the + whole turn diff). "Did this add more than was asked?" has no + answer after edit 1 of 12. Question design must name when the + evidence exists (`question-design.md`). + 3. **Banded confidence + fail-open.** Their product bands: ≥0.8 + repair in-session; 0.5–0.8 human note, agent silent; <0.5 + silence. That 0.8 is **their** operating point, not a universal + threshold (`mappings.md` §2 already forbids magic 0.8). Hooks + always exit 0, hard deadline; no key / no network → the edit + proceeds and the miss is logged. Soft judgment never sole hard + veto. + 4. **Rubric as editable artifact.** `.abide/rubric.json` quotes the + source instruction line. A wrong verdict is a rule rewrite. + `calibrate` scores rules against recent git history; `tune` + rewrites dead rules. False positives live in the question + (scope, criteria), not in the model. + 5. **Eval honesty / Harbor-adjacent.** Replay of 93 real Claude + Code sessions against each repo's own AGENTS.md (two private + repos + public `pr-lens`; hunks unpublished). Nothing is + re-run: diffs come from transcripts. Independent reviewer + (Claude, reading each rule's own text strictly; owner + spot-checked four first-pass comment flags and agreed). Replay + does **not** measure whether the agent repairs when told. + Harbor/jevals practice in the wild: frozen transcripts, phase + split, independent confirmation — not a Harbor taskset and not + a second jevals (`validation.md`). + + **Evidence (README + `benchmarks/replay/README.md`; not re-run; + Empirical as that named receipt).** 93 sessions with edits; 1,256 + edits judged; 147 turns judged; Jev cost **$0.22**; wall **~2 + min**. Edits flagged at 0.8+: **39 (3.1%)**. Turns flagged at + 0.8+: **15 (10%)**. After independent review: edits **10/39 = + 26%** precision; turns **11/15 = 73%**; all **21/54 = 39%**. + Author's gloss: 8 confirmed violations per 1,000 edits; 1 turn in + 13 ends with a confirmed violation of a rule no linter could + express; 11 of 93 sessions contained at least one. Recall probe + from 20 cleared hunks closest to the line (0.31–0.48): one real + miss (props-ordering). Most false positives from two edit-phase + rules (`plain-error-for-expected-failure`, `comment-volume`) — + fixable with `scope` and criteria, which is what calibrate / tune + exist for; the table is **before** either ran on the tightened + rubric. No turn-number drift: pr-lens-app flat ~2.5% through early + / mid / late turns; coldtea falls 5.8% → 1.3% because big + new-file writes happen early. Agents break these rules from the + first edit at a steady rate. Economics (README, author-measured + 2026-09-18, direct to TypeSafe): a check was ~2,500 tokens with an + ordinary LLM (cent+, seconds, prose to parse); Jev ~300 ms, + **$0.00004–0.00007** per check on this repo's 13 rules (1,000–1,600 + input tokens). Do not promote those $ / ms figures as class + constants. + + **Siblings — complementary, do not merge.** + + - **`doeixd/jev-pref`:** earlier watch; the contract Abide + productizes. Keep the contract quote; point here for compile / + calibrate / tune / replay / multi-host. + - **`24601/rh-guard`:** reward-hacking / eval-integrity on agent + tool use. Abide: project-instruction soft rules on diffs. Same + hook-host surface, different judgment class. Shared fail-open / + host-adapter lessons; no code dependency. + - **`suraj-phanindra/wellposed`:** request-shape lint still + upstream of any Score call (`tenbin` owns the skill). + - **`thevibeworks/jevgate`:** hard envelope owns safety; Abide is + soft residual judgment on residual soft rules. + - **`huntedman/JevLint`:** file-level convention Nouls; sibling, + not a substitute. + + **Placement.** Verifier over rules the project already wrote + (`mixed-architecture.md` preference lint; composition-algebra #9). + Pillar: selective classification / abstention + structural-prove ∩ + remainder. Hole: gate. Family: closed decision API (typed Score + per rule). Fail-open, banded. Eval path: replay + independent + review (named receipt above); not a substitute for jevals/Harbor + on *your* rubric. **Empirical** as README + dated replay. **Hypothesis** + that the same bands / precision transfer to *your* AGENTS.md. + Cards: `mixed-architecture.md`; `question-design.md`; + `validation.md`; `mappings.md` §2 / §18; `faq.md`; + `agent-self-assessment.md`; `toolbox-mapping.md`; + `methods-catalog.md`; `mental-models.md`. No wrapper. + +### Omni / Jev-omni + +Text/diff only today. Usage + measurement exemplar (replay harness, +precision by phase). Not a multimodal substrate and not a reason to +wait on Archer. diff --git a/research/refresh-log.md b/research/refresh-log.md index d728ad4..a29d6e7 100644 --- a/research/refresh-log.md +++ b/research/refresh-log.md @@ -410,6 +410,25 @@ formal-methods, applied-mappings, methods-catalog, toolbox. - notes.md §46; sources.json; findings.md batch #29. No wrapper. +## 2026-09-18 20:34 UTC — Abide soft-rule preference lint + +- Folded into open PR #2 (`cursor/augustus-store-envelope-00b4`). + Docs-only. Not a competing PR. Archer 27B drop still **WATCH**. +- HIGH: [`coldteadotai/abide`](https://github.com/coldteadotai/abide) + (MIT, created 2026-09-18). Productized Jev preference lint for + Claude Code / Codex / OpenCode. Soft AGENTS.md / CLAUDE.md rules → + one Score per rule on the diff (never the conversation); hard rules + stay with the linter (jevgate-family layering). Edit- vs turn-phase + observation window; banded confidence + fail-open; rubric as + artifact (calibrate/tune). Replay 93 sessions / 1,256 edits / 147 + turns; independent-reviewer precision edit ~26% / turn ~73% + (author-reported, before tune). Fuller path of jev-pref; + complementary to rh-guard. Text/diff only — not multimodal. +- Cards: mixed-architecture, question-design, validation, mappings + §2 / §18, faq, agent-self-assessment, toolbox, methods-catalog, + mental-models, formal-methods, composition-algebra, SKILL.md. +- notes.md §47; sources.json; findings.md batch #30. No wrapper. + diff --git a/research/sources.json b/research/sources.json index 9729a0d..c01aac2 100644 --- a/research/sources.json +++ b/research/sources.json @@ -1,6 +1,6 @@ { "refresh_cadence": "hourly", - "retrieved": "2026-09-18T20:03Z", + "retrieved": "2026-09-18T20:34Z", "sources": [ { "kind": "docs", @@ -318,7 +318,7 @@ "kind": "repo", "title": "doeixd/jev-pref", "url": "https://github.com/doeixd/jev-pref", - "note": "AGENTS.md \u2192 Jev linter. Contract: YOU define the rule; JEV classifies evidence; CODE maps outcome." + "note": "AGENTS.md \u2192 Jev linter. Contract: YOU define the rule; JEV classifies evidence; CODE maps outcome. Abide is the fuller productized path (notes.md \u00a747)." }, { "kind": "repo", @@ -798,7 +798,7 @@ "kind": "repo", "title": "24601/rh-guard", "url": "https://github.com/24601/rh-guard", - "note": "README HTTP 200. Coding-agent reward-hack / eval-integrity hooks. Different surface from GLiGuard LLM I/O safety and from jevgate. Do not copy install into Augustus." + "note": "README HTTP 200. Coding-agent reward-hack / eval-integrity hooks. Different surface from GLiGuard LLM I/O safety, jevgate, and Abide (project soft rules). Complementary fail-open / hook-host lessons; do not merge. Do not copy install into Augustus. notes.md \u00a747." }, { "kind": "social", @@ -1303,6 +1303,18 @@ "title": "yalindogusahin/jevpandas", "url": "https://github.com/yalindogusahin/jevpandas", "note": "pandas semantic index; noul/choice/score; LICENSE 404 this pass. notes.md \u00a746." + }, + { + "kind": "github", + "title": "coldteadotai/abide", + "url": "https://github.com/coldteadotai/abide", + "note": "MIT, created 2026-09-18, TypeScript, npm @coldtea/abide. Productized Jev preference lint (Claude Code / Codex / OpenCode). Soft AGENTS.md rules as one Score per rule on the diff; linter owns hard rules; bands + fail-open; compile/calibrate/tune/replay. Text/diff only. notes.md \u00a747." + }, + { + "kind": "github", + "title": "coldteadotai/abide benchmarks/replay", + "url": "https://github.com/coldteadotai/abide/blob/HEAD/benchmarks/replay/README.md", + "note": "Replay 93 Claude Code sessions, 1256 edits / 147 turns. Independent reviewer precision edit 26% / turn 73% before calibrate/tune. Author-reported; hunks unpublished. Harbor-adjacent, not a Harbor taskset. notes.md \u00a747." } ] } From d19541f6c54a1aff6c9ee60486462decbed13bbf Mon Sep 17 00:00:00 2001 From: Cursor Agent Date: Fri, 18 Sep 2026 20:46:50 +0000 Subject: [PATCH 05/43] Fold kev delta: Hub weights and NOTA training confrontation MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Do not rewrite §45. Hub jaredpalmer/kev-0.5b; PEFT FEATURE_EXTRACTION. Choice other must appear as a wrong alternative too — wellposed request lint is not enough. No invented NOTA rates. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com> --- .agents/skills/augustus/SKILL.md | 4 +-- .agents/skills/augustus/references/faq.md | 10 ++++-- .../augustus/references/judgment-class.md | 10 ++++-- .../augustus/references/mixed-architecture.md | 4 ++- .../augustus/references/question-design.md | 3 +- .../skills/augustus/references/validation.md | 6 +++- CHANGELOG.md | 8 +++++ docs/ecosystem.md | 2 +- research/archive/findings.md | 21 ++++++++++++ research/notes.md | 34 +++++++++++++++++++ research/refresh-log.md | 14 ++++++++ research/sources.json | 10 ++++-- 12 files changed, 114 insertions(+), 12 deletions(-) diff --git a/.agents/skills/augustus/SKILL.md b/.agents/skills/augustus/SKILL.md index b9dd905..b25df0c 100644 --- a/.agents/skills/augustus/SKILL.md +++ b/.agents/skills/augustus/SKILL.md @@ -1,6 +1,6 @@ --- name: augustus -description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill, open multimodal RLCD / blackwood), announced open decision-model (Watch — still not landed), constrained-AR (TypeAR, pcdServer, decision-token LoRA), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"wait for Archer vs open multimodal RLCD\", \"screenshot/DOM candidates → typed Choice\", \"eval path\", \"jevals\", \"Harbor taskset\", \"shared bake-off ECE/NLL/Brier\", \"LLM-as-judge is not the System One score\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge (allowlist *proves*; fail-open cannot block), \"missing other → confident wrong Choice\", \"lint the Jev request\", \"S1 reflex keeps control / optional S2 one-use advice\", \"soft AGENTS.md rules vs linter\" (Abide / jev-pref), \"edit-phase vs turn-phase observation window\", banded confidence fail-open preference lint, TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, or \"formally verify with Jev\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." +description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill, open multimodal RLCD / blackwood), announced open decision-model (Watch — still not landed), constrained-AR (TypeAR, pcdServer, decision-token LoRA), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"wait for Archer vs open multimodal RLCD\", \"screenshot/DOM candidates → typed Choice\", \"eval path\", \"jevals\", \"Harbor taskset\", \"shared bake-off ECE/NLL/Brier\", \"LLM-as-judge is not the System One score\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge (allowlist *proves*; fail-open cannot block), \"missing other → confident wrong Choice\", \"lint the Jev request\", \"training confronts Choice other / none-of-the-above\", \"S1 reflex keeps control / optional S2 one-use advice\", \"soft AGENTS.md rules vs linter\" (Abide / jev-pref), \"edit-phase vs turn-phase observation window\", banded confidence fail-open preference lint, TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, or \"formally verify with Jev\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." license: MIT metadata: version: 0.3.0 @@ -54,7 +54,7 @@ classical method you already trust, substitute it, classify the win "paraphrase brittleness", "allowlist then judge", "TOCTOU-of-Noul", "Jev inside the database / sqlite-jev", "Jev picks bitrate / join order / the model", "wait for Archer", "lint the request / missing - other", "soft AGENTS.md rules vs the linter", or "screenshot Choice / omni System One": + other", "training confronts Choice other / none-of-the-above", "soft AGENTS.md rules vs the linter", or "screenshot Choice / omni System One": read `references/faq.md`, then `references/mental-models.md`, then `references/mixed-architecture.md`, then diff --git a/.agents/skills/augustus/references/faq.md b/.agents/skills/augustus/references/faq.md index 5c6bca7..423dafe 100644 --- a/.agents/skills/augustus/references/faq.md +++ b/.agents/skills/augustus/references/faq.md @@ -89,7 +89,8 @@ on your own independent labels before you treat it as a decision API [jaredpalmer/kev](https://github.com/jaredpalmer/kev) is the laptop-local System One **API drop-in** on that same open path: Qwen2.5-0.5B LoRA + pointer, public gold not a Jev teacher, official SDK with a `base_url` -change. Use it for development and eval. Do not use 0.5B ID ECE as a +change. Hub weights: [`jaredpalmer/kev-0.5b`](https://huggingface.co/jaredpalmer/kev-0.5b) +(`notes.md` §45 delta). Use it for development and eval. Do not use 0.5B ID ECE as a knowledge or frontier substitute (`notes.md` §45). ## Open weights vs Jev vs constrained decoding vs encoder vs LoRA? @@ -325,7 +326,12 @@ escape hatch, broken state paths) before you trust the number. Recipe: email → `"support issue"` at 1.00 without `other`; overlapping options collapse to 0.19 — that failure is loud). `tenbin` still owns the design-time lint *skill*; Augustus owns the placement. -`question-design.md`; `notes.md` §46. +`question-design.md`; `notes.md` §46. Putting `"other"` on the request +is necessary and not sufficient for an open head you train: the +residual option must also appear as a **wrong** alternative, with +varied wording, or the hatch becomes a shortcut +([kev](https://github.com/jaredpalmer/kev) first-run lesson; +`none_of_the_above` eval; `notes.md` §45 delta). ## Should Jev live inside the database? diff --git a/.agents/skills/augustus/references/judgment-class.md b/.agents/skills/augustus/references/judgment-class.md index 66938d0..aafb049 100644 --- a/.agents/skills/augustus/references/judgment-class.md +++ b/.agents/skills/augustus/references/judgment-class.md @@ -558,8 +558,14 @@ README, MODEL_CARD, LICENSE, and release questions under a block-causal mask, one prefill, no decode. Architecture follows [Archer Hume's reconstruction](https://archerhume.com/posts/jevs-architecture-unmasked). Speaks TypeSafe `POST /v1/systemone`; official `typesafe-sdk` works -with a `base_url` change. Weights `kev-0.5b` (38 MB) on that release. -Not a how-to: do not copy serve flags, ports, or train commands. +with a `base_url` change. Weights `kev-0.5b` (38 MB) on that release +**and** on the Hub as [`jaredpalmer/kev-0.5b`](https://huggingface.co/jaredpalmer/kev-0.5b) +(`kev.publish`; `--run` accepts Hub ids; base still downloads on first +load). PEFT `task_type=FEATURE_EXTRACTION`; publish patches legacy +adapters. Not a how-to: do not copy serve flags, ports, or train +commands. **NOTA:** training must confront Choice `"other"` as a wrong +alternative too, with varied wording (`notes.md` §45 delta; +`question-design.md`). **Place it on the trained decision-only open path** next to Laya / Nimble / Archer Watch. It is the cleanest *runnable* productization of diff --git a/.agents/skills/augustus/references/mixed-architecture.md b/.agents/skills/augustus/references/mixed-architecture.md index 61bb41d..9763b5d 100644 --- a/.agents/skills/augustus/references/mixed-architecture.md +++ b/.agents/skills/augustus/references/mixed-architecture.md @@ -420,7 +420,9 @@ backbone or a second API skill. Laya: self-hostable, text-only, 512 tokens/question; vendor benches vs Jev are **claims**. Encoder open-jev: public gold, OOD drop. LoRA student: teacher-copy. **kev**: public gold, pointer readout, System One API drop-in; ID ECE only; not a teacher-copy -(`notes.md` §45). **blackwood-rlcd**: open multimodal RLCD, Jev-compatible +(`notes.md` §45). Hub fetch: +[`jaredpalmer/kev-0.5b`](https://huggingface.co/jaredpalmer/kev-0.5b). +**blackwood-rlcd**: open multimodal RLCD, Jev-compatible shim, CC BY-NC; Jev still leads general text; not Archer Watch (`notes.md` §46). Hume's 27B decision-model drop is **Watch**. Closed calibrated API vs open weights diff --git a/.agents/skills/augustus/references/question-design.md b/.agents/skills/augustus/references/question-design.md index 4983ffb..43a21c3 100644 --- a/.agents/skills/augustus/references/question-design.md +++ b/.agents/skills/augustus/references/question-design.md @@ -44,7 +44,7 @@ request, and treat a stale pin as a prior, never a setting. ## Criteria shape - Criteria are an extension of the instruction and must ask the same thing, in the same direction (a Noul whose `true` side describes "no" performs worse). -- Choice options: contrastive `what` / `not_for` / short concrete `examples` (instances, not descriptions of instances). Add an `other` / `none-of-the-above` when the list may not cover inputs. Skipping that hatch is not a style nit: the model will pick a listed option at confidence 1.00, and no downstream gate will see a problem (`notes.md` §46). +- Choice options: contrastive `what` / `not_for` / short concrete `examples` (instances, not descriptions of instances). Add an `other` / `none-of-the-above` when the list may not cover inputs. Skipping that hatch is not a style nit: the model will pick a listed option at confidence 1.00, and no downstream gate will see a problem (`notes.md` §46). Request-shape lint (wellposed / `tenbin`) puts the hatch on the offered set; **training must confront it as a wrong alternative too**, with varied wording, or the model learns "this wording ⇒ pick it" ([kev](https://github.com/jaredpalmer/kev) first-run shortcut; dedicated `none_of_the_above` eval; `notes.md` §45 delta). - Score levels (2–10): describe **situations**, one dimension each, each standing alone (Jev sees neither the level's number nor its neighbors — "worse than previous" means nothing). No numerals. Levels may be objects `{"summary", "signals"}`. Give a rare extreme its own level when code treats it differently. - Composite scoring: one Score per dimension, normalize by `len(criteria)-1`, weight and combine in code. Change policy by changing weights — never by rewriting questions. - Taxonomy walk: one Choice per tree level, walk in code; each option's value is its subtree (direct children + sample leaves); follow several branches when probabilities are close. @@ -55,6 +55,7 @@ request, and treat a stale pin as a prior, never a setting. | --- | --- | --- | | Wrong answers, high confidence | Instruction read literally | State exact condition; put boundary cases in criteria | | Wrong answers, **confidence ~1.00**, no `other` | Forced pick: the offered set does not cover the input; the model *must* choose | Add `other` / none-of-the-above. **Confidence gating cannot catch this** ([wellposed](https://github.com/suraj-phanindra/wellposed) live probe: unsubscribe email → `"support issue"` at 1.00 without `other`, `"other"` at 0.93 with it). Overlapping options collapse confidence (loud). `notes.md` §46. `tenbin` owns the lint skill | +| Residual `"other"` always picked (or never) | Training saw none-of-the-above only as the true label — a wording shortcut | Confront the hatch as a *wrong* alternative too; vary wording; eval present-vs-removed ([kev](https://github.com/jaredpalmer/kev) `none_of_the_above`; `notes.md` §45 delta). wellposed still owns request-shape lint | | Question names a state path that does not exist | Dead reference; the API still answers | Lint the request (walk JSON). Structural, not semantic. wellposed recipe; do not copy the CLI | | Low-confidence Choice | Options overlap / none fits | `what`/`not_for`/`examples`; add `other` | | Low-confidence Score | Overlapping levels, two dimensions, thin state | Distinct-situation levels; split question; add state field | diff --git a/.agents/skills/augustus/references/validation.md b/.agents/skills/augustus/references/validation.md index 6b4a8a6..b5319a0 100644 --- a/.agents/skills/augustus/references/validation.md +++ b/.agents/skills/augustus/references/validation.md @@ -44,7 +44,11 @@ Open reconstruction cousin: [`jaredpalmer/kev`](https://github.com/jaredpalmer/k — isolation packed vs separate max Δ 3.7e-6; secret-in-sibling p=0.03 vs in-state 0.99; permute argmax flips 7.4%; IIA log-odds shift mean 0.13; boundary forgery held. Those tests mirror Archer probes; they do -not prove kev = Jev (`notes.md` §45). +not prove kev = Jev (`notes.md` §45). Hub fetch path this pass: +[`jaredpalmer/kev-0.5b`](https://huggingface.co/jaredpalmer/kev-0.5b) +(`--run` accepts Hub ids). Dedicated `none_of_the_above` eval (true +option present vs removed) is the training-side cousin of wellposed's +request hatch; **no published rates this pass** (`notes.md` §45 delta). ## Offline eval: selective binary decisions diff --git a/CHANGELOG.md b/CHANGELOG.md index 3b6be1a..976a694 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -154,6 +154,14 @@ Pegged against [`typesafe-ai/skills` v0.5.7](https://github.com/typesafe-ai/skil Fuller productized path of the jev-pref contract. Complementary to rh-guard (eval-integrity vs project soft rules). Text/diff only — not multimodal. No hook how-to. +- kev delta (`research/notes.md` §45): Hub weights + [`jaredpalmer/kev-0.5b`](https://huggingface.co/jaredpalmer/kev-0.5b); + `--run` accepts Hub ids; PEFT `task_type=FEATURE_EXTRACTION` (publish + patches legacy adapters). HIGH question-design: confront Choice + `"other"` / none-of-the-above as a wrong alternative too, vary + wording, dedicated `none_of_the_above` eval (no published rates). + Cross-link wellposed request-shape lint. No species change. No + wrapper. - Effect-oriented loops (`notes.md` §28, `mappings.md` §19): Ward's ZIO client keeps Jev as the outer Choice and the handler as the effect. Not Effect.ts. GLiNER author: GLiNER2 "like jev" is GLiGuard diff --git a/docs/ecosystem.md b/docs/ecosystem.md index 8bc71dd..b6dd4eb 100644 --- a/docs/ecosystem.md +++ b/docs/ecosystem.md @@ -54,7 +54,7 @@ Per-keystroke launchers (104ms median, sequence-tagged staleness), firehose mode - **GLiGuard** (fastino-ai) — 0.3B GLiNER2 encoder, checkpoint `fastino/gliguard-LLMGuardrails-300M`. One bidirectional pass over a safety schema. Same interface shape as batched questions; different objective. Not a Jev weight clone. `judgment-class.md`; `notes.md` §30. - **DECRUX9812/openjev-lm** — Qwen2.5-0.5B+LoRA distilled from hosted Jev answers; 65/70 = 92.9% on 70 hand-labelled rows (one annotator, one domain, one seed) overnight on 6 vCPU, $0/call. Its 98.1% on fresh rows is teacher *agreement*, not gold. Receipts pattern: `notes.md` §25, §44. - **convaiinnovations/laya** — open Choice/Score/Noul head, text-only, 512 tok. Companion packaging this hour: [`laya-typed-decisions`](https://huggingface.co/convaiinnovations/laya-typed-decisions) (421.3M, acc 0.766 / Brier 0.066 unverified). Shared bake-off: [`pngwn/open-jev-laya-bench`](https://huggingface.co/datasets/pngwn/open-jev-laya-bench) (26+9 tasks, 11959 items; ECE/NLL/Brier; not TypeSafe Jev vs Laya). `notes.md` §18, §42, §46. -- **jaredpalmer/kev** — Qwen2.5-0.5B LoRA + pointer readout; Apache-2.0; `kev-0.5b` 38 MB on GitHub release v0.1.0. Runnable Archer reconstruction (`POST /v1/systemone`). Public gold, not a Jev teacher. Isolation exact; ID ECE 0.065 (0.031 after T); acc 0.799 / 1,350. Not a knowledge/frontier substitute. `notes.md` §45. +- **jaredpalmer/kev** — Qwen2.5-0.5B LoRA + pointer readout; Apache-2.0; Hub [`jaredpalmer/kev-0.5b`](https://huggingface.co/jaredpalmer/kev-0.5b) plus GitHub release tarball. Runnable Archer reconstruction (`POST /v1/systemone`). Public gold, not a Jev teacher. Isolation exact; ID ECE 0.065 (0.031 after T); acc 0.799 / 1,350. NOTA training must confront `"other"` as a wrong alternative (`notes.md` §45 delta). Not a knowledge/frontier substitute. - **BlackwoodAI/blackwood-rlcd** — open multimodal RLCD (image-text-to-text), Jev-compatible shim, CC BY-NC 4.0. Screenshot + marked candidates → Choice. Card: web acc 0.907 vs Jev 1.13 text-only 0.480; letter-shuffle 0.133 vs 0.587; ECE 0.037; ~200 ms H100. Jev still leads general text 0.850 vs 0.786. Not Archer Watch. `notes.md` §46. - **Foodoo1/Qwen3-14B-RLCD-Decision-LoRA** — decision-token QLoRA on Qwen3-14B under parallel constrained decoding. Held-out 200-case / 4-field: fraud_risk 64→95%, overall 85.2→98.8% at ~234 ms. Synthetic; not a financial product. `notes.md` §46. - **zmtomorrow/TypeAR** — constrained autoregressive decoding surface: typed fields on a pretrained open model, no retraining. Not a proper-scoring decision head. `research/notes.md` §32. diff --git a/research/archive/findings.md b/research/archive/findings.md index 268ff1b..021bac5 100644 --- a/research/archive/findings.md +++ b/research/archive/findings.md @@ -767,6 +767,27 @@ calibrate / tune / replay path; (bb) observation window (edit vs turn) is question design; (bc) banded fail-open means soft judgment is never the sole hard veto; (bd) false positives in the rubric, not the model. +## Batch #31 (2026-09-18) — kev delta (Hub + NOTA) + +Note: `research/notes.md` §45 delta. Docs-only. Folded into PR #2. +Not a rewrite of §45. No species change. No invented metrics. + +- **Hub weights.** [`jaredpalmer/kev-0.5b`](https://huggingface.co/jaredpalmer/kev-0.5b) + HTTP 200. `kev.publish`; `--run` accepts Hub ids; base still + downloads on first load. GitHub release tarball remains. 61★ this + pass (signal had 53★). +- **PEFT.** `task_type=FEATURE_EXTRACTION`; publish patches legacy + adapters with null task_type. Docs mention only kev-0.5b. +- **NOTA training (HIGH question-design).** First run learned "this + wording ⇒ pick it". Fix: add none-of-the-above as a wrong + alternative too; vary wording; dedicated `none_of_the_above` eval + (present vs removed). **No published rates.** wellposed still + owns request-shape lint; training must confront the residual + option. + +Cross-repo addition: (be) bake-off fetch path is a Hub id; (bf) +Choice `"other"` is a training confrontation, not only a request hatch. + diff --git a/research/notes.md b/research/notes.md index c28b8b8..d23f4ee 100644 --- a/research/notes.md +++ b/research/notes.md @@ -2398,6 +2398,40 @@ labels. Cards: `judgment-class.md`; FAQ; `validation.md`; (pointer-softmax is still a sensor); `optimizer-integration.md`. No wrapper. +### Delta (~14:35 Boise) — Hub weights, PEFT task_type, NOTA training + +Not a rewrite of §45. No architecture species change. HTTP 200 this +pass: Hub [`jaredpalmer/kev-0.5b`](https://huggingface.co/jaredpalmer/kev-0.5b) +(Apache-2.0 adapter; `peft` / `text-classification` card); GitHub +README now names that Hub id as the fetch path; release tarball +remains. GitHub this pass: **61★**, pushed ~20:29Z (signal had 53★). +Do not copy `--run` / publish flags into skill cards. + +1. **Hub weights.** `kev.publish` uploads the adapter; `--run` accepts + Hub ids (`jaredpalmer/kev-0.5b`); the Qwen2.5-0.5B base still + downloads on first load. Bake-off fetch path for jevals/Harbor + (`validation.md`). +2. **PEFT.** `LoraConfig(task_type="FEATURE_EXTRACTION")` (`kev/model.py`). + Publish patches legacy adapters that saved `task_type` null (Hub + warns; PEFT treats both the same on a bare backbone). Docs mention + only kev-0.5b. Pointer readout, not a new species. +3. **HIGH — `none_of_the_above` is a training question, not only a + request hatch.** First training run learned "this wording ⇒ pick + it" when NOTA appeared only as the correct answer (`kev/data.py` + comment). Fix: add NOTA as a **wrong alternative** too + (`p_none_distract`); vary wording (`NONE_OPTIONS`: "None of the + above" / "Something else" / "Not listed here" / …); dedicated + `test_none_of_the_above` (true option still present → little mass + on none; true option removed → pick none; a shortcut model picks + it in both cases). **No published rates this pass — do not + invent them.** Model card already augments with p=0.10 + true→`other: None of the above`; the *delta* is confronting the + hatch as a distractor as well. Cross-link wellposed / Choice + `"other"`: request-shape lint puts the residual option on the + offered set; **training must confront that option** or the hatch + becomes a wording shortcut. Cards: `question-design.md`; FAQ + forced Choice; `validation.md`. + ## 46. 14:03 Boise hourly — open multimodal RLCD, bake-off substrate, decision-token LoRA (2026-09-18) America/Boise 14:03 = 20:03 UTC. Docs-only fold into PR #2 diff --git a/research/refresh-log.md b/research/refresh-log.md index a29d6e7..9bd1615 100644 --- a/research/refresh-log.md +++ b/research/refresh-log.md @@ -429,6 +429,20 @@ mental-models, formal-methods, composition-algebra, SKILL.md. - notes.md §47; sources.json; findings.md batch #30. No wrapper. +## 2026-09-18 20:43 UTC — kev delta (Hub weights, NOTA training) + +- Folded into open PR #2 (`cursor/augustus-store-envelope-00b4`). + Docs-only. Not a rewrite of §45. No species change. +- Hub: [`jaredpalmer/kev-0.5b`](https://huggingface.co/jaredpalmer/kev-0.5b) + HTTP 200; `--run` accepts Hub ids; GitHub release tarball remains. + PEFT `task_type=FEATURE_EXTRACTION`; publish patches null task_type. + GitHub 61★ this pass. +- HIGH question-design: none-of-the-above must appear as a wrong + alternative too, varied wording; dedicated `none_of_the_above` eval + (no published rates). Cross-link wellposed / Choice `"other"`. +- Cards: question-design, faq, judgment-class, validation, ecosystem. +- notes.md §45 delta; sources.json; findings.md batch #31. No wrapper. + diff --git a/research/sources.json b/research/sources.json index c01aac2..93ad1db 100644 --- a/research/sources.json +++ b/research/sources.json @@ -1,6 +1,6 @@ { "refresh_cadence": "hourly", - "retrieved": "2026-09-18T20:34Z", + "retrieved": "2026-09-18T20:43Z", "sources": [ { "kind": "docs", @@ -1236,7 +1236,7 @@ "kind": "github", "title": "jaredpalmer/kev", "url": "https://github.com/jaredpalmer/kev", - "note": "Runnable Archer reconstruction: Qwen2.5-0.5B LoRA + pointer, POST /v1/systemone, Apache-2.0. Public gold, not a Jev teacher. notes.md \u00a745." + "note": "Runnable Archer reconstruction: Qwen2.5-0.5B LoRA + pointer, POST /v1/systemone, Apache-2.0. Public gold, not a Jev teacher. Hub weights jaredpalmer/kev-0.5b this pass. notes.md \u00a745 delta." }, { "kind": "github", @@ -1315,6 +1315,12 @@ "title": "coldteadotai/abide benchmarks/replay", "url": "https://github.com/coldteadotai/abide/blob/HEAD/benchmarks/replay/README.md", "note": "Replay 93 Claude Code sessions, 1256 edits / 147 turns. Independent reviewer precision edit 26% / turn 73% before calibrate/tune. Author-reported; hunks unpublished. Harbor-adjacent, not a Harbor taskset. notes.md \u00a747." + }, + { + "kind": "huggingface", + "title": "jaredpalmer/kev-0.5b", + "url": "https://huggingface.co/jaredpalmer/kev-0.5b", + "note": "Hub weights for kev. kev.publish; --run accepts Hub ids; base downloads on first load. PEFT FEATURE_EXTRACTION. Apache-2.0 adapter. notes.md \u00a745 delta." } ] } From b8afd59c193760e53944213bfdb20e452314957d Mon Sep 17 00:00:00 2001 From: Cursor Agent Date: Fri, 18 Sep 2026 21:04:59 +0000 Subject: [PATCH 06/43] Fold 14:52 Boise watch: extractive selection, local drop-in, observe-act MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Docs-only fold of novel categorization/scoring placements into PR #2: extractive quotes + offline redecide, pointer-not-generator, contract- compatible /v1/systemone (stub until hf), structured observe→decide→act, dataframe accessors (jevframe beside jevpandas), route≠memory, advisory sidecar, structure induction, collab-arm measurement, AST∩semantic lint. Archer remains Watch. No wrappers or copied how-tos. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com> --- .agents/skills/augustus/SKILL.md | 20 +- .../references/agent-self-assessment.md | 10 +- .../augustus/references/applied-mappings.md | 14 + .agents/skills/augustus/references/faq.md | 37 ++- .../augustus/references/judgment-class.md | 13 +- .../skills/augustus/references/mappings.md | 33 ++- .../augustus/references/mental-models.md | 4 + .../augustus/references/methods-catalog.md | 6 +- .../augustus/references/mixed-architecture.md | 23 +- .../augustus/references/question-design.md | 2 +- .../augustus/references/toolbox-mapping.md | 6 +- .../skills/augustus/references/validation.md | 16 +- CHANGELOG.md | 18 ++ README.md | 10 +- docs/ecosystem.md | 25 +- research/archive/findings.md | 46 +++ research/notes.md | 266 ++++++++++++++++++ research/refresh-log.md | 22 ++ research/sources.json | 104 ++++++- 19 files changed, 642 insertions(+), 33 deletions(-) diff --git a/.agents/skills/augustus/SKILL.md b/.agents/skills/augustus/SKILL.md index b25df0c..93d4c2d 100644 --- a/.agents/skills/augustus/SKILL.md +++ b/.agents/skills/augustus/SKILL.md @@ -1,6 +1,6 @@ --- name: augustus -description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill, open multimodal RLCD / blackwood), announced open decision-model (Watch — still not landed), constrained-AR (TypeAR, pcdServer, decision-token LoRA), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"wait for Archer vs open multimodal RLCD\", \"screenshot/DOM candidates → typed Choice\", \"eval path\", \"jevals\", \"Harbor taskset\", \"shared bake-off ECE/NLL/Brier\", \"LLM-as-judge is not the System One score\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge (allowlist *proves*; fail-open cannot block), \"missing other → confident wrong Choice\", \"lint the Jev request\", \"training confronts Choice other / none-of-the-above\", \"S1 reflex keeps control / optional S2 one-use advice\", \"soft AGENTS.md rules vs linter\" (Abide / jev-pref), \"edit-phase vs turn-phase observation window\", banded confidence fail-open preference lint, TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, or \"formally verify with Jev\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." +description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill, open multimodal RLCD / blackwood), announced open decision-model (Watch — still not landed), constrained-AR (TypeAR, pcdServer, decision-token LoRA), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"wait for Archer vs open multimodal RLCD\", \"screenshot/DOM candidates → typed Choice\", \"eval path\", \"jevals\", \"Harbor taskset\", \"shared bake-off ECE/NLL/Brier\", \"LLM-as-judge is not the System One score\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge (allowlist *proves*; fail-open cannot block), \"missing other → confident wrong Choice\", \"lint the Jev request\", \"training confronts Choice other / none-of-the-above\", \"S1 reflex keeps control / optional S2 one-use advice\", \"soft AGENTS.md rules vs linter\" (Abide / jev-pref), \"edit-phase vs turn-phase observation window\", banded confidence fail-open preference lint, \"extractive selection / pointer-not-generator\", \"local /v1/systemone drop-in (stub until hf scorer)\", \"dataframe-native semantic columns\", \"route≠memory\", \"advisory sidecar receipts\", \"structure induction over bags\", \"AST ∩ semantic lint\", TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, or \"formally verify with Jev\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." license: MIT metadata: version: 0.3.0 @@ -24,7 +24,7 @@ the official `typesafe-ai` skill plus the live docs own Jev integration contracts — read them before writing Jev API code. Neighbor skills `tenbin` (lint/measure) and `decision-first` (try-Jev-first habit) own their jobs. Do not collapse into a TypeSafe how-to, a Laya install, a kev -serve, a blackwood vLLM how-to, or a GLiClass or GLiNER tutorial. +serve, a blackwood vLLM how-to, a jev-local Docker install, or a GLiClass or GLiNER tutorial. Pick the **pillar** from the hole (expected utility, VOI, MCDA, signal detection, search/control, org/safety, formal methods), then the @@ -54,7 +54,7 @@ classical method you already trust, substitute it, classify the win "paraphrase brittleness", "allowlist then judge", "TOCTOU-of-Noul", "Jev inside the database / sqlite-jev", "Jev picks bitrate / join order / the model", "wait for Archer", "lint the request / missing - other", "training confronts Choice other / none-of-the-above", "soft AGENTS.md rules vs the linter", or "screenshot Choice / omni System One": + other", "training confronts Choice other / none-of-the-above", "soft AGENTS.md rules vs the linter", "screenshot Choice / omni System One", "extractive quotes / pointer not generator", "local drop-in vs stub scorer", or "route vs memory": read `references/faq.md`, then `references/mental-models.md`, then `references/mixed-architecture.md`, then @@ -115,15 +115,15 @@ classical method you already trust, substitute it, classify the win |---|---|---| | Mental models across domains (not SWE-only) | EU, abstention, VOI, MCDA, SDT, search/control, Leveson, NATM/Norman/snap-fit | `references/mental-models.md` | | Judgment-model class (Jev is exemplar, not monopoly) | Species: decide / locate (GLiNER) / categorize (GLiClass) / rank / perceive; open heads include encoder DeBERTa, LoRA distill, openjev-lm, kev | `references/judgment-class.md` | -| Open weights vs constrained decoding vs proprietary API | Three open paths: encoder open-jev / AR constrained decode (TypeAR + pcdServer; decision-token LoRA) / trained decision-only (Laya, Nimble, kev, **blackwood-rlcd** multimodal now, Archer Watch still Watch). Softmax over allowed tokens ≠ Noul | `references/judgment-class.md` (when-to-use table) | +| Open weights vs constrained decoding vs proprietary API | Three open paths: encoder open-jev / AR constrained decode (TypeAR + pcdServer; decision-token LoRA) / trained decision-only (Laya + ONNX port, Nimble, kev, **blackwood-rlcd** multimodal now, Archer Watch still Watch). Contract-compatible local `/v1/systemone` (jev-local) is a drop-in *surface*; default scorer is a stub until `hf`. Softmax over allowed tokens ≠ Noul | `references/judgment-class.md` (when-to-use table) | | Entropy as allocator (low / medium / high) | Typed low+medium decisions → System One marginals; high-entropy synthesis → frontier decoder. Product rhetoric, not a meter. **Hypothesis** | `references/judgment-class.md` | | Formal / semi-formal (proof vs judgment) | Sensor vs constraint vs searchlight; Alloy vs Apalache; DST trio; TOCTOU-of-Noul, AI×FM | `references/formal-methods.md` (one-screen: `references/formal-semi-formal.md`) | -| Mixed architecture (judgment model + LLM) | Provider judges, LLM writes, code owns control; not a stack replacement | `references/mixed-architecture.md` | +| Mixed architecture (judgment model + LLM) | Provider judges, LLM writes, code owns control; not a stack replacement. Advisory sidecar never changes host routing | `references/mixed-architecture.md` | | Context sieve | Relevance Noul per block; always-keep set in code; stub + recall key | `references/applied-mappings.md#1-context-sieve` | -| Exact-text keep / drop | Choice include/exclude/mixed over candidates code already holds | `references/applied-mappings.md#2-exact-text-keep--drop` | +| Exact-text keep / drop | Choice include/exclude/mixed over candidates code already holds. Extractive quotes / pointer-not-generator (model never writes the excerpt) | `references/applied-mappings.md#2-exact-text-keep--drop` | | Environment / harness triage | Scan every step for env failure; LLM autopsy only on flags | `references/applied-mappings.md#3-environment--harness-triage` | | Moderation and ranking | Hold-before-publish vs graded rerank; fail policy per action | `references/applied-mappings.md#4-moderation-and-ranking` | -| Skill / tool routing | Choice over a closed catalog + whether-anything-fits; code dispatches | `references/applied-mappings.md#5-skill--tool-routing` | +| Skill / tool routing | Choice over a closed catalog + whether-anything-fits; code dispatches. Route ≠ memory: cheap intent gate skips memory tours on easy routes | `references/applied-mappings.md#5-skill--tool-routing` | | Expensive observation router | Structural prove (text layer) ∩ remainder Noul (needs OCR?) | `references/applied-mappings.md#6-expensive-observation-router` | | Agent preference lint / semantic gates | Soft project rules as criteria; linter owns hard rules; one Score per rule on the diff; bands + fail-open; name the observation window (edit vs turn) | `references/mixed-architecture.md#preference-lint-and-gates` | | Dual orchestration (Jev ∩ LLM ∩ MCP) | Jev-as-tool vs Jev-as-outer-loop; schemas are exact state | `references/mixed-architecture.md#dual-orchestration-jev--llm--mcp` | @@ -132,7 +132,7 @@ classical method you already trust, substitute it, classify the win | Selective classification / decision theory | Thresholds from action costs, abstention paths | `references/mappings.md#2-probabilistic-judgments--cost-sensitive-decisions` | | Decision tables / circuits / state machines | Judgment predicates, code owns transitions | `references/mappings.md#3-semantic-predicates--decision-circuits` | | Retrieve + expensive relevance fn | Bounded rerank of a retrieved shortlist | `references/mappings.md#4-retrieval--bounded-semantic-reranking` (independent TREC DL2019 benchmark: Jev zero-shot best MAP 0.4748, nDCG@10 0.683 vs tuned monoBERT 0.718 — competitive, not dominant) | -| Store as semantic index (SQL / SQLite / zoxide) | Cheap exact predicates first; typed questions on the remainder. In-engine extension (sqlite-jev) vs CLI rewrite (jevql) vs path index (joxide) | `references/mappings.md#4-retrieval--bounded-semantic-reranking` | +| Store as semantic index (SQL / SQLite / zoxide / dataframe) | Cheap exact predicates first; typed questions on the remainder. In-engine extension (sqlite-jev) vs CLI rewrite (jevql) vs path index (joxide) vs dataframe columns (jevpandas / jevframe) | `references/mappings.md#4-retrieval--bounded-semantic-reranking` | | Soft judgment inside a hard envelope | Model may only match the deterministic policy or be more conservative (bitrate ABR; query-planner override-when-confident) | `references/mappings.md#12-runtime-assurance-sandwich-hypothesis`; `references/mappings.md#18-structural-prove--soft-remainder-hypothesis-as-domain-general-empirical-as-named-shapes` | | Value of information / gather as an act | Pay for another observation only if EV(decision) improves more than cost; abstain from calling *any* model when a regex already answers (meta-VOI) | `references/mappings.md#6-value-of-information--gather-as-an-enumerated-act` (**Hypothesis** until a labeled act/outcome log; 149-row receipt is Empirical as a shape) | | Signal detection / ROC | Criterion and operating point from costs and base rate, not accuracy | `references/mappings.md#7-signal-detection--criterion-not-accuracy` (**Hypothesis** for non-SWE plots) | @@ -150,8 +150,8 @@ classical method you already trust, substitute it, classify the win | Effect-oriented state-machine loops | Soft predicates on transitions; code owns the transition | `references/mappings.md#19-effect-oriented-state-machine-loops-hypothesis` (**Hypothesis**; ZIO client, not Effect.ts) | | Agent self-supervision / on-track detection | Pre-gate → output judge → done-check → supervisor nouls. S1 reflex keeps control; optional S2 is one-use advice | `references/agent-self-assessment.md` | | Optimizer/program frameworks (Ax, DSPy) | Typed fields → one provider request; judge metrics; threshold discipline. Ax and DSPy climb LM-program knobs only | `references/optimizer-integration.md` | -| Perception → decision pipeline / measure / hill-climb | Stages with a versioned state contract; frozen taskset; DSPy/Ax only on the LM-program slice. Specialist composition stays **Hypothesis**; open multimodal decide (blackwood-rlcd) is a named receipt. Same section as the row below | `references/validation.md#eval--hill-climb` | -| Eval & hill-climb | Decision-stage jevals hygiene; Harbor taskset × harness × runtime; one score-composition table. Shared bake-off exemplar: open-jev-laya-bench (ECE/NLL/Brier; LLM-as-judge is not the score) | `references/validation.md#eval--hill-climb` | +| Perception → decision pipeline / measure / hill-climb | Stages with a versioned state contract; frozen taskset; DSPy/Ax only on the LM-program slice. Specialist composition stays **Hypothesis**; open multimodal decide (blackwood-rlcd) is a named receipt. Structured observe→decide→verified-act (no screenshots) is a computer-use speed-layer receipt. Same section as the row below | `references/validation.md#eval--hill-climb` | +| Eval & hill-climb | Decision-stage jevals hygiene; Harbor taskset × harness × runtime; one score-composition table. Shared bake-off exemplar: open-jev-laya-bench (ECE/NLL/Brier; LLM-as-judge is not the score). Collab-arm curriculum: llm_autonomous vs scripted_plus_jev vs llm_plus_jev (Wilson / McNemar) | `references/validation.md#eval--hill-climb` | | (meta) Finding new mappings & applications | Toolbox sweep: judgment-shaped component of a known method, substituted + falsified | `references/toolbox-mapping.md` | | Named methods / operators / theorems | Substitution tiers: operand-judgments, preconditioned theorems, non-substitutable | `references/methods-catalog.md` | | (meta) Where a judgment model sits relative to any construct | 11 positions + logical-operator rules + position×construct traversal as the application generator | `references/composition-algebra.md` | diff --git a/.agents/skills/augustus/references/agent-self-assessment.md b/.agents/skills/augustus/references/agent-self-assessment.md index 3257d34..8b8cbb8 100644 --- a/.agents/skills/augustus/references/agent-self-assessment.md +++ b/.agents/skills/augustus/references/agent-self-assessment.md @@ -34,7 +34,12 @@ relying: foreman, pi-jev, pi-warden, winnow, fast-jev-compaction, jev-judgment. [jev-reflex-autonomy-lab](https://github.com/khordoo/jev-reflex-autonomy-lab): **S1 keeps control**; optional S2 is one-use advice on low confidence and does not fly the drone (`notes.md` §46). Experimental viz, not a - production supervisor. + production supervisor. Computer-use speed layer of the same split: + [solari-reflex](https://github.com/hitakshiA/solari-reflex) — one + structured observation → one typed decision → one verified action; + **no screenshots**; model output never becomes a selector. Harbor-style + task score (Stripe API / answer key). Author table vs Codex on Solari: + 60.2 s vs 194.9 s; 66 s vs 460 s; 24.2 s vs 98.4 s (`notes.md` §48). 6. **Context economy**: the context-sieve card (`references/applied-mappings.md#1-context-sieve`). Judge every large tool result with one relevance Noul before it enters context. Hide @@ -66,6 +71,9 @@ relying: foreman, pi-jev, pi-warden, winnow, fast-jev-compaction, jev-judgment. - Grounding of a generated claim: one Choice per claim–evidence pair (supports / contradicts / unrelated) + a confidence review flag; judge against the cited source text, never against another model's prose. + Pointer-not-generator: the model points at line ids; code copies + verbatim with place; *not found* is an answer + ([jev-reviewer](https://github.com/choxos/jev-reviewer); `notes.md` §48). - Self-report fidelity: compare the agent's claimed action with its actual trace via decomposed Nouls (right tool? args match schema? result matches call?). Escalate on low confidence; never auto-retry. diff --git a/.agents/skills/augustus/references/applied-mappings.md b/.agents/skills/augustus/references/applied-mappings.md index 52b0ec6..c1d89a4 100644 --- a/.agents/skills/augustus/references/applied-mappings.md +++ b/.agents/skills/augustus/references/applied-mappings.md @@ -86,6 +86,16 @@ picks among **letters drawn on the screenshot**; code still clicks (`notes.md` §46). Text-only cousin: [jev-e2e](https://github.com/perixtar/jev-e2e) — Jev selects observed controls; Playwright independently checks; a confident model cannot substitute for checked expectations. +**Extractive quotes (Empirical as named receipts, 2026-09-18 ~14:52):** +[testimonial-miner](https://github.com/AppitStudio/testimonial-miner) — +code numbers sentences; one broadcast (Choice/Noul/Score + per-sentence +Nouls); the model never writes; `redecide` retunes thresholds on the +log. [jev-reviewer](https://github.com/choxos/jev-reviewer) — the model +**points at line ids**; code copies verbatim quotes with place; *not +found* is an answer. Computer-use cousin: +[solari-reflex](https://github.com/hitakshiA/solari-reflex) — structured +observation → typed decision → verified act; **no screenshots**; model +output never becomes a selector (`notes.md` §48). **Counterexample**: "write the patch that matches this sentence" — that is generation. **Test**: every kept byte occurs in the input; mixed never auto-included; snapshot stale → abort, don't guess. @@ -204,6 +214,10 @@ shortlist; **not** `kevinpita/pi-jev-context` (sieve). [`ddfeyes/jev-mode`](https://github.com/ddfeyes/jev-mode) is the latency-class split: bulk triage/tag/route off the frontier context (synthetic 1,000: −77.8% tokens; accuracy claim is **parity**). +**Route ≠ memory** ([jev-hermes](https://github.com/de-niji/jev-hermes)): +a cheap intent Choice skips memory/tool *tours* on `calendar` / `mail` / +`status`; `complex` keeps memory. Savings are skipped tours, not +turning memory off (`notes.md` §48). Toolrouter / open JevRouter: **Hypothesis** until measured on *your* catalog. **Counterexample**: the agent looping "pick a tool, call it, pick again" with the provider as diff --git a/.agents/skills/augustus/references/faq.md b/.agents/skills/augustus/references/faq.md index 423dafe..86d4b6c 100644 --- a/.agents/skills/augustus/references/faq.md +++ b/.agents/skills/augustus/references/faq.md @@ -123,7 +123,11 @@ brittleness; compose with abstention and an allowlist gate is **WATCH** until weights, license, and evals exist (`notes.md` §31, §33). Open multimodal *decide* that already shipped: [blackwood-rlcd](https://huggingface.co/BlackwoodAI/blackwood-rlcd) -(CC BY-NC; not that drop; `notes.md` §46). Constrained decoding is §32; native serving is §42. Decision-token QLoRA on that graph: +(CC BY-NC; not that drop; `notes.md` §46). A local `POST /v1/systemone` +drop-in ([jev-local](https://github.com/us/jev-local)) is a **surface**, +not a fourth path — default scorer is a stub until `hf` (`notes.md` §48). +Laya ONNX port: [laya-onnx](https://huggingface.co/Mattepiu/laya-onnx) +(do not copy the inherited vs-Jev table). Constrained decoding is §32; native serving is §42. Decision-token QLoRA on that graph: [Foodoo1/Qwen3-14B-RLCD-Decision-LoRA](https://huggingface.co/Foodoo1/Qwen3-14B-RLCD-Decision-LoRA) (train the decision token, not prose; synthetic fraud receipt). Public logit dump for the read-the-letter graph: mini-jev-runs. "Smarter than Jev" is a claim. @@ -372,6 +376,37 @@ table). Specialist composition (SAM / OCR → text → Jev) remains valid. Do not wait, and do not treat screenshot-vs-Jev-text as the same input. `judgment-class.md`; `notes.md` §46. +## Is a local `/v1/systemone` the same as Jev? + +No — not until you know **which scorer** is behind the socket. +[jev-local](https://github.com/us/jev-local) is a **contract-compatible** +drop-in (`base_url`). The **default scorer is a deterministic stub** +and carries no intelligence. `JEVLOCAL_SCORER=hf` turns on a frozen-model +logprob scorer. Their README: an interface-compatible baseline, not a +reproduction of Jev's undisclosed model. kev is the other local +drop-in (trained pointer head, public gold). A green smoke test on the +stub is not a bake-off. `judgment-class.md`; `notes.md` §48. + +## Should the model write the quote / the citation / the click? + +No. Extractive keep/drop: code already holds the sentences, line ids, +or numbered controls; the model **selects**; code **copies or clicks**. +[testimonial-miner](https://github.com/AppitStudio/testimonial-miner) +assembles quotes from per-sentence Nouls and `redecide`s without new +calls. [jev-reviewer](https://github.com/choxos/jev-reviewer) points at +ids; *not found* is an answer. [solari-reflex](https://github.com/hitakshiA/solari-reflex) +never lets model output become a selector. Generation is only for +TYPE/prose when something must be written. `applied-mappings.md` §2; +`notes.md` §48. + +## Is routing the same as memory? + +No. A cheap intent gate can skip a memory/tool *tour* on +`calendar` / `mail` / `status` without turning memory off. +[jev-hermes](https://github.com/de-niji/jev-hermes): memory still +**writes**; `complex` still searches. Route ≠ memory. +`applied-mappings.md` §5; `notes.md` §48. + ## Is confidence a trained score? No — not on Hume's reconstruction, and not as a new contract. Re-read diff --git a/.agents/skills/augustus/references/judgment-class.md b/.agents/skills/augustus/references/judgment-class.md index aafb049..7ed0c4b 100644 --- a/.agents/skills/augustus/references/judgment-class.md +++ b/.agents/skills/augustus/references/judgment-class.md @@ -438,10 +438,21 @@ unverified — do not overwrite `notes.md` §18). Shared bake-off this hour: [`pngwn/open-jev-laya-bench`](https://huggingface.co/datasets/pngwn/open-jev-laya-bench) (26 + 9 tasks, 11,959 items; ECE/NLL/Brier; **not** TypeSafe Jev vs Laya; LLM-as-judge is not the score — `notes.md` §46, `validation.md`). +**Contract-compatible local surface, not a fourth path:** +[`us/jev-local`](https://github.com/us/jev-local) speaks `POST +/v1/systemone` so an SDK `base_url` drop-in works offline. **Default +scorer is a deterministic stub** until `JEVLOCAL_SCORER=hf`. A green +smoke test on the stub is not a local decision model (`notes.md` §48). +**ONNX replica of Laya:** +[`Mattepiu/laya-onnx`](https://huggingface.co/Mattepiu/laya-onnx) +(~15 ms CPU for one Noul, their card). Do not copy the inherited +vs-Jev accuracy table (`notes.md` §18). Reject: TypeAR or pcdServer scores as fail-closed P(permit); a LoRA student's agreement with Jev as independent gold; kev's ID ECE as a license to skip a held-out test on *your* workflow; shipping on "smarter than Jev"; +a green `/v1/systemone` smoke test on jev-local's stub as a bake-off; +copying laya-onnx vs-Jev rows as independent gold; thresholding [`jp-sns-jev7-estimator`](https://huggingface.co/kokuren/jp-sns-jev7-estimator) teacher scores as P(toxic) — the card says they are **not** calibrated, and `threat` F1@0.5 is 0.0000 on their table (`notes.md` §33). Domain-local @@ -451,7 +462,7 @@ universal ranking. Diffusion beating a decision head is Hypothesis. Detail: `research/notes.md` §33 (surfaces), §34 (marginals), §35 (Nimble), §36 (diffusion), §38 (entropy allocator, Hypothesis), §42 (pcdServer serving, meta-VOI, games), §45 (kev), §46 (blackwood, -laya-bench, decision-token LoRA). Before +laya-bench, decision-token LoRA), §48 (jev-local stub, laya-onnx). Before adopting a surface, the bake-off is a jevals-shaped suite and, for a product loop, a Harbor taskset (`validation.md`, Eval & hill-climb). The stage pipeline into that decision is the same file diff --git a/.agents/skills/augustus/references/mappings.md b/.agents/skills/augustus/references/mappings.md index d2dd507..135e354 100644 --- a/.agents/skills/augustus/references/mappings.md +++ b/.agents/skills/augustus/references/mappings.md @@ -42,7 +42,12 @@ code: shortlist, rank, filter — weights adjustable without re-inference ``` **Example**: research-reading map — extract reusability dimensions once, let -researchers re-rank and re-filter interactively. **Beyond SWE +researchers re-rank and re-filter interactively. +**Offline re-threshold (Empirical as named receipts):** +[testimonial-miner](https://github.com/AppitStudio/testimonial-miner) +`redecide` reapplies `Thresholds` to logged answers with **no new model +calls** — judge once, explore policy in code (`notes.md` §48). Same family +as firehose sliders. **Beyond SWE (Hypothesis until labeled):** vendor bid/no-bid (fit, urgency, risk Nouls; price and deadline exact); apartment shortlist (commute/light/ noise Scores; rent exact); hiring scorecard (evidence Nouls; labor-law @@ -162,9 +167,13 @@ local paths only; fail-open. Dataframe cousin this hour: [`yalindogusahin/jevpandas`](https://github.com/yalindogusahin/jevpandas) — `evaluate` / `filter` / `classify` / `score` / batched `ask` over a pandas frame; classify example includes `other`; failures never become -negative predictions; LICENSE absent this pass. Row contents leave the -store (same residency warning as AU health). Do not copy SQL, env, or CLI flags. -`notes.md` §42, §44, §46. +negative predictions; LICENSE absent this pass. Accessor sibling this +hour: [`ktaletsk/jevframe`](https://github.com/ktaletsk/jevframe) (MIT, +PyPI; pandas **and** Polars `.jev`; full `p__` columns; no silent +renormalize; one row per request). Same hole, two surfaces. Row +contents leave the store (same residency warning as AU health). Do not +copy SQL, env, or CLI flags. +`notes.md` §42, §44, §46, §48. ## 5. Hierarchy → bounded heuristic search @@ -384,6 +393,13 @@ immediate win missed once reversed; Fool's-mate confidence 31%/37% so a hole on the constrained-AR surface. Not a strength rating. `notes.md` §42; `validation.md`. +**Structure induction over a bag (Empirical as a *shape*, 2026-09-18):** +[`Joymfl/dag-jev`](https://github.com/Joymfl/dag-jev) — unordered items +in, pairwise "does i depend on j?" judgments, DAG in `petgraph`. Code +owns topology; the model does not emit edges. Experiment; empty README; +no metrics this pass (`notes.md` §48). Same hole as taxonomy beam (§5): +judgment is a pairwise (or Choice) classifier step, not the scheduler. + **Query planner as the envelope (author-reported, 2026-09-18):** [@mmalisper](https://x.com/mmalisper/status/2101001041903009987) on the Join Order Benchmark. Jev picking join order was **2× slower**. @@ -613,7 +629,14 @@ one Jev call on the remainder. Empty state was self-contradictory — that is why the refuse-empty rule exists. [`affirmitv/bitrate-advisor`](https://github.com/affirmitv/bitrate-advisor) is the same sandwich on a live encoder: policy proves the cap; Jev -judges only inside it (`notes.md` §44). +may only match it or be more conservative; missing the model returns +the policy's answer. Jev judges only inside it (`notes.md` §44). +Light sibling: +[`phin-tech/pi-jev-approver`](https://github.com/phin-tech/pi-jev-approver) +— regex `commandRules` prove allow/deny (a `deny` is a hard block); +typed Score/Nouls on the remainder; **fail-closed** without a key +(different polarity from jevgate). rh-guard-adjacent; light note only +(`notes.md` §48). [`coldteadotai/abide`](https://github.com/coldteadotai/abide) is the same *family* on project instructions: the **linter proves** lintable rules; Jev Scores only residual soft AGENTS.md rules; fail-open, banded diff --git a/.agents/skills/augustus/references/mental-models.md b/.agents/skills/augustus/references/mental-models.md index 8f0e34f..d91648d 100644 --- a/.agents/skills/augustus/references/mental-models.md +++ b/.agents/skills/augustus/references/mental-models.md @@ -400,6 +400,10 @@ Use these as *existence proofs of a position*. Write your own card. | Knowledge work | what to read next | on-question Noul + quality Score | library you hold | | Hiring | interview / reject / hold | evidence Nouls; veto rules in policy | labor law, scorecards you wrote | | Inbox | reply / snooze / archive | urgency Noul + aboutness Choice | send, calendar | +| Knowledge work | extract a quote / a cited fact | per-sentence or per-line-id Noul/Choice (**Empirical**: testimonial-miner, jev-reviewer) | verbatim join; place; human publish permission | +| Dataframe labeling | classify / score rows | Noul/Choice/Score + full `p__` (**Empirical** as jevframe / jevpandas *shape*) | pandas/Polars, thresholds in code | +| Computer-use speed | one verified act per step | operation + target Choice on numbered controls (**Empirical**: solari-reflex) | Guard check; deny-list absence; no screenshots | +| Agent turn | skip memory tour on easy intent | intent Choice (**Empirical**: jev-hermes) | Memory still writes; complex still searches | | Document / lab routing | which pages need the expensive observation | Noul on remainder after a text layer / recipe | local extract, merge order (**Empirical** as OCR-router *shape*) | | Shell / tool allowlist | unlisted remainder after a **proof** | five Nouls on unknown verbs | Proven/Refused in code; cannot block (**Empirical**: jevgate) | | SWE | residual AGENTS.md / CLAUDE.md rules | one Score per named instruction-file rule | linter owns hard rules; bands + fail-open (**Empirical**: Abide replay, `notes.md` §47) | diff --git a/.agents/skills/augustus/references/methods-catalog.md b/.agents/skills/augustus/references/methods-catalog.md index f88e33d..3528a73 100644 --- a/.agents/skills/augustus/references/methods-catalog.md +++ b/.agents/skills/augustus/references/methods-catalog.md @@ -42,6 +42,8 @@ judgment component is new). |---|---|---|---|---| | MCTS / PUCT | Prune invalid actions; priors P(s,a); leaf value V(s) | Batched Noul pruning + Choice priors + Score value — depth-capped where no simulator | Tree, budget, backprop, probes | **Empirical recipe** (jev-mcts: 24/24 vs 1/24 greedy; speculative depth 2) | | Beam search over taxonomies | Which branches deserve expansion | Choice distributions as branch priority; keep K paths where ambiguity is early | Frontier, budget, final selection | **Empirical recipe** (beam K=3 cookbook) | +| Structure induction over a bag | Pairwise "does i depend on j?" (or Choice over order) | One judgment per pair; DAG / scheduler in code | Topology, cycles, execution | **Empirical as a shape** (dag-jev experiment; empty README; no metrics, `notes.md` §48) | +| Collab-arm product loop | Scripted legal set vs LLM-propose vs unconstrained | Choice over legal actions; stop on low p rather than guess | Legality, Wilson/McNemar, ceiling flags | **Empirical as a harness shape** (jev-testbench; bake into jevals/Harbor, `notes.md` §48) | | Screening / Wald sequential tests | Pass / fail / keep-looking per candidate | One Noul gate per candidate in one batched request; budget in code | Sequential rule, stop boundaries | **Hypothesis** | | STPA / STAMP control structure | Sensor reading vs enforced constraint | Judgment as sensor; constraints in policy/code/interlock; STPA table if the sensor lies | The constraint, the actuator, the probe | **Contract** as ownership; **Hypothesis** as domain product (`mappings.md` §8) | | PufferLib / Ocean env contracts | Does this episode look like a known trainer-bug mode? | Cluster failing episodes; never "the policy is correct" | Seeded serial env, Ocean sanity, observed rewards | **Hypothesis** as placement; **Contract** that Ocean is not a comparative baseline (`formal-methods.md` DST trio) | @@ -64,8 +66,10 @@ judgment component is new). | Method | Judgment-shaped component | Jev substitution | Stays in code | Status | |---|---|---|---|---| -| Claim–evidence entailment (NLI) | supports / contradicts / not-established per claim–source pair | One Choice per pair + review flag; judge against the cited source text only | Quote extraction, citation graph, audit log | **Empirical recipe** (citation_check cookbook) | +| Claim–evidence entailment (NLI) | supports / contradicts / not-established per claim–source pair | One Choice per pair + review flag; judge against the cited source text only | Quote extraction, citation graph, audit log. When the answer *is* a span you already hold, **point** at line ids and copy verbatim — the model never writes the excerpt (jev-reviewer) | **Empirical recipe** (citation_check cookbook; jev-reviewer sample study, `notes.md` §48) | +| Extractive selection + offline re-threshold | Keep/drop over sentences or ids code already numbered | Per-item Noul/Choice + one broadcast; join in order; `redecide` on the log with no new calls | Numbering, header skip, thresholds, publish permission | **Empirical recipe** (testimonial-miner 8-request fixture; jev-reviewer; `notes.md` §48). Cousin of applied-mappings §2 | | Spec vs artifact conformance (model checking *mindset*) | Property holds / violated / unverifiable for a named requirement | One Noul/Score per requirement, batched; violated → named rule back into context (pi-warden / Abide shape). This is **not** TLC/Apalache/GNATprove | Requirement enumeration, enforcement, logging; the **linter** if the rule is lintable; the real checker if you have one | **Empirical recipe** (pi-warden: 6→0 rule breaks, 150 paired runs; jev-pref: YOU define the rule; Abide: productized compile/calibrate/tune/replay, `notes.md` §47). Ownership split: `formal-methods.md` | +| AST ∩ semantic lint | Semantic remainder after a parser already extracted units | Typed questions on Tree-sitter targets; do not execute scanned code | Parser, selection, fail-on; `tenbin` owns the lint *skill* | **Empirical as a shape** (jevscan 0.2.0rc4; not a calibration claim; `notes.md` §48) | | Alloy finder vs Apalache / TLC | Which bound, which counterexample, is the property tautological? | Triage instances/CEs; never "this spec looks right" | Analyzer / SMT / explicit-state engine | **Hypothesis** as product; **Contract** as ownership (`formal-methods.md` §2) | | Type-checking analog | Does this planned call match the schema/operation/target? | Decomposed nouls over {request, schema, trace}; never trust a Jev pass as authorization | Real validation of operation+target in code | **Empirical recipe** (validation.md self-monitoring) | diff --git a/.agents/skills/augustus/references/mixed-architecture.md b/.agents/skills/augustus/references/mixed-architecture.md index 9763b5d..34f4764 100644 --- a/.agents/skills/augustus/references/mixed-architecture.md +++ b/.agents/skills/augustus/references/mixed-architecture.md @@ -272,7 +272,12 @@ asynchronous — the reflex does not pause experimental drone viz, not a flight controller; GitHub license null this pass). Same Kahneman split as the toolbox row (S2 proposes, S1 discriminates; never the reverse). `notes.md` §46. -`agent-self-assessment.md`. +`agent-self-assessment.md`. **Route ≠ memory** is the same split on a +turn: [jev-hermes](https://github.com/de-niji/jev-hermes) cheap-gates +calendar/mail/status off the memory tour; complex keeps Honcho +(`notes.md` §48). **Advisory sidecar:** +[agent-workflow-typesafe-ai](https://github.com/ngallodev-software/agent-workflow-typesafe-ai) +emits `no_action` receipts and **never** changes host routing. **Effect-oriented loop (same author, later post).** Topology B inside an effect system @@ -391,7 +396,14 @@ decision-design card. Do not clone APIs from READMEs. | S1 reflex + optional S2 advice | Typed action Choice; planner one-use on low p | Collision, legality, the stick stays with S1 | jev-reflex-autonomy-lab (experimental) | | Decision-as-business-tool | Named judgment; gate is part of the result | Registry, arithmetic, hard guards | jev-decision-layer (unofficial) | | NL cases → checked e2e | Jev selects observed controls | Playwright expectations; PASS/FAIL/BLOCKED | jev-e2e (alpha) | -| Dataframe semantic index | Noul / Choice / Score per row | pandas, thresholds, never invent negatives | jevpandas | +| Extractive quotes / pointer evidence | Per-sentence or per-line-id Noul/Choice | Verbatim join; place; `redecide` / CSV; model never writes the excerpt | testimonial-miner; jev-reviewer | +| Structured observe → decide → act | Operation + target Choice on numbered controls | Guard check; deny-list absence; no screenshots; TYPE is the only generation | solari-reflex | +| Dataframe semantic columns | Noul / Choice / Score per row; full `p__` | pandas/Polars, indexes, never silent renormalize | jevpandas; jevframe (PyPI + Polars) | +| Route ≠ memory | Intent Choice before a turn | Config + flat tools on easy routes; memory stays on for hard ones | jev-hermes | +| Advisory sidecar receipts | Typed answers as `no_action` evidence | Host routing / executor / policy unchanged | agent-workflow-typesafe-ai | +| Structure induction over a bag | Pairwise dependency Noul/Choice | DAG / scheduler in code | dag-jev (experiment) | +| Simulated world control vs content | Intent / page-type Choice | Generator writes documents; Zod + deterministic compiler; SQLite world | jev-agentworld-web-simulator | +| AST ∩ semantic lint | Typed questions on Tree-sitter units | Parser, selection, fail-on; does not execute scanned code | jevscan (`tenbin` owns the lint skill) | | Model router | Requirement Scores; policy in code | Eligibility, cost/quality/latency objective | routeKit | | Bulk-judgment coprocessor | Choice/Noul off the frontier context | Counts, policy, fail-open gate | jev-mode | | Closed-catalog System One shell | Choice over host tools | Execute, arithmetic, credentials | jot | @@ -428,7 +440,12 @@ shim, CC BY-NC; Jev still leads general text; not Archer Watch decision-model drop is **Watch**. Closed calibrated API vs open weights is a self-eval tradeoff (`research/notes.md` §18, §33, §45). When-to-use axes: `judgment-class.md`. TypeSafe remains the documented *exemplar*, -not the class monopoly. GLiNER (locate) / GLiClass (categorize) / +not the class monopoly. **Local contract drop-in this hour:** +[`us/jev-local`](https://github.com/us/jev-local) speaks `/v1/systemone`; +**default scorer is a deterministic stub** until `JEVLOCAL_SCORER=hf` +(`notes.md` §48). **ONNX replica path:** +[`Mattepiu/laya-onnx`](https://huggingface.co/Mattepiu/laya-onnx) — do +not copy the inherited vs-Jev table. GLiNER (locate) / GLiClass (categorize) / GLiNER2.5 (local multi-head), listwise, and vision families: `judgment-class.md`. diff --git a/.agents/skills/augustus/references/question-design.md b/.agents/skills/augustus/references/question-design.md index 43a21c3..98077e6 100644 --- a/.agents/skills/augustus/references/question-design.md +++ b/.agents/skills/augustus/references/question-design.md @@ -17,7 +17,7 @@ request, and treat a stale pin as a prior, never a setting. 2. One question per judgment. Split any question that weighs two properties. 3. Pick the primitive whose answer code acts on directly. 4. Build the smallest state that answers every question; compute in code whatever code can compute. -5. Put every question sharing the state into **one request** (speculative fan-out — parallel questions cost little latency; code ignores unneeded answers). Second requests only when later data depends on an earlier answer. +5. Put every question sharing the state into **one request** (speculative fan-out — parallel questions cost little latency; code ignores unneeded answers). Second requests only when later data depends on an earlier answer. Extractive / pointer: number the candidates in **code**; ask per-id Noul/Choice; copy verbatim. "Not found" is an option. The model never writes the quote (`applied-mappings.md` §2; `notes.md` §48). 6. Combine in code: branches, weights, confidence gates. 7. Test on labeled examples; read `probabilities` on the misses; revise one or two questions at a time. diff --git a/.agents/skills/augustus/references/toolbox-mapping.md b/.agents/skills/augustus/references/toolbox-mapping.md index f5ab5c0..2154f9e 100644 --- a/.agents/skills/augustus/references/toolbox-mapping.md +++ b/.agents/skills/augustus/references/toolbox-mapping.md @@ -70,10 +70,12 @@ component; keep the rest of the method in code. | Search: value function | Score rubric as leaf value V(s) — only where a simulator validates outcomes; speculative depth hard-capped at 2 | **Empirical recipe** (jev-mcts fidelity split) | | Measurement theory: probe vs estimate | Only post-execution probes concede milestones; model estimates never do — "estimation wearing a measurement costume" is the rejection template | **Empirical recipe** (jev-mcts, pi-warden done-check) | | Experimental design: perturbation | Behavioral tests as the stats layer: candidate removal, option-order shuffle, letter-shuffle on screenshot Choice, distractor injection, boundary cases | **Contract-level** (validation.md); letter-shuffle receipt: blackwood-rlcd 0.133 vs Jev 1.13 0.587 on 300 web steps (`notes.md` §46) | +| Experimental design: collab arms | `llm_autonomous` vs `scripted_plus_jev` vs `llm_plus_jev`; Wilson + McNemar; the decision model is **not** a peer arm | **Empirical as a harness shape** (jev-testbench; bake into jevals/Harbor, `notes.md` §48) | | Discrete math: width vs depth | Fan out in width (parallel ≈ free), pay depth linearly; two-stage only when next options depend on an earlier answer | **Empirical recipe** (fan-out: 12.2× cheaper, 10× faster) | -| Psychology: Kahneman | System 2 generates/proposes (LLM), System 1 discriminates (Jev); never the reverse. S1 keeps control; optional S2 is one-use advice and does not fly | **Empirical recipe** (mcts-agent role split; 2026-09-18 mixed-architecture discourse; jev-reflex-autonomy-lab as a control-loop shape, `notes.md` §46) | +| Psychology: Kahneman | System 2 generates/proposes (LLM), System 1 discriminates (Jev); never the reverse. S1 keeps control; optional S2 is one-use advice and does not fly | **Empirical recipe** (mcts-agent role split; 2026-09-18 mixed-architecture discourse; jev-reflex-autonomy-lab as a control-loop shape, `notes.md` §46). **Route ≠ memory:** a cheap intent gate skips memory/tool *tours* on easy routes; memory still writes; complex still searches (jev-hermes, `notes.md` §48) | | IR / cascades | Cheap relevance / irrelevance before an expensive ranker or generator | **Empirical recipe** (RAG cookbook; jevprune; git-jev-stage; LlamaIndex Jev rerank) | -| Spec / lint | Project-defined semantic rules as predicates over a diff; linter owns hard rules | **Empirical recipe** (jev-pref contract; Abide productized path — replay 93 sessions, edit precision ~26% / turn ~73% before tune, `notes.md` §47; JevLint file-level Noul; pi-warden; snifftest unsure-band) | +| IR / extractive keep-drop | Number candidates in code; model selects; copy verbatim; `redecide` thresholds on the log with no new calls | **Empirical recipe** (testimonial-miner; jev-reviewer pointer-not-generator; `notes.md` §48) | +| Spec / lint | Project-defined semantic rules as predicates over a diff; linter owns hard rules. AST remainder: Tree-sitter units, then typed questions; do not execute scanned code | **Empirical recipe** (jev-pref contract; Abide productized path — replay 93 sessions, edit precision ~26% / turn ~73% before tune, `notes.md` §47; JevLint file-level Noul; pi-warden; snifftest unsure-band; jevscan AST∩semantic, `tenbin` owns the lint skill, `notes.md` §48) | | Formal methods / DST / safety | Judgment triages counterexamples, failing seeds, and named-rule conformance; proof/MC/DST stay with their tools. Alloy finder ≠ Apalache BMC ≠ Quint run. DST trio: Antithesis hypervisor / Resonate HQ Lean+oracle+SDK (durable async) / PufferLib env+seed. Noul is a sensor, not a discharged PO. Semi-formal diagrams are vocabularies, not enforcers | **Hypothesis as product**, **Contract** as ownership (matching `mappings.md` §8 and `methods-catalog.md`; worked shape pi-warden — `formal-methods.md`, `formal-semi-formal.md`) | | Decision analysis: VOI | Gather as an enumerated act; pay iff expected decision-loss drop > cost | **Hypothesis** as calculator (`mappings.md` §6; `mental-models.md`) | | Signal detection | Noul as evidence variable; criterion from costs and base rate; ROC/PR on your labels | **Hypothesis** for non-SWE plots (`mappings.md` §7) | diff --git a/.agents/skills/augustus/references/validation.md b/.agents/skills/augustus/references/validation.md index b5319a0..8fec564 100644 --- a/.agents/skills/augustus/references/validation.md +++ b/.agents/skills/augustus/references/validation.md @@ -321,6 +321,7 @@ Rules: | LM-program knobs only | DSPy/Ax (narrow) | never primary System One calibration score | | Reward-hack / eval gaming | [rh-guard](https://github.com/24601/rh-guard) | structural deny + System One sidecar | | Project soft-rule lint | [Abide](https://github.com/coldteadotai/abide) | Score per rule on the diff; bands; fail-open; replay + independent review | +| Collab / computer-use product loop | [jev-testbench](https://github.com/ufx7/jev-testbench); [solari-reflex](https://github.com/hitakshiA/solari-reflex) | Wilson/McNemar arms; independently checked task time | rh-guard is a reward-hack hook, a different surface from jevgate and from Abide (eval-integrity vs allowlist-remainder vs project soft @@ -333,12 +334,25 @@ confirmation), not a Harbor taskset and not a jevals substitute. Turn-phase soft rules held up better; false positives mostly fixable in the rubric. Text/diff only. +**Harbor-style computer-use receipt this hour:** +[solari-reflex](https://github.com/hitakshiA/solari-reflex) scores +the *task* (Stripe API / answer key), not a paragraph judge. Observe +→ decide → verified act; no screenshots. Author table vs Codex on +the same Solari machines: 60.2 s vs 194.9 s; 66 s vs 460 s; 24.2 s +vs 98.4 s (`notes.md` §48). **Collab-arm curriculum:** +[jev-testbench](https://github.com/ufx7/jev-testbench) — +`llm_autonomous` vs `scripted_plus_jev` vs `llm_plus_jev`; Wilson + +McNemar; Jev is not a peer arm. Bake into jevals/Harbor hygiene, do +not copy the harness. + ### Bake-off mandate Before adopting proprietary Jev vs Laya vs TypeAR vs Nimble vs kev vs blackwood-rlcd vs Archer vs openjev-lm, run a jevals-shaped labeled suite (or an equivalent with this hygiene) and, for a product loop, a Harbor taskset. A design -card with no eval path is incomplete. +card with no eval path is incomplete. A green smoke test on +[jev-local](https://github.com/us/jev-local)'s **default stub** is not +that bake-off (`notes.md` §48). **Shared bake-off exemplar (Empirical as that named receipt, not a ranking).** [`pngwn/open-jev-laya-bench`](https://huggingface.co/datasets/pngwn/open-jev-laya-bench) diff --git a/CHANGELOG.md b/CHANGELOG.md index 976a694..9ebdcd8 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -162,6 +162,24 @@ Pegged against [`typesafe-ai/skills` v0.5.7](https://github.com/typesafe-ai/skil wording, dedicated `none_of_the_above` eval (no published rates). Cross-link wellposed request-shape lint. No species change. No wrapper. +- Hourly ~14:52 Boise fold (`research/notes.md` §48): Archer still + Watch. Extractive selection + offline `redecide` + ([testimonial-miner](https://github.com/AppitStudio/testimonial-miner)); + pointer-not-generator + ([jev-reviewer](https://github.com/choxos/jev-reviewer)). Local + `/v1/systemone` drop-in ([jev-local](https://github.com/us/jev-local); + default scorer is a stub until `hf`). Observe→decide→verified-act, + no screenshots ([solari-reflex](https://github.com/hitakshiA/solari-reflex); + 60.2/194.9, 66/460, 24.2/98.4 s vs Codex on Solari). Dataframe + accessor sibling ([jevframe](https://github.com/ktaletsk/jevframe); + note jevpandas). Route ≠ memory (jev-hermes). Advisory sidecar + (agent-workflow-typesafe-ai). Structure induction (dag-jev experiment). + Decision-for-control / generator-for-content (jev-agentworld-web-simulator). + Collab arms + Wilson/McNemar (jev-testbench). AST ∩ semantic (jevscan; + `tenbin` owns lint). Light Pi gate (pi-jev-approver). Laya ONNX port + ([laya-onnx](https://huggingface.co/Mattepiu/laya-onnx); do not copy + vs-Jev table). Spotcheck: SemIf 1551★; jevlike 866★; tracker + 20:12:57Z still lists Laya, not Blackwood. No wrapper. - Effect-oriented loops (`notes.md` §28, `mappings.md` §19): Ward's ZIO client keeps Jev as the outer Choice and the handler as the effect. Not Effect.ts. GLiNER author: GLiNER2 "like jev" is GLiGuard diff --git a/README.md b/README.md index 245671b..2cf4a0a 100644 --- a/README.md +++ b/README.md @@ -31,7 +31,8 @@ never launder a Noul as a proof. - `.agents/skills/augustus/references/judgment-class.md` — the class (Jev exemplar, not monopoly): open heads (Laya, kev, encoder DeBERTa, LoRA distill), constrained-AR (TypeAR, pcdServer), announced decision-model (Watch), - open multimodal RLCD (blackwood-rlcd; not Archer), + open multimodal RLCD (blackwood-rlcd; not Archer), Laya ONNX port, + contract-compatible local `/v1/systemone` (stub until hf scorer), GLiNER/GLiClass species (locate vs categorize vs local multi-head), listwise vs decision objectives, vision scoring, when-to-use axes, agent-architecture portents @@ -48,10 +49,10 @@ never launder a Noul as a proof. placement: judgment-class model + LLM + code; preference lint; provider (Jev default / other family with self-eval) - `.agents/skills/augustus/references/applied-mappings.md` — context sieve, - exact-text keep/drop, env triage, moderation/ranking, skill routing + exact-text keep/drop (extractive / pointer-not-generator), env triage, moderation/ranking, skill routing (route ≠ memory) - `.agents/skills/augustus/references/faq.md` — "just classification", stack replacement, Jev vs open head vs encoder vs LoRA vs constrained AR vs kev vs blackwood, - wait-for-Archer, missing-other confident-wrong, soft project rules vs linter (Abide), GLiNER vs GLiClass vs CLIP, LLM-as-judge, in-engine vs CLI store, + wait-for-Archer, missing-other confident-wrong, soft project rules vs linter (Abide), extractive/pointer-not-generator, local drop-in vs stub scorer, route ≠ memory, GLiNER vs GLiClass vs CLIP, LLM-as-judge, in-engine vs CLI store, hard envelope (bitrate / planner), not-another-how-to - `.agents/skills/augustus/references/mappings.md` — classical-method mappings with boundaries, counterexamples, acceptance tests (including @@ -60,7 +61,8 @@ never launder a Noul as a proof. recipes, Jev-for-skills (routing, self-monitoring, testing, modularity, frontmatter), and Eval & hill-climb (jevals hygiene + Harbor taskset; open-jev-laya-bench as ECE/NLL/Brier bake-off exemplar; Abide replay as - Harbor-adjacent soft-rule measurement) + Harbor-adjacent soft-rule measurement; solari-reflex Harbor-style + computer-use; jev-testbench collab arms) - `.agents/skills/augustus/references/boundary-audit.md` — existing-system insertion: fit test, opportunity map, smallest boundary, red flags - `.agents/skills/augustus/scripts/evaluate_decisions.py` — offline evaluator diff --git a/docs/ecosystem.md b/docs/ecosystem.md index b6dd4eb..626a72e 100644 --- a/docs/ecosystem.md +++ b/docs/ecosystem.md @@ -29,7 +29,7 @@ weekdays. Jev is the densest public corpus, not the class monopoly. - **khordoo/jev-reflex-autonomy-lab** — S1 Jev reflex keeps control; optional S2 planner is one-use advice on low confidence. Experimental viz, not a flight controller. `notes.md` §46. - **perixtar/jev-e2e** — NL cases; Jev selects observed controls; Playwright independently checks. PASS/FAIL/BLOCKED. Alpha. `notes.md` §46. - **Wany-i/jev-decision-layer** — business decision tool; caller names the judgment; `gate` is part of the result. Unofficial. `notes.md` §46. -- **yalindogusahin/jevpandas** — pandas semantic index; noul/choice/score; LICENSE absent this pass. `notes.md` §46. +- **yalindogusahin/jevpandas** — pandas semantic index; noul/choice/score; LICENSE absent this pass. `notes.md` §46. Accessor sibling: **ktaletsk/jevframe** (PyPI; pandas and Polars `.jev`; full `p__`). `notes.md` §48. - **Friedjof/jev-mobile** — durable Android worker + Mobile MCP; Jev sees prevalidated candidates only. `notes.md` §33. - **jcpsimmons/jev-macos-loop** — Apple-silicon computer-use; local OmniParser/OCR/AX; text-only Jev. Finder demo independently verified. - **rajdhakad9826/routeKit** — Jev estimates task requirements; policy engine selects the LLM. Jev does not pick the model. @@ -53,7 +53,7 @@ Per-keystroke launchers (104ms median, sequence-tagged staleness), firehose mode - **GLiNER / GLiNER2.5 / GLiClass** — species map: locate spans vs categorize the sequence vs local multi-head (fastino-ai GLiNER2.5 CPU-first). Peer of Jev, not a footnote. `references/judgment-class.md`. Author primary source: GLiNER2 "like jev" is schema-conditioned categorize (GLiGuard), not a Noul (`notes.md` §28). 36× Browser Use claim is a tweet (`notes.md` §25). - **GLiGuard** (fastino-ai) — 0.3B GLiNER2 encoder, checkpoint `fastino/gliguard-LLMGuardrails-300M`. One bidirectional pass over a safety schema. Same interface shape as batched questions; different objective. Not a Jev weight clone. `judgment-class.md`; `notes.md` §30. - **DECRUX9812/openjev-lm** — Qwen2.5-0.5B+LoRA distilled from hosted Jev answers; 65/70 = 92.9% on 70 hand-labelled rows (one annotator, one domain, one seed) overnight on 6 vCPU, $0/call. Its 98.1% on fresh rows is teacher *agreement*, not gold. Receipts pattern: `notes.md` §25, §44. -- **convaiinnovations/laya** — open Choice/Score/Noul head, text-only, 512 tok. Companion packaging this hour: [`laya-typed-decisions`](https://huggingface.co/convaiinnovations/laya-typed-decisions) (421.3M, acc 0.766 / Brier 0.066 unverified). Shared bake-off: [`pngwn/open-jev-laya-bench`](https://huggingface.co/datasets/pngwn/open-jev-laya-bench) (26+9 tasks, 11959 items; ECE/NLL/Brier; not TypeSafe Jev vs Laya). `notes.md` §18, §42, §46. +- **convaiinnovations/laya** — open Choice/Score/Noul head, text-only, 512 tok. Companion packaging this hour: [`laya-typed-decisions`](https://huggingface.co/convaiinnovations/laya-typed-decisions) (421.3M, acc 0.766 / Brier 0.066 unverified). Shared bake-off: [`pngwn/open-jev-laya-bench`](https://huggingface.co/datasets/pngwn/open-jev-laya-bench) (26+9 tasks, 11959 items; ECE/NLL/Brier; not TypeSafe Jev vs Laya). ONNX replica: [`Mattepiu/laya-onnx`](https://huggingface.co/Mattepiu/laya-onnx) (~15 ms CPU; do not copy vs-Jev table). `notes.md` §18, §42, §46, §48. - **jaredpalmer/kev** — Qwen2.5-0.5B LoRA + pointer readout; Apache-2.0; Hub [`jaredpalmer/kev-0.5b`](https://huggingface.co/jaredpalmer/kev-0.5b) plus GitHub release tarball. Runnable Archer reconstruction (`POST /v1/systemone`). Public gold, not a Jev teacher. Isolation exact; ID ECE 0.065 (0.031 after T); acc 0.799 / 1,350. NOTA training must confront `"other"` as a wrong alternative (`notes.md` §45 delta). Not a knowledge/frontier substitute. - **BlackwoodAI/blackwood-rlcd** — open multimodal RLCD (image-text-to-text), Jev-compatible shim, CC BY-NC 4.0. Screenshot + marked candidates → Choice. Card: web acc 0.907 vs Jev 1.13 text-only 0.480; letter-shuffle 0.133 vs 0.587; ECE 0.037; ~200 ms H100. Jev still leads general text 0.850 vs 0.786. Not Archer Watch. `notes.md` §46. - **Foodoo1/Qwen3-14B-RLCD-Decision-LoRA** — decision-token QLoRA on Qwen3-14B under parallel constrained decoding. Held-out 200-case / 4-field: fraud_risk 64→95%, overall 85.2→98.8% at ~234 ms. Synthetic; not a financial product. `notes.md` §46. @@ -117,6 +117,27 @@ clones — they own lint/eval and try-Jev-first habit. Entropy allocator decisions; a frontier write only for high-entropy synthesis (`judgment-class.md`; `research/notes.md` §38). +### Hourly ~14:52 Boise (extractive / local surface / speed layer) + +Patterns, not a catalog. `notes.md` §48. TypeSafe Jev is the exemplar in +the READMEs, not a monopoly. + +- **AppitStudio/testimonial-miner** — extractive selection + multi-question broadcast + offline `redecide`. Model never writes the quote. +- **choxos/jev-reviewer** — pointer-not-generator: line ids; verbatim copy with place; *not found* is an answer. +- **us/jev-local** — contract-compatible `POST /v1/systemone`. Default scorer is a **stub** until `JEVLOCAL_SCORER=hf`. +- **hitakshiA/solari-reflex** — observe → decide → verified act; no screenshots. Author table vs Codex on Solari ~3–7× wall. +- **ktaletsk/jevframe** — pandas/Polars `.jev` accessor; full `p__`; sibling of jevpandas. +- **de-niji/jev-hermes** — route ≠ memory: cheap intent gate skips memory tours. +- **ngallodev-software/agent-workflow-typesafe-ai** — advisory sidecar receipts; never changes host routing (Apache-2.0). +- **Joymfl/dag-jev** — structure induction over a bag (experiment; empty README; no metrics). +- **knowlet/jev-agentworld-web-simulator** — decision for control, generator for content; SQLite world. +- **ufx7/jev-testbench** — collab arms (`llm_autonomous` / `scripted_plus_jev` / `llm_plus_jev`); Wilson / McNemar. +- **alexykn/jevscan** — Tree-sitter ∩ typed questions. `tenbin` owns the lint skill. +- **phin-tech/pi-jev-approver** — Pi shell gate; fail-closed without a key. Light rh-guard-adjacent note. +- **Mattepiu/laya-onnx** — Laya ONNX port (~15 ms CPU). Do not copy the vs-Jev table. + +Spotcheck this pass (not a fold): SemIf **1551★** (+60 vs awesome claim 1491); jevlike **866★**. Awesomejev 488/21644 not re-derived (public snapshot still 410 / 10,093). Tracker lastModified **2026-09-18T20:12:57Z**; Laya listed; Blackwood not. Archer still Watch. + See `references/mixed-architecture.md` in the skill. Class-level family choice: `references/judgment-class.md`. Proof vs judgment (Alloy vs Apalache; DST trio Antithesis / Resonate HQ / PufferLib): diff --git a/research/archive/findings.md b/research/archive/findings.md index 021bac5..e8b38b4 100644 --- a/research/archive/findings.md +++ b/research/archive/findings.md @@ -788,6 +788,52 @@ Not a rewrite of §45. No species change. No invented metrics. Cross-repo addition: (be) bake-off fetch path is a Hub id; (bf) Choice `"other"` is a training confrontation, not only a request hatch. +## Batch #32 (2026-09-18 ~14:52 Boise) — extractive / local surface / speed layer + +Note: `research/notes.md` §48. Docs-only. Folded into PR #2. Not a +second PR. Archer still Watch. No invented metrics. Do not re-fold +blackwood-rlcd, open-jev-laya-bench, decision-token LoRA, jevgate, +wellposed, jev-reflex-autonomy-lab, Abide, kev, jevpandas, +bitrate-advisor. + +- **AppitStudio/testimonial-miner (Empirical as README fixture).** MIT. + Gmail → numbered sentences in code → one broadcast (Choice/Noul/Score + + per-sentence Nouls). Model never writes. `redecide` retunes + thresholds on the log. 8-request fixture: 5 candidates / 3 rejected + / 3 header skips. Offline tests use fakes. +- **choxos/jev-reviewer (Empirical as sample study).** MIT. Pointer at + line ids; code copies verbatim with place. *Not found* is an answer. + 712-line article: 1q 10 req / 1.2–2 s; 9q 17 / 2.3 s. Spot checks, + not a validation study. +- **us/jev-local (Contract as surface; Hypothesis on your labels).** + `POST /v1/systemone` drop-in. Default scorer is a **deterministic + stub** until `JEVLOCAL_SCORER=hf`. LICENSE absent this pass. Not a + Jev reproduction. +- **hitakshiA/solari-reflex (Empirical as named table).** MIT. + Observe → decide → verified act; no screenshots. Vs Codex on Solari: + 60.2 s vs 194.9 s; 66 s vs 460 s; 24.2 s vs 98.4 s (~3–7×). +- **ktaletsk/jevframe (Empirical as shape).** MIT, PyPI. pandas and + Polars `.jev`; full `p__`; no silent renormalize. Sibling of + jevpandas, not a re-fold. +- **MED:** jev-hermes (route ≠ memory); agent-workflow-typesafe-ai + (advisory sidecar, Apache-2.0); dag-jev (structure induction; + experiment; no metrics); jev-agentworld-web-simulator (decision + control / generator content); jev-testbench (collab arms; LICENSE + absent); jevscan (AST ∩ semantic; `tenbin` owns lint); pi-jev-approver + (fail-closed without key; light note); Mattepiu/laya-onnx (~15 ms + CPU; do not copy vs-Jev table). +- **Spotcheck:** SemIf 1551★; jevlike 866★; Awesomejev 488/21644 not + re-derived; tracker lastModified 2026-09-18T20:12:57Z; Laya listed; + Blackwood not. + +Cross-repo addition: (bg) extractive keep/drop + offline re-threshold +is judge-once/re-policy; (bh) pointer-not-generator is citation +integrity; (bi) `/v1/systemone` drop-in is a surface — stub ≠ scorer; +(bj) observe→decide→verified-act needs no screenshots; (bk) route ≠ +memory; (bl) advisory sidecar never changes host routing; (bm) collab +arms belong in the measurement curriculum. + + diff --git a/research/notes.md b/research/notes.md index d23f4ee..c75505c 100644 --- a/research/notes.md +++ b/research/notes.md @@ -2777,3 +2777,269 @@ guidance. Text/diff only today. Usage + measurement exemplar (replay harness, precision by phase). Not a multimodal substrate and not a reason to wait on Archer. + +## 48. Extractive selection, pointer-not-generator, local contract drop-in (2026-09-18 ~14:52 Boise) + +America/Boise ~14:52 = 20:52 UTC (run 205210). Docs-only fold into +open PR #2 (`cursor/augustus-store-envelope-00b4`). Not a competing +PR. Archer 27B drop still **WATCH**. Identity lock vs `typesafe-ai` / +`tenbin` / `decision-first` holds. No wrapper, no install.sh, no +copied ports / Docker / PyPI how-tos. No invented metrics. Do not +re-fold blackwood-rlcd, open-jev-laya-bench, decision-token LoRA, +jevgate, wellposed, jev-reflex-autonomy-lab, Abide, kev, jevpandas, +bitrate-advisor. + +Backend-agnostic: these are categorization / scoring *placements* +(extractive keep/drop, pointer evidence, contract-compatible local +scorer, structured observe→decide→act, dataframe columns, route≠memory, +advisory sidecar, structure induction, AST∩semantic). TypeSafe Jev is +the exemplar in the READMEs, not a monopoly. + +### HIGH + +1. **[`AppitStudio/testimonial-miner`](https://github.com/AppitStudio/testimonial-miner)** + (MIT, Python, created 2026-09-18T20:42Z, 0★ this pass). Gmail/IMAP + → numbered sentences in **code** → **one** typed request per email + (Choice kind + app, Nouls for user/praise/problem/English, Score + quotability, **one Noul per sentence**). **The model never writes + text.** The stored quote is the sender's sentences, selected by + those Nouls and joined in order. All thresholds live in + `Thresholds` and `redecide` reapplies them to logged answers + **without new model calls**. Header rules skip newsletters / + outbound / quoted replies before any call. Pattern: **extractive + selection + multi-question broadcast + offline re-thresholding.** + Cousin of exact-text keep/drop (`applied-mappings.md` §2), not a + testimonial product and not a Gmail how-to. + + **Evidence (README; not re-run; Empirical as that named receipt).** + Developed against `typesafe-sdk` 0.7.0 / `jev-1.13.0`. Their + *product* bars (not class constants): candidate if user ≥0.5, + praise ≥0.6, quality ≥1.5 on 0–3; borderline praise ≥0.35 / + quality ≥1.0; quote sentences ≥0.6. Live fixtures (8 requests, + ~22k input tokens): **5** candidates, **3** rejected, **3** header + skips. Offline tests use fakes and say nothing about model + accuracy. Author cost gloss: ~$0.042 / M input; typical email + 2–3k tokens; 10k judged emails ~$1. Do not copy IMAP setup. + Cards: `applied-mappings.md` §2; `mappings.md` §2; `methods-catalog.md`. + +2. **[`choxos/jev-reviewer`](https://github.com/choxos/jev-reviewer)** + (MIT, JavaScript, 1★, created 2026-09-18T14:47Z). Systematic-review + data extraction: PDF/Office/HTML/CSV stay in the **browser**; + segmenter assigns line ids (`A001`…); the decision model **points + at ids**; **code copies verbatim quotes with file and place**. + Nothing is paraphrased, so nothing can be invented. Two-pass: + screening Choice "which line answers q?" (+ none) over chunks, + then verifying Nouls "does this line itself answer q?". **Not + found is an answer.** Speculative fan-out: every question against + the same text. Pattern: **pointer-not-generator** for evidence + synthesis / citation integrity. Same species as keep/drop over + candidates code already holds; GLiNER locate is the cousin when + the answer *is* a span the encoder proposes. + + **Evidence (README; sample study, Sep 2026; not a validation + study).** 17-page article + 12-page analysis plan + CONSORT, + 712 lines: 1 question 10 req / 1.2–2 s / $0.0016; 9 questions + 17 / 2.3 s / $0.0052; 18-question template 27 / 4.6 s / $0.0101. + Chunks of 12k characters matched 7k with a third fewer requests. + Treat as spot checks. Scanned PDFs need OCR first (their limit). + Do not copy the relay. Cards: `applied-mappings.md` §2; + `methods-catalog.md` claim–evidence; `mental-models.md`. + +3. **[`us/jev-local`](https://github.com/us/jev-local)** + (LICENSE absent this pass; Python; 0★; created 2026-09-18T19:28Z). + Local `POST /v1/systemone` drop-in (Docker / `pip` / SDK + `base_url`). Open-weights *path*, no waitlist. **Default scorer is + a deterministic stub and carries no intelligence** until + `JEVLOCAL_SCORER=hf`. That flag is load-bearing: a green smoke + test on the stub is not a local Jev. Pattern: **contract-compatible + local scorer for offline/dev**, beside kev's Hub drop-in and + Laya's native head — not a fourth species. Their README: interface- + compatible baseline, **not a reproduction of Jev's undisclosed + model or training**. Softmax ignores level ordering; option order + can shift logits. Do not copy `install.sh`, ports, or model ids + into skill cards. + + **Evidence (README; Empirical as *their* named tables, Hypothesis + on *your* labels).** Official `typesafe-sdk==0.6.0` with only + `base_url` pointed at the server (4B backend) is the drop-in + proof. When `hf` is on they publish Wilson-CI / ECE / NLL tables + on templated sets (set1 / set2 / set3); do not re-promote those + rows as a ranking of TypeSafe Jev. Head-to-head vs published + jev-1.13.0 on 5 questions: 4/5 top-answer agree; payout + billing-vs-technical is an ambiguous prior, not a prompt bug. + Cards: `judgment-class.md`; `faq.md`; `validation.md`. + +4. **[`hitakshiA/solari-reflex`](https://github.com/hitakshiA/solari-reflex)** + (MIT, TypeScript, 0★, created 2026-09-18T15:24Z). Computer-use + **speed layer** on Solari (browser + Linux desktop): **one + structured observation → one typed decision → one verified + action**. **No screenshots in the loop.** Observation is + numbered controls / a11y tree / visible text (Calc: used visible + rows, never 2³¹ cells). Decision: which operation, which target + (speculative, same request), done/blocked. Write: a small model + only for TYPE as strict JSON. Act: guard check, refuse covered + controls, pipeline input; model output **never** becomes a + selector, coordinate, or script. Deny lists are **absent from the + question**, not merely disfavoured. Sibling of + `browser-use/jev-ultrafast` / jev-use / typesafe-computer-use. + Pattern: **perception → decision → act** with Harbor-style + measurement (score the task; independently checked by the app / + Stripe API / answer key). + + **Evidence (README + [solari-fast-showcase](https://github.com/hitakshiA/solari-fast-showcase); + vs Codex CLI GPT-6 Astra on the same Solari machines; not re-run).** + Stripe Checkout (qty 2, promo, card): **60.2 s**, $0.011 vs + **194.9 s**, 34 tool calls. Six different Stripe checkouts: + **66 s**, $0.064 vs **460 s**, 86 tool calls. 30 Calc expenses: + **24.2 s**, $0.0008 vs **98.4 s**, 77 tool calls. Ratios on that + table sit in ~3–7×. Their step table: decide ~400 ms / ~$0.0001; + TYPE write ~600 ms. **0.6** is *their* Advisor handoff, not a + universal threshold. Do not copy npm git-install. Cards: + `mixed-architecture.md`; `validation.md`; `applied-mappings.md` §2; + `agent-self-assessment.md`. + +5. **[`ktaletsk/jevframe`](https://github.com/ktaletsk/jevframe)** + (MIT, Python 3.10+, PyPI `jevframe`, pandas **and** Polars, + created 2026-09-18T17:06Z, 0★). `.jev` accessor: `noul` / `choice` + / `score` / `evaluate` with **full probability columns** (`p__…`), + preserved row order/indexes, one input row per request, questions + about a row share the request, default `max_concurrency` 16. + **No result is thresholded or silently renormalized.** `score` is + the expected zero-based level, not a probability. Sibling of + [`yalindogusahin/jevpandas`](https://github.com/yalindogusahin/jevpandas) + (`notes.md` §46): both are dataframe-native semantic columns; + jevpandas is the store-as-index cousin (`evaluate`/`filter`/ + `classify`/`score`); jevframe is the accessor + Polars + packed + struct layout. Neither is SQL `jev()`. Pattern: **dataframe-native + semantic columns** (class, not vendor). v0: no chat, no generated + records, no custom dtypes. Do not copy the client. Cards: + `mappings.md` §4; `mixed-architecture.md` gallery. + +### MED (pointers, not cards of their own) + +6. **[`de-niji/jev-hermes`](https://github.com/de-niji/jev-hermes)** + (MIT, Python, 0★). Intent **gate before a Hermes turn**: + `calendar` / `mail` / `status` → config + flat tools, **no memory + search spam that turn**; `complex` / people / prefs / "what did + we…" → memory + normal agent. Memory providers still **write** in + the background. Token savings come from skipping long tool/memory + *tours*, **not from turning memory off**. Pattern: **route ≠ + memory**; S1 gate preserves S2+memory for hard routes. OpenRouter + `POST /api/alpha/decisions` (same surface as jev-decision-layer). + Do not copy `docker cp`. Cards: `applied-mappings.md` §5; + `mixed-architecture.md` dual orchestration. + +7. **[`ngallodev-software/agent-workflow-typesafe-ai`](https://github.com/ngallodev-software/agent-workflow-typesafe-ai)** + (Apache-2.0, Python, 0★). Advisory-only host plugin. Projects + bounded redacted evidence into typed questions; normalizes answers + into versioned secret-free **semantic receipts**. **Never changes + host routing**, executor, model policy, lifecycle, evaluation, + review, or acceptance. Missing credentials / SDK / uncertain + answers / failures → distinct `no_action` outcomes. Pattern: + **soft sidecar receipts** (fail-open evidence). Complementary to + Abide (Abide is in-session Score on a diff; this is advisory + metadata the host may ignore). Do not copy the TOML. + +8. **[`Joymfl/dag-jev`](https://github.com/Joymfl/dag-jev)** + (Rust + petgraph; README empty this pass; 0★; created + 2026-09-18T20:42Z). GitHub description: DAG from unordered items + via Jev. Source: pairwise "does task i depend on j?" over a bag + of numbered steps (reads/writes in `input.txt`); answers intended + to build a `petgraph`. Experiment; graph wiring incomplete in + `main.rs` this pass; **no metrics**. Pattern: **structure + induction over bags** — code owns the DAG; the model only answers + pairwise (or Choice) dependency questions. Do not clone the + request builder. Cards: `mappings.md` §9 / §5. + +9. **[`knowlet/jev-agentworld-web-simulator`](https://github.com/knowlet/jev-agentworld-web-simulator)** + (MIT, TypeScript, 0★). Fictional web: Jev Choice for **search + intent** and (same request) **layout / palette**; an + OpenAI-compatible generator writes `SearchDocument` / + `PageDocument`; Zod validates; a **deterministic** compiler emits + json-render spec; SQLite is the world. Model cannot add + components or handlers. Mock mode does **not** silent-fallback + to live. Pattern: **decision for control, generator for content** + in a simulated world. Live smoke (when run) is an integration + test (budget **3 Jev + 3 generator**), **not** calibration or + world-consistency. Offline CI: 27 unit/contract + 2 Chromium + browse tests; [CI #2](https://github.com/knowlet/jev-agentworld-web-simulator/actions/runs/35391369860) + on `b1dd6d7` passed. Do not copy `.env`. Cards: + `mixed-architecture.md`. + +10. **[`ufx7/jev-testbench`](https://github.com/ufx7/jev-testbench)** + (LICENSE absent this pass; TypeScript; 0★). Two tools: black-box + **determinism / latency / context / concurrency** (`src/bench/`); + collaboration harness (`src/collab/`) with three arms — + `llm_autonomous` (unconstrained; illegal actions tracked), + `scripted_plus_jev` (code enumerates legal actions; Choice + picks; **no LLM**), `llm_plus_jev` (LLM proposes; Choice + arbitrates on the legal set; low p escalates then **stops** + rather than guessing). **Jev is not a peer arm.** Reports + Wilson intervals, exact McNemar per level (p < 0.05 **and** ≥5 + discordant pairs), ceiling-effect flags, cost/latency **per + model**. Pattern: bake this into the jevals/Harbor **measurement + curriculum**, not a second product. Grid-task demo uses a mock + LLM to prove the harness; swap before trusting a real model. + Cards: `validation.md`. + +11. **[`alexykn/jevscan`](https://github.com/alexykn/jevscan)** + (MIT, Python 3.12+, `0.2.0rc4`, 0★). Tree-sitter extracts + lexical units (Python/Rust/Perl/TS/JS); typed questions on + those targets; **does not execute or import the scanned code**. + Release candidate: **not a claim of calibrated semantic + accuracy.** Pattern: **structural AST + semantic judgment + compose** (sibling of riff / JevLint / Abide). `tenbin` still + owns the lint *skill*; this is a recipe of AST∩remainder, not + a second Augustus skill. Do not copy YAML. Cards: + `toolbox-mapping.md` spec/lint; `mixed-architecture.md`. + +12. **[`phin-tech/pi-jev-approver`](https://github.com/phin-tech/pi-jev-approver)** + (LICENSE present this pass; TypeScript; 0★). Pi **shell safety + gate**. Facts (git branch, path scope, registry-publish regex) + computed in code; typed Score `risk_level` + Nouls + Choice + `primary_concern` on the remainder. `commandRules` regex + **short-circuits** allow/deny — a `deny` rule is a hard block + no LLM or human prompt can overturn. **No key → fail closed.** + Optional LLM escalation can only reduce how often you are + asked, never replace the human as last resort. Author + side-by-side (their `jev-test` project): classification + **~500 ms** vs **~2.5–3.5 s** chat-model. Live-tested: `aws s3 + rm` 1.72/2, `ec2 terminate-instances` 1.80/2 in the ask band; + read verbs matching `action: allow` never called the model. + **Light note only** — rh-guard-adjacent (coding-agent tool + gate), different remainder from jevgate (unlisted verbs) and + Abide (soft project rules). Do not copy the Pi install. + +13. **[`Mattepiu/laya-onnx`](https://huggingface.co/Mattepiu/laya-onnx)** + (Apache-2.0 ONNX export of [`convaiinnovations/laya`](https://huggingface.co/convaiinnovations/laya); + Hub likes **1** this pass; updated 2026-09-18T10:45Z). Open + **replica deployment path**: non-autoregressive marker-token + head in onnxruntime. Card example: **~15 ms on CPU** for one + Noul. **Do not copy the inherited vs-Jev accuracy table** — + those rows remain vendor claims (`notes.md` §18, + `judgment-class.md`). Pattern: ONNX/runtime port of a trained + decision-only head, beside Laya native and kev Hub. Not Archer. + +### Spotcheck (not a fold) + +- **[`TheoLeeCJ/SemIf`](https://github.com/TheoLeeCJ/SemIf):** **1551★** + this pass (2026-09-18T20:51Z). Watch cited awesome claim **1491** + → **+60**. Independent; not affiliated with Jev/TypeSafe. +- **[`vinnylarouge/jevlike`](https://github.com/vinnylarouge/jevlike):** + **866★** this pass. +- **Awesomejev 488 / 21644:** watch cited "unchanged." This pass did + **not** independently re-derive 488/21644. + [awesomejev.com](https://awesomejev.com/) still showed **410 + entries / 10,093 stars** (refreshed 2026-09-17) in the public + snapshot fetched here. Do not invent a new count. +- **Tracker** [`multimodalart/jev-reproductions-tracker`](https://huggingface.co/spaces/multimodalart/jev-reproductions-tracker): + Hub `lastModified` **2026-09-18T20:12:57Z**. Space `models[]` + still lists **`convaiinnovations/laya`**. **`BlackwoodAI/blackwood-rlcd` + is not in that array.** Archer remains a *promised* trained-head + card. Omni decide that shipped (blackwood) is still off the + tracker this pass. + +Cards: `judgment-class.md`; `validation.md`; `applied-mappings.md` +§2 / §5; `mappings.md` §4 / §9; `mixed-architecture.md`; `faq.md`; +`mental-models.md`; `methods-catalog.md`; `toolbox-mapping.md`; +`agent-self-assessment.md`. No wrapper. diff --git a/research/refresh-log.md b/research/refresh-log.md index 9bd1615..c9fe7c0 100644 --- a/research/refresh-log.md +++ b/research/refresh-log.md @@ -443,6 +443,28 @@ - Cards: question-design, faq, judgment-class, validation, ecosystem. - notes.md §45 delta; sources.json; findings.md batch #31. No wrapper. +## 2026-09-18 20:52 UTC — extractive / local surface / speed layer (~14:52 Boise) + +- Folded into open PR #2 (`cursor/augustus-store-envelope-00b4`). + Docs-only. Not a competing PR. Archer 27B drop still **WATCH**. +- HIGH: testimonial-miner (extractive + `redecide`); jev-reviewer + (pointer-not-generator); jev-local (`/v1/systemone`; stub until hf); + solari-reflex (observe→act, no screenshots; 60.2/194.9, 66/460, + 24.2/98.4 s); jevframe (pandas/Polars accessor; sibling of jevpandas). +- MED: jev-hermes (route ≠ memory); agent-workflow-typesafe-ai + (advisory sidecar); dag-jev (structure induction, experiment); + jev-agentworld-web-simulator; jev-testbench (collab arms); + jevscan (AST ∩ semantic); pi-jev-approver (light); laya-onnx + (do not copy vs-Jev table). +- Spotcheck: SemIf 1551★; jevlike 866★; tracker 20:12:57Z lists Laya, + not Blackwood. Awesomejev 488/21644 not re-derived. +- Cards: SKILL.md, applied-mappings §2/§5, mappings §1/§4/§9/§18, + mixed-architecture, mental-models, validation, faq, judgment-class, + methods-catalog, toolbox, agent-self-assessment, question-design, + ecosystem, CHANGELOG, README. +- notes.md §48; sources.json; findings.md batch #32. No wrapper. + + diff --git a/research/sources.json b/research/sources.json index 93ad1db..4f6a5d5 100644 --- a/research/sources.json +++ b/research/sources.json @@ -1,6 +1,6 @@ { "refresh_cadence": "hourly", - "retrieved": "2026-09-18T20:43Z", + "retrieved": "2026-09-18T20:52Z", "sources": [ { "kind": "docs", @@ -1321,6 +1321,108 @@ "title": "jaredpalmer/kev-0.5b", "url": "https://huggingface.co/jaredpalmer/kev-0.5b", "note": "Hub weights for kev. kev.publish; --run accepts Hub ids; base downloads on first load. PEFT FEATURE_EXTRACTION. Apache-2.0 adapter. notes.md \u00a745 delta." + }, + { + "kind": "github", + "title": "AppitStudio/testimonial-miner", + "url": "https://github.com/AppitStudio/testimonial-miner", + "note": "MIT. Extractive quotes: numbered sentences in code; one broadcast; model never writes; redecide without new calls. notes.md \u00a748." + }, + { + "kind": "github", + "title": "choxos/jev-reviewer", + "url": "https://github.com/choxos/jev-reviewer", + "note": "MIT. Pointer-not-generator: line ids; verbatim copy with place; not found is an answer. notes.md \u00a748." + }, + { + "kind": "github", + "title": "us/jev-local", + "url": "https://github.com/us/jev-local", + "note": "POST /v1/systemone drop-in. Default scorer is a deterministic stub until JEVLOCAL_SCORER=hf. LICENSE absent this pass. notes.md \u00a748." + }, + { + "kind": "github", + "title": "hitakshiA/solari-reflex", + "url": "https://github.com/hitakshiA/solari-reflex", + "note": "MIT. Observe\u2192decide\u2192verified act; no screenshots. Vs Codex on Solari: 60.2/194.9, 66/460, 24.2/98.4 s. notes.md \u00a748." + }, + { + "kind": "github", + "title": "hitakshiA/solari-fast-showcase", + "url": "https://github.com/hitakshiA/solari-fast-showcase", + "note": "Companion table for solari-reflex vs Codex on Solari. notes.md \u00a748." + }, + { + "kind": "github", + "title": "ktaletsk/jevframe", + "url": "https://github.com/ktaletsk/jevframe", + "note": "MIT, PyPI. pandas and Polars .jev accessor; full p__ columns; no silent renormalize. Sibling of jevpandas. notes.md \u00a748." + }, + { + "kind": "github", + "title": "de-niji/jev-hermes", + "url": "https://github.com/de-niji/jev-hermes", + "note": "MIT. Route \u2260 memory: cheap intent gate skips memory tours on calendar/mail/status. notes.md \u00a748." + }, + { + "kind": "github", + "title": "ngallodev-software/agent-workflow-typesafe-ai", + "url": "https://github.com/ngallodev-software/agent-workflow-typesafe-ai", + "note": "Apache-2.0. Advisory sidecar receipts; never changes host routing. notes.md \u00a748." + }, + { + "kind": "github", + "title": "Joymfl/dag-jev", + "url": "https://github.com/Joymfl/dag-jev", + "note": "Structure induction over a bag; pairwise dependency; petgraph. README empty; experiment; no metrics. notes.md \u00a748." + }, + { + "kind": "github", + "title": "knowlet/jev-agentworld-web-simulator", + "url": "https://github.com/knowlet/jev-agentworld-web-simulator", + "note": "MIT. Decision for control, generator for content; SQLite world. CI #2 passed. notes.md \u00a748." + }, + { + "kind": "github", + "title": "ufx7/jev-testbench", + "url": "https://github.com/ufx7/jev-testbench", + "note": "Collab arms llm_autonomous / scripted_plus_jev / llm_plus_jev; Wilson/McNemar. LICENSE absent this pass. notes.md \u00a748." + }, + { + "kind": "github", + "title": "alexykn/jevscan", + "url": "https://github.com/alexykn/jevscan", + "note": "MIT. Tree-sitter \u2229 typed questions; 0.2.0rc4; does not execute scanned code. tenbin owns the lint skill. notes.md \u00a748." + }, + { + "kind": "github", + "title": "phin-tech/pi-jev-approver", + "url": "https://github.com/phin-tech/pi-jev-approver", + "note": "Pi shell safety gate; regex commandRules prove allow/deny; fail-closed without a key. Light rh-guard-adjacent note. notes.md \u00a748." + }, + { + "kind": "huggingface", + "title": "Mattepiu/laya-onnx", + "url": "https://huggingface.co/Mattepiu/laya-onnx", + "note": "Apache-2.0 ONNX export of convaiinnovations/laya. ~15 ms CPU for one Noul (their card). Do not copy inherited vs-Jev table. notes.md \u00a748." + }, + { + "kind": "huggingface", + "title": "multimodalart/jev-reproductions-tracker", + "url": "https://huggingface.co/spaces/multimodalart/jev-reproductions-tracker", + "note": "Hub lastModified 2026-09-18T20:12:57Z this pass. models[] still lists convaiinnovations/laya; BlackwoodAI/blackwood-rlcd absent. notes.md \u00a748." + }, + { + "kind": "github", + "title": "TheoLeeCJ/SemIf (star spotcheck)", + "url": "https://github.com/TheoLeeCJ/SemIf", + "note": "1551 stars this pass (2026-09-18T20:51Z); +60 vs awesome claim 1491. Independent. notes.md \u00a748." + }, + { + "kind": "github", + "title": "vinnylarouge/jevlike (star spotcheck)", + "url": "https://github.com/vinnylarouge/jevlike", + "note": "866 stars this pass. notes.md \u00a748." } ] } From 6a6d0208f9ffae747e8d9c12a04708658f6e41b6 Mon Sep 17 00:00:00 2001 From: Cursor Agent Date: Fri, 18 Sep 2026 22:03:54 +0000 Subject: [PATCH 07/43] Fold 15:52 Boise watch: boundary map, Harbor bake-off, dual-process Independent when-it-holds atlas (extractable-from-state), DMB frozen protocol vs constrained LLMs, jevals-data recompute-from-logs, and Kahneman S1/S2 cascade. Brief: ARC combinatorial negative, packed open-LLM System One, tiny SAN local drop-in, kev star delta. Archer still Watch. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com> --- .agents/skills/augustus/SKILL.md | 12 +- .../references/agent-self-assessment.md | 6 + .../augustus/references/applied-mappings.md | 11 +- .agents/skills/augustus/references/faq.md | 48 +++- .../augustus/references/judgment-class.md | 38 ++- .../skills/augustus/references/mappings.md | 23 ++ .../augustus/references/mental-models.md | 63 +++++ .../augustus/references/methods-catalog.md | 6 +- .../augustus/references/mixed-architecture.md | 28 +++ .../augustus/references/question-design.md | 2 + .../augustus/references/toolbox-mapping.md | 5 +- .../skills/augustus/references/validation.md | 44 +++- CHANGELOG.md | 19 ++ README.md | 20 +- docs/ecosystem.md | 19 +- research/archive/findings.md | 52 ++++ research/notes.md | 227 ++++++++++++++++++ research/refresh-log.md | 25 ++ research/sources.json | 56 ++++- 19 files changed, 677 insertions(+), 27 deletions(-) diff --git a/.agents/skills/augustus/SKILL.md b/.agents/skills/augustus/SKILL.md index 93d4c2d..cfdb2cd 100644 --- a/.agents/skills/augustus/SKILL.md +++ b/.agents/skills/augustus/SKILL.md @@ -1,6 +1,6 @@ --- name: augustus -description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill, open multimodal RLCD / blackwood), announced open decision-model (Watch — still not landed), constrained-AR (TypeAR, pcdServer, decision-token LoRA), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"wait for Archer vs open multimodal RLCD\", \"screenshot/DOM candidates → typed Choice\", \"eval path\", \"jevals\", \"Harbor taskset\", \"shared bake-off ECE/NLL/Brier\", \"LLM-as-judge is not the System One score\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge (allowlist *proves*; fail-open cannot block), \"missing other → confident wrong Choice\", \"lint the Jev request\", \"training confronts Choice other / none-of-the-above\", \"S1 reflex keeps control / optional S2 one-use advice\", \"soft AGENTS.md rules vs linter\" (Abide / jev-pref), \"edit-phase vs turn-phase observation window\", banded confidence fail-open preference lint, \"extractive selection / pointer-not-generator\", \"local /v1/systemone drop-in (stub until hf scorer)\", \"dataframe-native semantic columns\", \"route≠memory\", \"advisory sidecar receipts\", \"structure induction over bags\", \"AST ∩ semantic lint\", TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, or \"formally verify with Jev\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." +description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill, open multimodal RLCD / blackwood), announced open decision-model (Watch — still not landed), constrained-AR (TypeAR, pcdServer, decision-token LoRA), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"wait for Archer vs open multimodal RLCD\", \"screenshot/DOM candidates → typed Choice\", \"eval path\", \"jevals\", \"Harbor taskset\", \"shared bake-off ECE/NLL/Brier\", \"LLM-as-judge is not the System One score\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge (allowlist *proves*; fail-open cannot block), \"missing other → confident wrong Choice\", \"lint the Jev request\", \"training confronts Choice other / none-of-the-above\", \"S1 reflex keeps control / optional S2 one-use advice\", \"soft AGENTS.md rules vs linter\" (Abide / jev-pref), \"edit-phase vs turn-phase observation window\", banded confidence fail-open preference lint, \"extractive selection / pointer-not-generator\", \"local /v1/systemone drop-in (stub until hf scorer)\", \"dataframe-native semantic columns\", \"route≠memory\", \"advisory sidecar receipts\", \"structure induction over bags\", \"AST ∩ semantic lint\", \"extractable-from-state / retrieve first\", \"decision-model vs constrained-LLM bake-off\", \"dual-process S1 decide / S2 generate\", \"combinatorial grid ≠ extractive\", TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, or \"formally verify with Jev\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." license: MIT metadata: version: 0.3.0 @@ -54,7 +54,7 @@ classical method you already trust, substitute it, classify the win "paraphrase brittleness", "allowlist then judge", "TOCTOU-of-Noul", "Jev inside the database / sqlite-jev", "Jev picks bitrate / join order / the model", "wait for Archer", "lint the request / missing - other", "training confronts Choice other / none-of-the-above", "soft AGENTS.md rules vs the linter", "screenshot Choice / omni System One", "extractive quotes / pointer not generator", "local drop-in vs stub scorer", or "route vs memory": + other", "training confronts Choice other / none-of-the-above", "soft AGENTS.md rules vs the linter", "screenshot Choice / omni System One", "extractive quotes / pointer not generator", "local drop-in vs stub scorer", "route vs memory", "when does it hold / extractable from state", "decision model vs constrained LLM", "dual-process S1/S2", or "combinatorial grid vs extractive": read `references/faq.md`, then `references/mental-models.md`, then `references/mixed-architecture.md`, then @@ -113,12 +113,12 @@ classical method you already trust, substitute it, classify the win | Familiar method | Judgment shape | Detail | |---|---|---| -| Mental models across domains (not SWE-only) | EU, abstention, VOI, MCDA, SDT, search/control, Leveson, NATM/Norman/snap-fit | `references/mental-models.md` | +| Mental models across domains (not SWE-only) | EU, abstention, VOI, MCDA, SDT, search/control, Leveson, NATM/Norman/snap-fit; **extractable-from-state boundary map** (self-contained vs needs outside knowledge) | `references/mental-models.md` | | Judgment-model class (Jev is exemplar, not monopoly) | Species: decide / locate (GLiNER) / categorize (GLiClass) / rank / perceive; open heads include encoder DeBERTa, LoRA distill, openjev-lm, kev | `references/judgment-class.md` | -| Open weights vs constrained decoding vs proprietary API | Three open paths: encoder open-jev / AR constrained decode (TypeAR + pcdServer; decision-token LoRA) / trained decision-only (Laya + ONNX port, Nimble, kev, **blackwood-rlcd** multimodal now, Archer Watch still Watch). Contract-compatible local `/v1/systemone` (jev-local) is a drop-in *surface*; default scorer is a stub until `hf`. Softmax over allowed tokens ≠ Noul | `references/judgment-class.md` (when-to-use table) | +| Open weights vs constrained decoding vs proprietary API | Three open paths: encoder open-jev / AR constrained decode (TypeAR + pcdServer; decision-token LoRA; packed one-forward logprob on open LLMs) / trained decision-only (Laya + ONNX port, Nimble, kev, **blackwood-rlcd** multimodal now, Archer Watch still Watch). Local `/v1/systemone` surfaces: jev-local (stub until `hf`), kev (trained pointer), von (tiny SAN — not a replica). Softmax over allowed tokens ≠ Noul | `references/judgment-class.md` (when-to-use table) | | Entropy as allocator (low / medium / high) | Typed low+medium decisions → System One marginals; high-entropy synthesis → frontier decoder. Product rhetoric, not a meter. **Hypothesis** | `references/judgment-class.md` | | Formal / semi-formal (proof vs judgment) | Sensor vs constraint vs searchlight; Alloy vs Apalache; DST trio; TOCTOU-of-Noul, AI×FM | `references/formal-methods.md` (one-screen: `references/formal-semi-formal.md`) | -| Mixed architecture (judgment model + LLM) | Provider judges, LLM writes, code owns control; not a stack replacement. Advisory sidecar never changes host routing | `references/mixed-architecture.md` | +| Mixed architecture (judgment model + LLM) | Provider judges, LLM writes, code owns control; not a stack replacement. Advisory sidecar never changes host routing. Dual-process: S1 decides, S2 generates (routing accuracy unmeasured). Internals ≠ FSM; placement is a **component node**. Browser-use strength = DOM-as-text + speculative fan-out, not vision | `references/mixed-architecture.md` | | Context sieve | Relevance Noul per block; always-keep set in code; stub + recall key | `references/applied-mappings.md#1-context-sieve` | | Exact-text keep / drop | Choice include/exclude/mixed over candidates code already holds. Extractive quotes / pointer-not-generator (model never writes the excerpt) | `references/applied-mappings.md#2-exact-text-keep--drop` | | Environment / harness triage | Scan every step for env failure; LLM autopsy only on flags | `references/applied-mappings.md#3-environment--harness-triage` | @@ -151,7 +151,7 @@ classical method you already trust, substitute it, classify the win | Agent self-supervision / on-track detection | Pre-gate → output judge → done-check → supervisor nouls. S1 reflex keeps control; optional S2 is one-use advice | `references/agent-self-assessment.md` | | Optimizer/program frameworks (Ax, DSPy) | Typed fields → one provider request; judge metrics; threshold discipline. Ax and DSPy climb LM-program knobs only | `references/optimizer-integration.md` | | Perception → decision pipeline / measure / hill-climb | Stages with a versioned state contract; frozen taskset; DSPy/Ax only on the LM-program slice. Specialist composition stays **Hypothesis**; open multimodal decide (blackwood-rlcd) is a named receipt. Structured observe→decide→verified-act (no screenshots) is a computer-use speed-layer receipt. Same section as the row below | `references/validation.md#eval--hill-climb` | -| Eval & hill-climb | Decision-stage jevals hygiene; Harbor taskset × harness × runtime; one score-composition table. Shared bake-off exemplar: open-jev-laya-bench (ECE/NLL/Brier; LLM-as-judge is not the score). Collab-arm curriculum: llm_autonomous vs scripted_plus_jev vs llm_plus_jev (Wilson / McNemar) | `references/validation.md#eval--hill-climb` | +| Eval & hill-climb | Decision-stage jevals hygiene; Harbor taskset × harness × runtime; one score-composition table. Shared bake-off exemplar: open-jev-laya-bench (ECE/NLL/Brier; LLM-as-judge is not the score). Harbor-style frozen protocol vs constrained LLMs: DMB (accuracy/calibration/latency/cost; raw logs). Feedstock: jevals-data CC-BY-4.0 boards + JSONL (recompute-from-logs). Collab-arm curriculum: llm_autonomous vs scripted_plus_jev vs llm_plus_jev (Wilson / McNemar). Negative: combinatorial grid assembly ≠ extractive (ARC-AGI Direct Jev 4/400) | `references/validation.md#eval--hill-climb` | | (meta) Finding new mappings & applications | Toolbox sweep: judgment-shaped component of a known method, substituted + falsified | `references/toolbox-mapping.md` | | Named methods / operators / theorems | Substitution tiers: operand-judgments, preconditioned theorems, non-substitutable | `references/methods-catalog.md` | | (meta) Where a judgment model sits relative to any construct | 11 positions + logical-operator rules + position×construct traversal as the application generator | `references/composition-algebra.md` | diff --git a/.agents/skills/augustus/references/agent-self-assessment.md b/.agents/skills/augustus/references/agent-self-assessment.md index 8b8cbb8..3dc94f6 100644 --- a/.agents/skills/augustus/references/agent-self-assessment.md +++ b/.agents/skills/augustus/references/agent-self-assessment.md @@ -40,6 +40,12 @@ relying: foreman, pi-jev, pi-warden, winnow, fast-jev-compaction, jev-judgment. **no screenshots**; model output never becomes a selector. Harbor-style task score (Stripe API / answer key). Author table vs Codex on Solari: 60.2 s vs 194.9 s; 66 s vs 460 s; 24.2 s vs 98.4 s (`notes.md` §48). + Productized Kahneman cascade for *any* cheap-decide / expensive-write + loop (business/life, not only SWE): + [dual-process-ai](https://github.com/taro1985/dual-process-ai) — + `conf ≥ τ` S1 decides else S2 generates; routing fails open; safety + fails closed; **routing accuracy not measured**; keyword fallback is + not S1 (`notes.md` §49). Tune τ on your escalation log. 6. **Context economy**: the context-sieve card (`references/applied-mappings.md#1-context-sieve`). Judge every large tool result with one relevance Noul before it enters context. Hide diff --git a/.agents/skills/augustus/references/applied-mappings.md b/.agents/skills/augustus/references/applied-mappings.md index c1d89a4..183decd 100644 --- a/.agents/skills/augustus/references/applied-mappings.md +++ b/.agents/skills/augustus/references/applied-mappings.md @@ -96,6 +96,13 @@ found* is an answer. Computer-use cousin: [solari-reflex](https://github.com/hitakshiA/solari-reflex) — structured observation → typed decision → verified act; **no screenshots**; model output never becomes a selector (`notes.md` §48). +**DOM-as-text + fan-out (Empirical as atlas browser-use *shape*):** a +screenshot task translated into a structured DOM snapshot as `state`, +then speculative questions over numbered candidates — not vision +(`notes.md` §49; `mental-models.md` §boundary). +**Axis check:** if the kept byte / cited fact / click target is not +already in the candidates you numbered, this card does not apply — +retrieve or parse first; do not ask recall. **Counterexample**: "write the patch that matches this sentence" — that is generation. **Test**: every kept byte occurs in the input; mixed never auto-included; snapshot stale → abort, don't guess. @@ -248,4 +255,6 @@ the OCR bill it authorises**; 9 false-skips vs 28 for rules-only OCR, so a watermark talks a scan into "has text." **Test**: planted scans are sent; planted born-digital pages are not billed; page order preserved. Re-measure on *your* documents. Same sandwich as jevgate (Proven / Refused / -Unknown). +Unknown). Same VOI as retrieve-then-state: if the answer is not in the +cheap text layer, **pay for the passage / OCR**, then judge +(`mental-models.md` §boundary; atlas history suite). diff --git a/.agents/skills/augustus/references/faq.md b/.agents/skills/augustus/references/faq.md index 86d4b6c..8348704 100644 --- a/.agents/skills/augustus/references/faq.md +++ b/.agents/skills/augustus/references/faq.md @@ -376,6 +376,44 @@ table). Specialist composition (SAM / OCR → text → Jev) remains valid. Do not wait, and do not treat screenshot-vs-Jev-text as the same input. `judgment-class.md`; `notes.md` §46. +## When does a decision model hold? + +When the answer is **extractable from the state you feed it** +(classification, citation/paraphrase/reversed-meaning with claim + +quote both given, sarcasm whose trigger is in the text, DOM-as-text +fan-out). It **fails, often confidently**, on recall without a +supporting passage. Atlas history suite (`notes.md` §49): Case A +wrong @ 0.90 with no context; Case B near-flat 0.07 (correct by luck); +Case C — same item as B — right @ 0.97 once the passage is in +`state`. Retrieve first. Overlapping categories can still be +**dangerous-high** (DAIR Emotion 48% acc / mean conf 0.819). Not a +leaderboard; one axis. `mental-models.md` §boundary. + +## Are the internals a state machine? + +No. Distributed LM understanding, not hand-written transitions +(paraphrase vs reversed-meaning contrast). **Placement** is a +component node in *your* code — that instinct is right; "it *is* a +state machine" is not. `mappings.md` §3; `mixed-architecture.md`. + +## Decision model or a constrained LLM? + +Neither class won on quality in DMB v2 (`notes.md` §49). Decision +models win the *axes you actually buy*: p50 ~264–276 ms, ~$0.07/1k +on banking, schema-valid, flat to 255 options — and **fail at 256+**. +Constrained LLMs handle 512 and (most of them) admit ignorance on +no-good items; they are slower and cost more. Spam: two OpenAI models +sat *below* the majority baseline. Re-run on *your* labels. Do not +merge Banking77 87% / 76.3% / 79.67% across protocols. +`judgment-class.md`; `validation.md`. + +## Can many cell-wise Choices solve a combinatorial grid? + +Not in the Direct Jev ARC-AGI-1 experiment: **4/400 (1%)**, ~$2.32. +Dimensions ~90%; complete grids rarely. Combinatorial assembly ≠ +extractive keep/drop. Search or a program stays in code. +`mappings.md` §9; `notes.md` §49. + ## Is a local `/v1/systemone` the same as Jev? No — not until you know **which scorer** is behind the socket. @@ -384,8 +422,14 @@ drop-in (`base_url`). The **default scorer is a deterministic stub** and carries no intelligence. `JEVLOCAL_SCORER=hf` turns on a frozen-model logprob scorer. Their README: an interface-compatible baseline, not a reproduction of Jev's undisclosed model. kev is the other local -drop-in (trained pointer head, public gold). A green smoke test on the -stub is not a bake-off. `judgment-class.md`; `notes.md` §48. +drop-in (trained pointer head, public gold). [von](https://github.com/wfzyx/von) +is a **tiny SAN** (14 MB needle; authored144 52.6%) at the extreme of +the speed/econ class — not the stub, not kev, **not a calibrated Jev +replica**. Do not copy its vs-Jev table. +[open-alternative-jev](https://github.com/ikermoel/open-alternative-jev) +packs one-forward logprobs on an open LLM you already have (RACE-H +92.9% @ 4.55 q/s); **not a Jev reproduction**. A green smoke test on +the stub is not a bake-off. `judgment-class.md`; `notes.md` §48, §49. ## Should the model write the quote / the citation / the click? diff --git a/.agents/skills/augustus/references/judgment-class.md b/.agents/skills/augustus/references/judgment-class.md index 7ed0c4b..12f9b08 100644 --- a/.agents/skills/augustus/references/judgment-class.md +++ b/.agents/skills/augustus/references/judgment-class.md @@ -447,12 +447,45 @@ smoke test on the stub is not a local decision model (`notes.md` §48). [`Mattepiu/laya-onnx`](https://huggingface.co/Mattepiu/laya-onnx) (~15 ms CPU for one Noul, their card). Do not copy the inherited vs-Jev accuracy table (`notes.md` §18). +**Packed one-forward on an open LLM (constrained-AR / logprob path, +not a Jev reproduction):** +[`ikermoel/open-alternative-jev`](https://github.com/ikermoel/open-alternative-jev) +(Apache-2.0). RACE-H packed **92.9% @ 4.55 q/s** on Qwen3.6-27B +8-bit; interference 6–9%; temperature scaling on *your* labels. +Space demo. Economics of packing a shared state, not trained +decision-only (`notes.md` §49). +**Tiny SAN local surface (extreme speed/econ class, not a replica):** +[`wfzyx/von`](https://github.com/wfzyx/von) — 14 MB Needle; `POST +/v1/systemone`; sub-15 ms CPU *claim* / ~38 ms embed in their table; +authored144 needle **52.6%**. Distinguish from jev-local's **stub** +and kev's trained pointer. **Do not copy the vs-Jev ranking table.** + +**When to use a decision model vs a constrained LLM (Harbor-style +bake-off, not a quality ranking).** +[`nibzard/decision-model-benchmark`](https://github.com/nibzard/decision-model-benchmark) +v2 (`notes.md` §49): no class wins on *accuracy*. Pick from axes +you actually need: + +| Need | Lean decision-model (Jev-class) | Lean constrained LLM | +|---|---|---| +| Latency / $ at schema-valid enums | p50 ~264–276 ms; S1 ~$0.07/1k; 0% malformed this run | Thinking-mode seconds and $0.19–$2.48/1k; some malformed | +| Choice sets ≤255 | Flat latency 2→255; **fails at 256+** (`400 Too many choices`) | Handles 512 | +| Honesty / abstain on no-good items | S5 admits 49.7%; ECE 0.246 — measure, do not assume | Most 97.3–100% (gpt-5.4-mini 64.7%) | +| Position stability | S4 flip 13% this run | Up to 37% | +| Mid-pack banking accuracy | 76.3% this protocol (atlas/jev-benchmarks 87%; jevals.com 79.67% — **do not merge**) | gpt-oss-120b 81.3%; glm-5.3 80.4% | + +Quality is not the reason to skip System One. Speed, cost, schema, +and the Choice cap are. Always re-run on *your* labels. Feedstock +for recomputing named boards: [`Jevals/jevals-data`](https://github.com/Jevals/jevals-data) +(CC-BY-4.0; `validation.md`). Reject: TypeAR or pcdServer scores as fail-closed P(permit); a LoRA student's agreement with Jev as independent gold; kev's ID ECE as a license to skip a held-out test on *your* workflow; shipping on "smarter than Jev"; a green `/v1/systemone` smoke test on jev-local's stub as a bake-off; -copying laya-onnx vs-Jev rows as independent gold; +copying laya-onnx or von vs-Jev rows as independent gold; treating von's +14 MB needle (52.6% authored144) or open-alternative-jev as a Jev +reproduction; thresholding [`jp-sns-jev7-estimator`](https://huggingface.co/kokuren/jp-sns-jev7-estimator) teacher scores as P(toxic) — the card says they are **not** calibrated, and `threat` F1@0.5 is 0.0000 on their table (`notes.md` §33). Domain-local @@ -462,7 +495,8 @@ universal ranking. Diffusion beating a decision head is Hypothesis. Detail: `research/notes.md` §33 (surfaces), §34 (marginals), §35 (Nimble), §36 (diffusion), §38 (entropy allocator, Hypothesis), §42 (pcdServer serving, meta-VOI, games), §45 (kev), §46 (blackwood, -laya-bench, decision-token LoRA), §48 (jev-local stub, laya-onnx). Before +laya-bench, decision-token LoRA), §48 (jev-local stub, laya-onnx), +§49 (boundary map; DMB vs constrained LLMs; von; open-alternative-jev). Before adopting a surface, the bake-off is a jevals-shaped suite and, for a product loop, a Harbor taskset (`validation.md`, Eval & hill-climb). The stage pipeline into that decision is the same file diff --git a/.agents/skills/augustus/references/mappings.md b/.agents/skills/augustus/references/mappings.md index 135e354..118f001 100644 --- a/.agents/skills/augustus/references/mappings.md +++ b/.agents/skills/augustus/references/mappings.md @@ -91,6 +91,12 @@ on-question?"; "call this lead / nurture / drop" — same act/abstain/ gather table, costs written in hours or dollars, threshold per *action*. VOI: pay for the full PDF or the customer call only if expected decision change beats the cost (`mental-models.md` §VOI, §decision). +**Dual-process cascade (Empirical as a productized metaphor, routing +accuracy unmeasured):** +[dual-process-ai](https://github.com/taro1985/dual-process-ai) — +`confidence ≥ τ` → S1 decides; else escalate to S2 (generate). Routing +fails open; safety fails closed. Keyword fallback without a key is not +equivalent S1. Tune τ on *your* escalation log (`notes.md` §49). **Counterexample**: a flat Choice over three fine categories may still name a harmless best pick — low confidence need not veto a low-stakes preference. **Test**: cost/coverage curve on held-out slices; score the @@ -115,6 +121,9 @@ NOT `P(A∧B)`; operation+target head pairs can be invalid (browser-use asks both heads per request but executes only the matching target after validation); relational judgments ("does passage support claim?") must stay one question, not two split classifications; exact computation stays in code. +**Internals are not a state machine; placement is a component node** +(atlas: paraphrase vs reversed-meaning are LM understanding; "a node in +your state machine" is the architectural instinct — `notes.md` §49). **Example**: game director — Jev judges whether player dialogue is conciliatory or threatening; code enforces inventory, prerequisites, @@ -276,6 +285,12 @@ prose → frontier. Same 149 business rows: Jev 79.9% vs Haiku 4.5 83.2%; Jev 1.6× faster, not 20–200×; Jev confidence monotonic, Haiku inverts in 0.80–0.95. If you do not *branch on confidence*, use whatever you already have (`notes.md` §42). +**Retrieve-then-state (Empirical as an axis proof, not a knowledge +estimate):** if the answer is not in `state`, **buy the passage first**, +then ask. Atlas history suite: wrong @ 0.90 without context → right @ +0.97 with the passage (`notes.md` §49; `mental-models.md` §boundary). +That observation is VOI with a named receipt. Do not rely on bare +recall. **Beyond SWE (Hypothesis until you log act/outcome pairs):** full PDF vs abstract; customer call vs CRM fields that already fail a hard rule (credit limit is exact); blood test vs @@ -400,6 +415,14 @@ owns topology; the model does not emit edges. Experiment; empty README; no metrics this pass (`notes.md` §48). Same hole as taxonomy beam (§5): judgment is a pairwise (or Choice) classifier step, not the scheduler. +**Combinatorial grid assembly ≠ extractive keep/drop (Empirical as a +negative):** +[`simonmesmith/jev-arc-agi-v1-experiment`](https://github.com/simonmesmith/jev-arc-agi-v1-experiment) +— Direct Jev cell-wise Choice on ARC-AGI-1: **4/400 (1%)**. Dimensions +~90%; complete grids rarely. Many small extractive decisions do not +add up to a consistent transformation. Search / a program / a +simulator stay in code (`notes.md` §49). + **Query planner as the envelope (author-reported, 2026-09-18):** [@mmalisper](https://x.com/mmalisper/status/2101001041903009987) on the Join Order Benchmark. Jev picking join order was **2× slower**. diff --git a/.agents/skills/augustus/references/mental-models.md b/.agents/skills/augustus/references/mental-models.md index d91648d..82f7ea9 100644 --- a/.agents/skills/augustus/references/mental-models.md +++ b/.agents/skills/augustus/references/mental-models.md @@ -52,6 +52,7 @@ until you label *your* cases. | Crossover metaphors | NATM, snap-fit, Norman, Kent, Shirky | This file §crossover | | Formal / semi-formal | Proof vs DST vs judgment | `formal-methods.md`, `formal-semi-formal.md` | | Class / family / objective | Decide vs locate vs categorize vs rank vs perceive | `judgment-class.md` species map | +| Boundary map / extractable-from-state | Self-contained in fed state vs needs outside knowledge | This file §boundary; atlas receipts `notes.md` §49 | Pick the pillar from the hole, then the family, then the vendor. @@ -105,6 +106,59 @@ SWE examples are **Empirical** (git-jev-stage, OpenSmoke). The others are curve on held-out *your* cases; score the fallback (escalation is not automatically correct). `mappings.md` §2. +## Boundary map: extractable from state (placement judgment) + +Primary mental model this hour +([jev-capability-atlas](https://github.com/Zaious/jev-capability-atlas); +independent unofficial receipts, not a leaderboard; `notes.md` §49). +Before picking a family or a vendor, place the *task*: + +> Is the correct answer fully recoverable from the `state` you hand +> the model, or does it require outside knowledge that is not in +> `state`? + +| Self-contained (in the state) | Not self-contained (needs outside knowledge) | +|---|---| +| Classify / route / gate over text you already hold | Trivia / recall with no supporting passage | +| Citation / paraphrase / reversed-meaning given claim + quote | Score that needs comparison against a whole field | +| Sarcasm / entailment whose trigger is in the given text | Overlapping blurred categories (dangerous-high ECE) | +| DOM snapshot / numbered candidates → Choice | Combinatorial assembly (grid cells that must agree) | + +**Empirical as that named axis, not as a knowledge-breadth estimate.** +History suite (N=3, single annotator, Chinese history; Case A ground +truth itself contested): common-knowledge item **wrong @ 0.90** with +no context (Yongzheng; Kangxi by popular convention); obscure item +near-flat **0.07** without a passage (correct by luck; informal rerun +wrong @ 0.08) → **right @ 0.97** with the passage in `state` +(Xianfeng, 0.98 mass). Teaching: **bare memory is unreliable; reading +comprehension over supplied text is reliable.** Retrieve first; put +the passage in `state`. Do not treat the atlas 30-second slogan as +the table. + +**Placement, not internals.** The model is not a state machine under +the hood (distributed LM understanding: `paraphrase_support` and +`reversed_meaning_high_overlap` both judged correctly). It *is* +correctly used as a **component node** in *your* program — code owns +transitions (`mappings.md` §3). Confidence is a **statistic from the +distribution** (RLCD trains the distribution; Choice `confidence` is +how peaked it is), not a second trained correctness score. +Calibration is **population-level** and can fail **dangerous-high**: +DAIR Emotion via jev-benchmarks — 48% acc, mean conf **0.819**, 16% +of items p(correct)=0. Overlapping categories, overconfident. Plot +reliability on *your* labels before you threshold. + +**Browser-use is this axis, not vision.** Strength = DOM-as-text + +speculative fan-out over candidates code already numbered — a visual +task translated into extractive text. Not screenshots. Same +component-node placement as lizard-agent / solari-reflex +(`applied-mappings.md` §2; `mixed-architecture.md`). + +**Does not:** merge Banking77 87% (atlas/jev-benchmarks) with DMB +76.3% or jevals.com 79.67% into one ranking — protocol / n / split +(`validation.md`, `notes.md` §49). Combinatorial grids are not +extractive keep/drop (ARC-AGI Direct Jev 4/400). FAQ: when-it-holds; +state-machine; retrieve-first. + ## Calibration and cost-sensitive thresholds A number you can threshold is a *decision* number only after you check @@ -132,6 +186,11 @@ two existing options in every block, and reversing order moved a probability across a ~0.9 threshold. Property-test both (`validation.md`). Correctness is not that confidence field: report both, on held-out cases (`validation.md`, Eval & hill-climb). Stimulus design, not a proof. +Atlas receipt of the same arithmetic: DAIR Emotion mean conf 0.819 at +48% acc (`notes.md` §49) — population calibration can fail +dangerous-high on overlapping labels. DMB S5: jev admits-ignorance +49.7% vs most constrained LLMs 97.3–100% (ECE 0.246). Do not skip +the honesty suite because in-distribution ECE looked fine. For a calibrated binary p and unequal error costs, the Bayes threshold is `t = C_FP / (C_FP + C_FN)` when you act vs not @@ -408,6 +467,10 @@ Use these as *existence proofs of a position*. Write your own card. | Shell / tool allowlist | unlisted remainder after a **proof** | five Nouls on unknown verbs | Proven/Refused in code; cannot block (**Empirical**: jevgate) | | SWE | residual AGENTS.md / CLAUDE.md rules | one Score per named instruction-file rule | linter owns hard rules; bands + fail-open (**Empirical**: Abide replay, `notes.md` §47) | | Screenshot candidates → act | lettered elements code already marked | Choice over those letters | Click in code (**Empirical** as blackwood-rlcd *shape*; CC BY-NC) | +| Browser / DOM candidates → act | numbered elements from a **text** snapshot | Choice / Nouls over those ids (**Empirical** as atlas browser-use *shape*: DOM-as-text + fan-out, not vision) | Click in code; no screenshots | +| Knowledge / recall | fact that is not in the document | **Do not ask.** Retrieve the passage first; then a self-contained Choice (**Empirical**: history suite A wrong@0.90 → C right@0.97) | Index, citation, the passage in `state` | +| Dual-process cascade | cheap classify / route vs write | S1 typed decision + τ; S2 generates only on low conf (**Empirical as a productized metaphor**; routing accuracy **unmeasured** — dual-process-ai) | Safety still fail-closed in code | +| Combinatorial puzzle | whole grid / program that must be consistent | **Rejected as extractive.** Cell-wise Choice assembly is not keep/drop (ARC-AGI Direct Jev 4/400) | Search, a program, a simulator | | Moderation | hold before publish | hazard Nouls (**Empirical** as family) | block/review policy | | Phishing / fraud screen | hold vs deliver | SDT criterion on a Noul | blocklist, SPF/DKIM exact (**Hypothesis**) | | Personal ops | cook done / not | "looks done" Noul | thermometer probe | diff --git a/.agents/skills/augustus/references/methods-catalog.md b/.agents/skills/augustus/references/methods-catalog.md index 3528a73..6e970de 100644 --- a/.agents/skills/augustus/references/methods-catalog.md +++ b/.agents/skills/augustus/references/methods-catalog.md @@ -33,7 +33,8 @@ judgment component is new). | Neyman–Pearson / selective classification | Decision threshold under error costs | One threshold per action, set on split A, reported on split B; abstention path | Loss model, ROC analysis | **Contract + empirical** (confidence-routing; evaluator script) | | Value of information (EVPI / EVSI) | Whether another observation is worth its cost | Gather as an enumerated act; pay iff expected decision-loss drop > cost | Cost of the observation; the loss table | **Hypothesis** as a numeric calculator; **Contract** as the placement (`mappings.md` §6) | | Signal detection (Green & Swets) | Evidence variable + criterion | Noul as noisy evidence; t from costs and base rate; ROC/PR on your labels | Operating point, base-rate tracking | **Hypothesis** for non-SWE plots; **Empirical** as moderation *shape* (`mappings.md` §7) | -| Reliability calibration (Platt/temperature) | Raw scores → calibrated probabilities | Noul is natively calibrated **in-distribution only**; verify with reliability bins on your own population; re-fit a correction out-of-distribution | Calibration fitting, binning | **Empirical recipe** (ECE 0.0313 in-distribution; 32% OOD collapse — Archer Hume) | +| Reliability calibration (Platt/temperature) | Raw scores → calibrated probabilities | Noul is natively calibrated **in-distribution only**; verify with reliability bins on your own population; re-fit a correction out-of-distribution | Calibration fitting, binning | **Empirical recipe** (ECE 0.0313 in-distribution; 32% OOD collapse — Archer Hume). Atlas: DAIR Emotion dangerous-high (48% / 0.819); DMB S5 ECE 0.246 (`notes.md` §49) | +| Frozen-protocol bake-off vs constrained LLMs | Same items, accuracy + ECE + latency + cost + honesty | Decision-model as one contender class, not the score | Protocol, raw logs, baselines | **Empirical as Harbor/jevals practice** (DMB v2; jevals-data CC-BY-4.0 recompute-from-logs; `notes.md` §49) | | Survey scoring / psychometrics | Rubric level judgment with defined anchors | Score with concrete level descriptions; probabilities read beside every score | Weighted aggregation, reliability analysis | **Contract** (score docs: split composite judgments) | ## Search, planning & operations research @@ -66,8 +67,9 @@ judgment component is new). | Method | Judgment-shaped component | Jev substitution | Stays in code | Status | |---|---|---|---|---| -| Claim–evidence entailment (NLI) | supports / contradicts / not-established per claim–source pair | One Choice per pair + review flag; judge against the cited source text only | Quote extraction, citation graph, audit log. When the answer *is* a span you already hold, **point** at line ids and copy verbatim — the model never writes the excerpt (jev-reviewer) | **Empirical recipe** (citation_check cookbook; jev-reviewer sample study, `notes.md` §48) | +| Claim–evidence entailment (NLI) | supports / contradicts / not-established per claim–source pair | One Choice per pair + review flag; judge against the cited source text only | Quote extraction, citation graph, audit log. When the answer *is* a span you already hold, **point** at line ids and copy verbatim — the model never writes the excerpt (jev-reviewer) | **Empirical recipe** (citation_check cookbook; jev-reviewer sample study, `notes.md` §48). Atlas: paraphrase_support and reversed_meaning_high_overlap correctly judged when both texts are in `state` (`notes.md` §49) | | Extractive selection + offline re-threshold | Keep/drop over sentences or ids code already numbered | Per-item Noul/Choice + one broadcast; join in order; `redecide` on the log with no new calls | Numbering, header skip, thresholds, publish permission | **Empirical recipe** (testimonial-miner 8-request fixture; jev-reviewer; `notes.md` §48). Cousin of applied-mappings §2 | +| Combinatorial grid / program synthesis | Consistent whole-object from many cells | **Rejected as extractive.** Cell-wise Choice does not assemble ARC grids (4/400 Direct Jev) | Search, a program, a simulator | **Empirical as a negative** (`notes.md` §49) | | Spec vs artifact conformance (model checking *mindset*) | Property holds / violated / unverifiable for a named requirement | One Noul/Score per requirement, batched; violated → named rule back into context (pi-warden / Abide shape). This is **not** TLC/Apalache/GNATprove | Requirement enumeration, enforcement, logging; the **linter** if the rule is lintable; the real checker if you have one | **Empirical recipe** (pi-warden: 6→0 rule breaks, 150 paired runs; jev-pref: YOU define the rule; Abide: productized compile/calibrate/tune/replay, `notes.md` §47). Ownership split: `formal-methods.md` | | AST ∩ semantic lint | Semantic remainder after a parser already extracted units | Typed questions on Tree-sitter targets; do not execute scanned code | Parser, selection, fail-on; `tenbin` owns the lint *skill* | **Empirical as a shape** (jevscan 0.2.0rc4; not a calibration claim; `notes.md` §48) | | Alloy finder vs Apalache / TLC | Which bound, which counterexample, is the property tautological? | Triage instances/CEs; never "this spec looks right" | Analyzer / SMT / explicit-state engine | **Hypothesis** as product; **Contract** as ownership (`formal-methods.md` §2) | diff --git a/.agents/skills/augustus/references/mixed-architecture.md b/.agents/skills/augustus/references/mixed-architecture.md index 34f4764..0ed0f25 100644 --- a/.agents/skills/augustus/references/mixed-architecture.md +++ b/.agents/skills/augustus/references/mixed-architecture.md @@ -78,6 +78,30 @@ from the page; Jev picks among visible elements; code enforces prices/dates; answers are *located*, never composed. The day the task needs a paragraph, an LLM re-enters — that is mixed architecture, not a different religion. +**Component node, not internals-as-FSM.** Place the judgment model as a +node in *your* state machine; do not describe its internals as one. +Atlas teaching (`notes.md` §49): paraphrase_support and +reversed_meaning_high_overlap are distributed LM understanding, not +keyword rules. Code still owns transitions (`mappings.md` §3). + +**Browser-use strength is DOM-as-text + speculative fan-out, not +vision.** Translate a screenshot task into a structured DOM snapshot +as `state`; fan out over candidate elements in one call. Text-only +models then sit in their strong zone (extractable from fed state). +Same placement as lizard-agent / solari-reflex; not blackwood +screenshot-in. + +**Dual-process cascade (Kahneman productized; routing accuracy +unmeasured).** +[`taro1985/dual-process-ai`](https://github.com/taro1985/dual-process-ai) +(MIT): S1 typed decision + confidence; S2 generates only when +`confidence < τ`. Routing **fails open**; safety **fails closed**. +Degraded keyword mode without a key is **not** equivalent S1. Crossover +for business/life, not only SWE: cheap classify/route/gate on S1; +write/reason on S2. Same split as jev-reflex-autonomy-lab (S1 keeps +control) and jev-hermes (route ≠ memory). Do not copy hooks. Tune τ +on *your* escalation log. + ## It's just classification Canonical answer: `references/faq.md`. Short form: classification is not @@ -278,6 +302,10 @@ calendar/mail/status off the memory tour; complex keeps Honcho (`notes.md` §48). **Advisory sidecar:** [agent-workflow-typesafe-ai](https://github.com/ngallodev-software/agent-workflow-typesafe-ai) emits `no_action` receipts and **never** changes host routing. +**Productized Kahneman cascade:** +[dual-process-ai](https://github.com/taro1985/dual-process-ai) — S1 +decides, S2 writes; routing accuracy **not measured**; keyword +fallback is not S1 (`notes.md` §49). **Effect-oriented loop (same author, later post).** Topology B inside an effect system diff --git a/.agents/skills/augustus/references/question-design.md b/.agents/skills/augustus/references/question-design.md index 98077e6..a90b8fa 100644 --- a/.agents/skills/augustus/references/question-design.md +++ b/.agents/skills/augustus/references/question-design.md @@ -54,6 +54,8 @@ request, and treat a stale pin as a prior, never a setting. | Symptom | Likely cause | Fix | | --- | --- | --- | | Wrong answers, high confidence | Instruction read literally | State exact condition; put boundary cases in criteria | +| Wrong answers, high confidence, **nothing in state that could answer** | Bare recall / missing evidence | **Retrieve first**; put the passage in `state`. Atlas history: wrong@0.90 without context → right@0.97 with passage. Confidence gating on recall is not enough (Case A was 0.90 *and wrong*). `notes.md` §49 | +| Wrong answers, **dangerous-high** ECE on overlapping labels | Population calibration failed (blurred categories) | Do not threshold. DAIR Emotion: 48% acc / mean conf 0.819 / 16% p(correct)=0. Plot reliability on *your* labels | | Wrong answers, **confidence ~1.00**, no `other` | Forced pick: the offered set does not cover the input; the model *must* choose | Add `other` / none-of-the-above. **Confidence gating cannot catch this** ([wellposed](https://github.com/suraj-phanindra/wellposed) live probe: unsubscribe email → `"support issue"` at 1.00 without `other`, `"other"` at 0.93 with it). Overlapping options collapse confidence (loud). `notes.md` §46. `tenbin` owns the lint skill | | Residual `"other"` always picked (or never) | Training saw none-of-the-above only as the true label — a wording shortcut | Confront the hatch as a *wrong* alternative too; vary wording; eval present-vs-removed ([kev](https://github.com/jaredpalmer/kev) `none_of_the_above`; `notes.md` §45 delta). wellposed still owns request-shape lint | | Question names a state path that does not exist | Dead reference; the API still answers | Lint the request (walk JSON). Structural, not semantic. wellposed recipe; do not copy the CLI | diff --git a/.agents/skills/augustus/references/toolbox-mapping.md b/.agents/skills/augustus/references/toolbox-mapping.md index 2154f9e..042bbc1 100644 --- a/.agents/skills/augustus/references/toolbox-mapping.md +++ b/.agents/skills/augustus/references/toolbox-mapping.md @@ -70,9 +70,12 @@ component; keep the rest of the method in code. | Search: value function | Score rubric as leaf value V(s) — only where a simulator validates outcomes; speculative depth hard-capped at 2 | **Empirical recipe** (jev-mcts fidelity split) | | Measurement theory: probe vs estimate | Only post-execution probes concede milestones; model estimates never do — "estimation wearing a measurement costume" is the rejection template | **Empirical recipe** (jev-mcts, pi-warden done-check) | | Experimental design: perturbation | Behavioral tests as the stats layer: candidate removal, option-order shuffle, letter-shuffle on screenshot Choice, distractor injection, boundary cases | **Contract-level** (validation.md); letter-shuffle receipt: blackwood-rlcd 0.133 vs Jev 1.13 0.587 on 300 web steps (`notes.md` §46) | +| Experimental design: frozen protocol bake-off | Decision-model vs constrained LLMs vs deterministic baselines; accuracy + ECE + latency + cost + honesty; raw logs; recompute | **Empirical as Harbor/jevals practice** (DMB v2; jevals-data CC-BY-4.0 boards + JSONL; `notes.md` §49). Do not merge Banking77 across protocols | +| Experimental design: extractable-from-state axis | Same question with vs without a supporting passage; citation paraphrase vs reversed-meaning | **Empirical as a boundary map** (jev-capability-atlas history suite N=3; `notes.md` §49). Qualitative, not a knowledge-breadth estimate | +| Experimental design: combinatorial negative | Cell-wise Choice assembly of a grid vs extractive keep/drop | **Empirical as a negative** (ARC-AGI Direct Jev 4/400; `notes.md` §49) | | Experimental design: collab arms | `llm_autonomous` vs `scripted_plus_jev` vs `llm_plus_jev`; Wilson + McNemar; the decision model is **not** a peer arm | **Empirical as a harness shape** (jev-testbench; bake into jevals/Harbor, `notes.md` §48) | | Discrete math: width vs depth | Fan out in width (parallel ≈ free), pay depth linearly; two-stage only when next options depend on an earlier answer | **Empirical recipe** (fan-out: 12.2× cheaper, 10× faster) | -| Psychology: Kahneman | System 2 generates/proposes (LLM), System 1 discriminates (Jev); never the reverse. S1 keeps control; optional S2 is one-use advice and does not fly | **Empirical recipe** (mcts-agent role split; 2026-09-18 mixed-architecture discourse; jev-reflex-autonomy-lab as a control-loop shape, `notes.md` §46). **Route ≠ memory:** a cheap intent gate skips memory/tool *tours* on easy routes; memory still writes; complex still searches (jev-hermes, `notes.md` §48) | +| Psychology: Kahneman | System 2 generates/proposes (LLM), System 1 discriminates (Jev); never the reverse. S1 keeps control; optional S2 is one-use advice and does not fly. Productized cascade: `conf ≥ τ` → S1 decides else S2 writes; routing fails open / safety fails closed; **routing accuracy unmeasured**; keyword fallback ≠ S1 | **Empirical recipe** (mcts-agent role split; 2026-09-18 mixed-architecture discourse; jev-reflex-autonomy-lab as a control-loop shape, `notes.md` §46; [dual-process-ai](https://github.com/taro1985/dual-process-ai) as a business/life cascade, `notes.md` §49). **Route ≠ memory:** a cheap intent gate skips memory/tool *tours* on easy routes; memory still writes; complex still searches (jev-hermes, `notes.md` §48) | | IR / cascades | Cheap relevance / irrelevance before an expensive ranker or generator | **Empirical recipe** (RAG cookbook; jevprune; git-jev-stage; LlamaIndex Jev rerank) | | IR / extractive keep-drop | Number candidates in code; model selects; copy verbatim; `redecide` thresholds on the log with no new calls | **Empirical recipe** (testimonial-miner; jev-reviewer pointer-not-generator; `notes.md` §48) | | Spec / lint | Project-defined semantic rules as predicates over a diff; linter owns hard rules. AST remainder: Tree-sitter units, then typed questions; do not execute scanned code | **Empirical recipe** (jev-pref contract; Abide productized path — replay 93 sessions, edit precision ~26% / turn ~73% before tune, `notes.md` §47; JevLint file-level Noul; pi-warden; snifftest unsure-band; jevscan AST∩semantic, `tenbin` owns the lint skill, `notes.md` §48) | diff --git a/.agents/skills/augustus/references/validation.md b/.agents/skills/augustus/references/validation.md index 8fec564..6a706f6 100644 --- a/.agents/skills/augustus/references/validation.md +++ b/.agents/skills/augustus/references/validation.md @@ -309,7 +309,10 @@ Rules: LLM-as-judge is not the primary score for a calibrated System One. Shared bake-off exemplar: [open-jev-laya-bench](https://huggingface.co/datasets/pngwn/open-jev-laya-bench) - (ECE/NLL/Brier). + (ECE/NLL/Brier). Harbor-style class bake-off: + [DMB](https://github.com/nibzard/decision-model-benchmark) (v2). + Public log feedstock: + [jevals-data](https://github.com/Jevals/jevals-data) (CC-BY-4.0). ### Composition @@ -345,14 +348,49 @@ vs 98.4 s (`notes.md` §48). **Collab-arm curriculum:** McNemar; Jev is not a peer arm. Bake into jevals/Harbor hygiene, do not copy the harness. +**Harbor-style frozen protocol vs constrained LLMs (Empirical as that +named receipt, not a ranking).** +[`nibzard/decision-model-benchmark`](https://github.com/nibzard/decision-model-benchmark) +(DMB): jev vs 8 constrained LLMs vs keyword/majority/random; five +suites; **$28.34**; raw logs. **`results/v2/v2.md` is the report of +record.** Protocol frozen before the run; negative results ship; +unknown usage is never a measured zero; later runs replace cells +whole. jev S1 banking **76.3%**, S2 spam **93.0%**, S3 **100%*** at +valid coverage **72.7%** (225 failed = 256+ Choice cap), S4 flip +**13%**, S5 admits-ignorance **49.7%** / ECE **0.246**; p50 +**264–276 ms**; S1 cost/1k **$0.07**. No class wins on quality. +Do not copy `uv`. Do not merge this Banking77 with atlas 87% or +jevals.com 79.67% (`notes.md` §49). + +**Feedstock / recompute-from-logs (not a third ranking).** +[`Jevals/jevals-data`](https://github.com/Jevals/jevals-data) +(CC-BY-4.0): release boards + per-decision JSONL + suite files for +[jevals.com](https://jevals.com). 2026-09-18 board, suite 0.1.0, 8 +systems (banking77 / helpsteer2 / pubmedqa). Formulas: +https://jevals.com/methodology/. Jev on *this* board (n=300×5): +banking77 acc **0.7967**, ECE **0.0981**, p50 **467 ms**, cost/1k +**$0.043**. Cite the release; recompute from logs; do not dump the +board as a ranking. + +**Negative: combinatorial assembly ≠ extractive keep/drop.** +[`simonmesmith/jev-arc-agi-v1-experiment`](https://github.com/simonmesmith/jev-arc-agi-v1-experiment) +— Direct Jev on ARC-AGI-1 public eval **4/400 (1%)**, 1.125% +task-weighted, ~$2.32, 10 min. Cell-wise Choice; dimensions ~90%; +rarely a complete grid. A frozen Harbor-shaped protocol that +falsifies "many small decisions add up to a puzzle." + ### Bake-off mandate Before adopting proprietary Jev vs Laya vs TypeAR vs Nimble vs kev vs -blackwood-rlcd vs Archer vs openjev-lm, run a jevals-shaped labeled suite (or an equivalent +blackwood-rlcd vs Archer vs openjev-lm vs a constrained LLM vs von vs +open-alternative-jev, run a jevals-shaped labeled suite (or an equivalent with this hygiene) and, for a product loop, a Harbor taskset. A design card with no eval path is incomplete. A green smoke test on [jev-local](https://github.com/us/jev-local)'s **default stub** is not -that bake-off (`notes.md` §48). +that bake-off (`notes.md` §48). von's 14 MB needle at 52.6% authored144 +is not that bake-off either (`notes.md` §49). DMB is the frozen-protocol +exemplar for decision-model vs constrained-LLM vs baselines; jevals-data +is the public log feedstock. Do not promote a vendor table into a ranking. **Shared bake-off exemplar (Empirical as that named receipt, not a ranking).** [`pngwn/open-jev-laya-bench`](https://huggingface.co/datasets/pngwn/open-jev-laya-bench) diff --git a/CHANGELOG.md b/CHANGELOG.md index 9ebdcd8..fcc0eb3 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -180,6 +180,25 @@ Pegged against [`typesafe-ai/skills` v0.5.7](https://github.com/typesafe-ai/skil ([laya-onnx](https://huggingface.co/Mattepiu/laya-onnx); do not copy vs-Jev table). Spotcheck: SemIf 1551★; jevlike 866★; tracker 20:12:57Z still lists Laya, not Blackwood. No wrapper. +- Hourly ~15:52 Boise fold (`research/notes.md` §49): Archer still + Watch. X discourse blocked. Boundary map / extractable-from-state + ([jev-capability-atlas](https://github.com/Zaious/jev-capability-atlas); + history suite A wrong@0.90 / B 0.07 / C right@0.97; component node; + dangerous-high ECE; DOM-as-text + fan-out). Harbor-style bake-off vs + constrained LLMs + ([DMB](https://github.com/nibzard/decision-model-benchmark) v2: jev + banking 76.3% / spam 93.0% / 256+ cap; p50 264–276 ms; $0.07/1k; no + class wins on quality). Feedstock + ([jevals-data](https://github.com/Jevals/jevals-data) CC-BY-4.0; + recompute-from-logs; 2026-09-18 board). Dual-process S1 decide / S2 + generate ([dual-process-ai](https://github.com/taro1985/dual-process-ai); + routing accuracy unmeasured). Combinatorial ≠ extractive (ARC-AGI + Direct Jev 4/400). Packed one-forward open LLM + ([open-alternative-jev](https://github.com/ikermoel/open-alternative-jev) + RACE-H 92.9% @ 4.55 q/s; not a Jev reproduction). Tiny SAN local + surface ([von](https://github.com/wfzyx/von) 14 MB; not a replica). + kev light delta **100★**. Do not merge Banking77 87% / 76.3% / + 79.67%. No wrapper. - Effect-oriented loops (`notes.md` §28, `mappings.md` §19): Ward's ZIO client keeps Jev as the outer Choice and the handler as the effect. Not Effect.ts. GLiNER author: GLiNER2 "like jev" is GLiGuard diff --git a/README.md b/README.md index 2cf4a0a..d65adee 100644 --- a/README.md +++ b/README.md @@ -27,14 +27,16 @@ never launder a Noul as a proof. - `.agents/skills/augustus/SKILL.md` — working protocol + decision-design card - `.agents/skills/augustus/references/mental-models.md` — cross-domain frames (EU, abstention, VOI, MCDA, SDT, search/control, Leveson, - NATM/snap-fit/Norman); not SWE-only + NATM/snap-fit/Norman); not SWE-only. Extractable-from-state boundary + map (self-contained vs needs outside knowledge) - `.agents/skills/augustus/references/judgment-class.md` — the class (Jev exemplar, not monopoly): open heads (Laya, kev, encoder DeBERTa, LoRA distill), constrained-AR (TypeAR, pcdServer), announced decision-model (Watch), open multimodal RLCD (blackwood-rlcd; not Archer), Laya ONNX port, - contract-compatible local `/v1/systemone` (stub until hf scorer), + contract-compatible local `/v1/systemone` (stub until hf scorer; also kev pointer / von tiny SAN — not replicas), GLiNER/GLiClass species (locate vs categorize vs local multi-head), - listwise vs decision objectives, vision scoring, when-to-use axes, + listwise vs decision objectives, vision scoring, when-to-use axes + (including decision-model vs constrained LLM), agent-architecture portents - `.agents/skills/augustus/references/formal-methods.md` — judgment vs proof ownership; Alloy Analyzer vs Apalache (finder ≠ BMC ≠ @@ -47,12 +49,13 @@ never launder a Noul as a proof. alias of the FM pillar - `.agents/skills/augustus/references/mixed-architecture.md` — default placement: judgment-class model + LLM + code; preference lint; provider - (Jev default / other family with self-eval) + (Jev default / other family with self-eval); dual-process S1 decide / S2 + generate; component node; DOM-as-text + fan-out - `.agents/skills/augustus/references/applied-mappings.md` — context sieve, exact-text keep/drop (extractive / pointer-not-generator), env triage, moderation/ranking, skill routing (route ≠ memory) - `.agents/skills/augustus/references/faq.md` — "just classification", stack replacement, Jev vs open head vs encoder vs LoRA vs constrained AR vs kev vs blackwood, - wait-for-Archer, missing-other confident-wrong, soft project rules vs linter (Abide), extractive/pointer-not-generator, local drop-in vs stub scorer, route ≠ memory, GLiNER vs GLiClass vs CLIP, LLM-as-judge, in-engine vs CLI store, + wait-for-Archer, missing-other confident-wrong, soft project rules vs linter (Abide), extractive/pointer-not-generator, local drop-in vs stub scorer, route ≠ memory, when-it-holds / extractable-from-state, decision-model vs constrained LLM, dual-process S1/S2, combinatorial grid ≠ extractive, GLiNER vs GLiClass vs CLIP, LLM-as-judge, in-engine vs CLI store, hard envelope (bitrate / planner), not-another-how-to - `.agents/skills/augustus/references/mappings.md` — classical-method mappings with boundaries, counterexamples, acceptance tests (including @@ -60,9 +63,12 @@ never launder a Noul as a proof. - `.agents/skills/augustus/references/validation.md` — design gate, eval recipes, Jev-for-skills (routing, self-monitoring, testing, modularity, frontmatter), and Eval & hill-climb (jevals hygiene + Harbor taskset; - open-jev-laya-bench as ECE/NLL/Brier bake-off exemplar; Abide replay as + open-jev-laya-bench as ECE/NLL/Brier bake-off exemplar; DMB as + Harbor-style frozen protocol vs constrained LLMs; jevals-data as + CC-BY-4.0 recompute-from-logs feedstock; Abide replay as Harbor-adjacent soft-rule measurement; solari-reflex Harbor-style - computer-use; jev-testbench collab arms) + computer-use; jev-testbench collab arms; ARC-AGI Direct Jev as + combinatorial-≠-extractive negative) - `.agents/skills/augustus/references/boundary-audit.md` — existing-system insertion: fit test, opportunity map, smallest boundary, red flags - `.agents/skills/augustus/scripts/evaluate_decisions.py` — offline evaluator diff --git a/docs/ecosystem.md b/docs/ecosystem.md index 626a72e..93bcc53 100644 --- a/docs/ecosystem.md +++ b/docs/ecosystem.md @@ -46,7 +46,7 @@ Per-keystroke launchers (104ms median, sequence-tagged staleness), firehose mode - **arnabgho/rlcd-lite** — GRPO + Brier proper-scoring-rule reward → calibrated decisions; binary reward doesn't calibrate. - **stephanj/parallelConstraintDecoding** — whole JSON schema of booleans/enums in two forward passes (prefill → parallel masked fields). - **Foadsf/jev-for-engineers**, **AbdelStark/jev-benchmarks**, **BrendanH18/jev-lab** — measurement discipline and cost/latency visibility. -- **dayhaysoos/jevals** — local MIT workbench: labeled cases (Noul / Choice / Score), compare runs, WebMCP + agent skill. Empirical acceptance-test surface for Hypothesis mapping cards; complements `evaluate_decisions.py`. Not affiliated with TypeSafe. Pointer: `research/notes.md` §24. Hygiene and the Harbor substrate: `validation.md` Eval & hill-climb (`notes.md` §40). Shared bake-off exemplar this hour: [`pngwn/open-jev-laya-bench`](https://huggingface.co/datasets/pngwn/open-jev-laya-bench) (ECE/NLL/Brier; LLM-as-judge is not the score; `notes.md` §46). +- **dayhaysoos/jevals** — local MIT workbench: labeled cases (Noul / Choice / Score), compare runs, WebMCP + agent skill. Empirical acceptance-test surface for Hypothesis mapping cards; complements `evaluate_decisions.py`. Not affiliated with TypeSafe. Pointer: `research/notes.md` §24. Hygiene and the Harbor substrate: `validation.md` Eval & hill-climb (`notes.md` §40). Shared bake-off exemplar: [`pngwn/open-jev-laya-bench`](https://huggingface.co/datasets/pngwn/open-jev-laya-bench) (ECE/NLL/Brier; LLM-as-judge is not the score; `notes.md` §46). Harbor-style frozen protocol vs constrained LLMs: [`nibzard/decision-model-benchmark`](https://github.com/nibzard/decision-model-benchmark) (DMB v2; jev banking 76.3% / spam 93.0% / 256+ cap; p50 264–276 ms; $0.07/1k; `notes.md` §49). Feedstock: [`Jevals/jevals-data`](https://github.com/Jevals/jevals-data) (CC-BY-4.0 boards + JSONL; recompute-from-logs; 2026-09-18 board). Boundary map (not a leaderboard): [`Zaious/jev-capability-atlas`](https://github.com/Zaious/jev-capability-atlas). Combinatorial negative: [`simonmesmith/jev-arc-agi-v1-experiment`](https://github.com/simonmesmith/jev-arc-agi-v1-experiment) (Direct Jev 4/400). - **jeiel85/jevscope** — local-first visual debugger + JSONL regression for Choice/Score/Noul; compare two definitions; policy buckets are JevScope-derived. Sits next to jevals. Pointer: `research/notes.md` §25. ### Local / open heads & GLi\* species @@ -54,7 +54,9 @@ Per-keystroke launchers (104ms median, sequence-tagged staleness), firehose mode - **GLiGuard** (fastino-ai) — 0.3B GLiNER2 encoder, checkpoint `fastino/gliguard-LLMGuardrails-300M`. One bidirectional pass over a safety schema. Same interface shape as batched questions; different objective. Not a Jev weight clone. `judgment-class.md`; `notes.md` §30. - **DECRUX9812/openjev-lm** — Qwen2.5-0.5B+LoRA distilled from hosted Jev answers; 65/70 = 92.9% on 70 hand-labelled rows (one annotator, one domain, one seed) overnight on 6 vCPU, $0/call. Its 98.1% on fresh rows is teacher *agreement*, not gold. Receipts pattern: `notes.md` §25, §44. - **convaiinnovations/laya** — open Choice/Score/Noul head, text-only, 512 tok. Companion packaging this hour: [`laya-typed-decisions`](https://huggingface.co/convaiinnovations/laya-typed-decisions) (421.3M, acc 0.766 / Brier 0.066 unverified). Shared bake-off: [`pngwn/open-jev-laya-bench`](https://huggingface.co/datasets/pngwn/open-jev-laya-bench) (26+9 tasks, 11959 items; ECE/NLL/Brier; not TypeSafe Jev vs Laya). ONNX replica: [`Mattepiu/laya-onnx`](https://huggingface.co/Mattepiu/laya-onnx) (~15 ms CPU; do not copy vs-Jev table). `notes.md` §18, §42, §46, §48. -- **jaredpalmer/kev** — Qwen2.5-0.5B LoRA + pointer readout; Apache-2.0; Hub [`jaredpalmer/kev-0.5b`](https://huggingface.co/jaredpalmer/kev-0.5b) plus GitHub release tarball. Runnable Archer reconstruction (`POST /v1/systemone`). Public gold, not a Jev teacher. Isolation exact; ID ECE 0.065 (0.031 after T); acc 0.799 / 1,350. NOTA training must confront `"other"` as a wrong alternative (`notes.md` §45 delta). Not a knowledge/frontier substitute. +- **jaredpalmer/kev** — Qwen2.5-0.5B LoRA + pointer readout; Apache-2.0; Hub [`jaredpalmer/kev-0.5b`](https://huggingface.co/jaredpalmer/kev-0.5b) plus GitHub release tarball. Runnable Archer reconstruction (`POST /v1/systemone`). Public gold, not a Jev teacher. Isolation exact; ID ECE 0.065 (0.031 after T); acc 0.799 / 1,350. NOTA training must confront `"other"` as a wrong alternative (`notes.md` §45 delta). **100★** this pass (light activity delta; `notes.md` §49). Not a knowledge/frontier substitute. +- **ikermoel/open-alternative-jev** — packed one-forward logprob System One on open LLMs (HF + vLLM). Apache-2.0. **Not a Jev reproduction.** RACE-H 92.9% @ 4.55 q/s on Qwen3.6-27B 8-bit; interference 6–9%. Space demo. `notes.md` §49. +- **wfzyx/von** — 14 MB Needle SAN; local `POST /v1/systemone`; sub-15 ms CPU claim / ~38 ms embed. authored144 needle 52.6% — **not a calibrated Jev replica**. Distinguish from jev-local stub and kev pointer. Do not copy vs-Jev table. `notes.md` §49. - **BlackwoodAI/blackwood-rlcd** — open multimodal RLCD (image-text-to-text), Jev-compatible shim, CC BY-NC 4.0. Screenshot + marked candidates → Choice. Card: web acc 0.907 vs Jev 1.13 text-only 0.480; letter-shuffle 0.133 vs 0.587; ECE 0.037; ~200 ms H100. Jev still leads general text 0.850 vs 0.786. Not Archer Watch. `notes.md` §46. - **Foodoo1/Qwen3-14B-RLCD-Decision-LoRA** — decision-token QLoRA on Qwen3-14B under parallel constrained decoding. Held-out 200-case / 4-field: fraud_risk 64→95%, overall 85.2→98.8% at ~234 ms. Synthetic; not a financial product. `notes.md` §46. - **zmtomorrow/TypeAR** — constrained autoregressive decoding surface: typed fields on a pretrained open model, no retraining. Not a proper-scoring decision head. `research/notes.md` §32. @@ -138,6 +140,19 @@ the READMEs, not a monopoly. Spotcheck this pass (not a fold): SemIf **1551★** (+60 vs awesome claim 1491); jevlike **866★**. Awesomejev 488/21644 not re-derived (public snapshot still 410 / 10,093). Tracker lastModified **2026-09-18T20:12:57Z**; Laya listed; Blackwood not. Archer still Watch. +### Hourly ~15:52 Boise (boundary map / Harbor bake-off / dual-process) + +Patterns, not a catalog. `notes.md` §49. TypeSafe Jev is the exemplar, not a monopoly. Archer still Watch. X discourse blocked this hour. + +- **Zaious/jev-capability-atlas** — when-it-holds map with API receipts. Axis: extractable from fed state vs needs outside knowledge. History suite table (A wrong@0.90 / B near-flat 0.07 / C right@0.97). Component node ≠ internals-as-FSM. Dangerous-high ECE (DAIR Emotion). Browser-use = DOM-as-text + fan-out, not vision. +- **nibzard/decision-model-benchmark (DMB)** — frozen protocol: jev vs 8 constrained LLMs vs baselines. v2 report of record. jev banking 76.3% / spam 93.0% / 256+ cap; p50 264–276 ms; $0.07/1k. No class wins on quality. +- **Jevals/jevals-data** — CC-BY-4.0 boards + JSONL. Recompute-from-logs. 2026-09-18 board (do not merge Banking77 with DMB/atlas). +- **taro1985/dual-process-ai** — Kahneman S1 decide / S2 generate. Routing accuracy unmeasured. Keyword fallback ≠ S1. +- **simonmesmith/jev-arc-agi-v1-experiment** — Direct Jev 4/400 (1%). Combinatorial ≠ extractive. +- **ikermoel/open-alternative-jev** — packed one-forward; RACE-H 92.9% @ 4.55 q/s. Not a Jev reproduction. +- **wfzyx/von** — 14 MB SAN local drop-in. Distinguishes from jev-local stub / kev pointer. Do not copy vs-Jev table. +- **jaredpalmer/kev** — light delta: **100★** this pass. No species rewrite. + See `references/mixed-architecture.md` in the skill. Class-level family choice: `references/judgment-class.md`. Proof vs judgment (Alloy vs Apalache; DST trio Antithesis / Resonate HQ / PufferLib): diff --git a/research/archive/findings.md b/research/archive/findings.md index e8b38b4..931ca31 100644 --- a/research/archive/findings.md +++ b/research/archive/findings.md @@ -833,6 +833,58 @@ integrity; (bi) `/v1/systemone` drop-in is a surface — stub ≠ scorer; memory; (bl) advisory sidecar never changes host routing; (bm) collab arms belong in the measurement curriculum. +## Batch #33 (2026-09-18 ~15:52 Boise) — boundary map / Harbor bake-off / dual-process + +Note: `research/notes.md` §49. Docs-only. Folded into PR #2. Not a +second PR. Archer still Watch. X discourse blocked this hour. No +invented metrics. Do not re-fold §48 items. + +- **Zaious/jev-capability-atlas (Empirical as an axis, not a + knowledge-breadth estimate).** README MIT / GitHub SPDX + NOASSERTION. Unofficial. Extractable-from-state vs needs-outside + knowledge. History suite table: A wrong@0.90 (Yongzheng; Kangxi by + popular convention; GT contested); B near-flat 0.07 (luck); + C right@0.97 with passage. N=3, single annotator. Internals ≠ FSM; + placement is a component node. Confidence is a distribution + statistic (RLCD). Dangerous-high ECE: DAIR Emotion 48% / 0.819 / + 16% p(correct)=0. Browser-use = DOM-as-text + speculative fan-out, + not vision. +- **nibzard/decision-model-benchmark (Empirical as Harbor/jevals + practice).** LICENSE absent this pass. Frozen protocol; v2 report + of record; $28.34. jev banking 76.3%, spam 93.0%, S3 100%* at + 72.7% valid coverage (256+ cap), S4 flip 13%, S5 admits 49.7% / + ECE 0.246; p50 264–276 ms; S1 $0.07/1k. No class wins on quality. + Do not merge Banking77 87% / 76.3% / 79.67%. +- **Jevals/jevals-data (Contract as feedstock).** CC-BY-4.0. Boards + + JSONL + suites. 2026-09-18 board, suite 0.1.0, 8 systems. Jev + banking77 acc 0.7967 / ECE 0.0981 / p50 467 ms / $0.043/1k on + *this* board. Recompute-from-logs; not a ranking. +- **taro1985/dual-process-ai (Empirical as a productized metaphor).** + MIT. Kahneman S1 decide / S2 generate. Routing fails open; safety + fails closed. Routing accuracy **unmeasured**. Keyword fallback ≠ + S1. +- **simonmesmith/jev-arc-agi-v1-experiment (Empirical as a + negative).** LICENSE absent. Direct Jev 4/400 (1%), 1.125%, ~$2.32, + 10 min. Cell-wise Choice. Combinatorial ≠ extractive. +- **ikermoel/open-alternative-jev (Empirical as packed-logprob + economics).** Apache-2.0. Not a Jev reproduction. RACE-H 92.9% @ + 4.55 q/s; interference 6–9%. +- **wfzyx/von (Empirical as extreme speed/econ surface).** + Apache-2.0. 14 MB SAN; authored144 52.6%. Not jev-local stub, not + kev, not a Jev replica. Do not copy vs-Jev table. +- **jaredpalmer/kev (light delta).** 100★ this pass. No species + rewrite. + +Cross-repo addition: (bn) extractable-from-state is the placement +axis; (bo) retrieve-then-state is VOI with a named receipt; (bp) +Harbor-style class bake-off includes constrained LLMs and +baselines; (bq) public JSONL boards are feedstock, not rankings; +(br) dual-process is S1 decide / S2 generate with unmeasured +routing accuracy; (bs) combinatorial assembly is not extractive +keep/drop; (bt) local `/v1/systemone` has three surfaces (stub / +pointer / tiny SAN). + + diff --git a/research/notes.md b/research/notes.md index c75505c..5381ddc 100644 --- a/research/notes.md +++ b/research/notes.md @@ -3043,3 +3043,230 @@ Cards: `judgment-class.md`; `validation.md`; `applied-mappings.md` §2 / §5; `mappings.md` §4 / §9; `mixed-architecture.md`; `faq.md`; `mental-models.md`; `methods-catalog.md`; `toolbox-mapping.md`; `agent-self-assessment.md`. No wrapper. + +## 49. Boundary map, Harbor bake-off vs constrained LLMs, dual-process (2026-09-18 ~15:52 Boise) + +America/Boise ~15:52 = 21:52 UTC (archive 214908). Docs-only fold into +open PR #2 (`cursor/augustus-store-envelope-00b4`). Not a competing +PR. Archer 27B drop still **WATCH** (still ~2026-09-19; not landed). +X MCP flap blocked discourse this hour — no tweets invented. Identity +lock vs `typesafe-ai` / `tenbin` / `decision-first` holds. No wrapper, +no install.sh, no copied `uv` / `von serve` / pip how-tos. No invented +metrics. Do not re-fold §48 items, blackwood-rlcd, open-jev-laya-bench, +decision-token LoRA, jevgate, wellposed, jev-reflex-autonomy-lab, Abide, +kev species, jevpandas, bitrate-advisor. + +Backend-agnostic: these are **placement / measurement** cards +(extractable-from-state axis; Harbor-style frozen-protocol bake-off; +recompute-from-logs feedstock; Kahneman S1 decide / S2 generate; +combinatorial assembly ≠ extractive keep/drop; packed one-forward open +LLM economics; tiny non-AR local surface). TypeSafe Jev is the +exemplar in the READMEs, not a monopoly. kev is a **star/activity +delta only** this hour (100★ this pass; user cited 97). + +### HIGH + +1. **[`Zaious/jev-capability-atlas`](https://github.com/Zaious/jev-capability-atlas)** + (README claims MIT; GitHub SPDX **NOASSERTION** this pass; Python; + created 2026-09-18T21:30:40Z; 0★; unofficial, not TypeSafe). + Independent **when-it-holds map with real API receipts**, not a + leaderboard. They cite [jev-benchmarks](https://github.com/AbdelStark/jev-benchmarks) + and thaiexam charts; they do not redo them. Primary mental model + this hour: **place the task on one axis** — + + > Is the correct answer fully recoverable from the `state` you hand + > the model, or does it require outside knowledge that is not in + > `state`? + + **Self-contained (strong):** classification they cite from + jev-benchmarks (AG News 91%, Banking77 **87%** — a *different* + protocol from DMB 76.3% and jevals.com 79.67%; do not merge); + citation support-checking with claim + quote both given + (`paraphrase_support` and `reversed_meaning_high_overlap` both + correctly judged); sarcasm when the trigger is in the given text. + **Not self-contained (fails, often confidently):** pure recall + without a supporting passage; scoring that needs a whole-field + comparison; genuinely overlapping categories. + + **History suite** (`suites/history-recall-context/`, receipt + `runs/2026-09-19.json`; N=3, single annotator, Chinese history + only — qualitative axis proof, **not** a knowledge-breadth + estimate). Cite the table, not the README's 30-second slogan: + + | Case | Question | Context? | Pick | Conf | Dist | Correct? | + |---|---|---|---|---|---|---| + | A | 2nd Qing emperor | No | Yongzheng | 0.90 | 0.04 / 0.93 / 0.03 | ✗ (Kangxi by popular convention; ground truth itself contested) | + | B | 7th Qing (obscure) | No | Xianfeng | 0.07 | 0.25 / 0.37 / 0.38 | ✓ by luck (near-flat) | + | C | Same as B | Passage in state | Xianfeng | 0.97 | 0.00 / 0.02 / 0.98 | ✓ | + + Informal unsaved B rerun picked Daoguang at 0.08 — both near-flat. + Teaching: **bare memory is unreliable; reading comprehension over + supplied text is reliable.** Conflating the two misjudges the risk. + Retrieve first; put the passage in `state`. + + **Not a state machine internally** (distributed LM understanding — + paraphrase vs reversed-meaning contrast). **Correctly placed as a + component node** in *your* code ("a node in your state machine" is + the architectural instinct; "its internals are a state machine" is + not). Confidence is a **statistic from the distribution** (RLCD + trains the distribution; `confidence` is arithmetic on how peaked + it is), not a second trained correctness score. Calibration is + **population-level** and can fail **dangerous-high**: DAIR Emotion + via jev-benchmarks — 48% acc, mean conf **0.819**, 16% of items + p(correct)=0. That is the "blind guessing" concern that actually + lands: overlapping blurred categories, overconfident. + + **Browser-use strength is placement, not vision.** Third-party + `jev-ultrafast` (Browser Use): Google Flights 9.5s → 7.1s; 12-task + vs Playwright MCP 1.5× faster / 1.6× cheaper, comparable accuracy; + standalone loop ~1.8s / $0.0005 / 97% (their figures, not + re-run). Jev is text-only. What held: **DOM snapshot as `state`** + (a visual task translated into extractive text) + **speculative + fan-out** over candidate elements. Not screenshots. Same + component-node placement as lizard-agent / solari-reflex. + + Pattern: **boundary map / placement judgment.** Cards: + `mental-models.md` (primary); `faq.md`; `question-design.md`; + `applied-mappings.md` §2 / §6; `mappings.md` §3 / §6; `mixed-architecture.md`. + +2. **[`nibzard/decision-model-benchmark`](https://github.com/nibzard/decision-model-benchmark) + (DMB)** (LICENSE **absent** this pass; Python; created + 2026-09-18T21:48:07Z; 0★; no vendor sponsorship). Independent + **frozen protocol**: typesafe:jev vs **8 constrained LLMs** vs + **3 deterministic baselines** (keyword / majority / random). Five + suites, 60 cells, **$28.34** measured spend; every raw log + published. **`results/v2/v2.md` is the report of record** (v1 + majority-prior defect corrected; mixed protocol versions merged + by explicit `--allow-protocol-mix`). Harbor-style measurement + exemplar: protocol frozen before the run; negative results ship; + recompute from logs; unknown usage is never a measured zero; later + runs replace cells whole. Do not copy `uv` how-to. + + **jev (v2, named receipt, not a ranking):** + + | Suite | Acc | Notes | + |---|---|---| + | S1 banking77 77-way | **76.3%** | ECE 0.083; cost/1k **$0.07**; p50 **274 ms** | + | S2 SMS spam | **93.0%** | ECE 0.249 (overconfident vs GLM 0.042) | + | S3 cardinality | **100%*** | valid coverage **72.7%** (225 failed = **256+ Choice cap**, `400 Too many choices`); 254–255 still 100% | + | S4 order permute | **76.7%** | flip **13%** (worst LLM 37%) | + | S5 honesty | **14.3%** | admits-ignorance **49.7%** vs most LLMs 97.3–100% (gpt-5.4-mini **64.7%**); ECE **0.246**; mean conf on no-good **0.543** | + + p50 across jev cells **264–276 ms**, flat 2→255 options. Fastest + *measured* vs thinking-mode LLMs is **10–16×**, **1.2×** vs + gpt-oss-120b on Cerebras — **not** the vendor 40–200× claim + against LLMs left in their slowest default. gpt-oss-120b banking + **81.3%**; glm-5.3 **80.4%** / spam **94.9%**. **No class wins on + quality.** Axes that *do* separate: latency, cost, schema-validity + (jev 0% malformed), Choice cap, honesty. Two OpenAI models sit + *below* the 87.7% majority baseline on spam. + + **Do not collapse Banking77:** atlas/jev-benchmarks **87%**, DMB + **76.3%**, jevals.com 2026-09-18 **79.67%** — n / split / protocol. + Cite named receipts. Pattern: **Harbor/jevals practice + + when-to-use-vs-constrained-LLM table.** Cards: `validation.md`; + `judgment-class.md`. + +3. **[`Jevals/jevals-data`](https://github.com/Jevals/jevals-data)** + (CC-BY-4.0; created 2026-09-18T21:36:00Z; 0★). Release boards + + per-decision JSONL run logs + suite files behind + [jevals.com](https://jevals.com). Cite "Jevals (jevals.com), + release \". **2026-09-18** board: suite **0.1.0**, **8 + systems**, tasks banking77 / helpsteer2 / pubmedqa. Formulas at + https://jevals.com/methodology/. **Recompute-from-logs pattern**, + not a third ranking to merge with DMB or atlas. + + Jev on *this* board (native probabilities; n=300 × 5 repeats; + **not** DMB n): banking77 acc **0.7967**, ECE **0.0981**, p50 + **467 ms**, cost/1k **$0.043**; helpsteer2 Score acc **0.4127** + (label-prior 0.4167 — barely above chance on that primitive); + pubmedqa Noul acc **0.9127**, ECE **0.0504**, p50 **438 ms**. Do + not dump the board as a ranking. Feedstock for Harbor/jevals: + frozen suite files, item ids pointing at public datasets (item + text not republished), run header + per-decision rows + (`item_id`, `epoch`, `target`, `order_seed`, `output`, + `usage`, `cost_usd`, `seconds`, `malformed`, `refusal`, + `retries`). Cards: `validation.md`. + +4. **[`taro1985/dual-process-ai`](https://github.com/taro1985/dual-process-ai)** + (MIT; Python; created 2026-09-18T20:55:08Z; 0★). Explicit + **Kahneman S1 (Jev) / S2 (Gemini)** design pattern. S1: typed + decision + confidence, ~70–500 ms. S2: free-form text. Mechanism: + `confidence ≥ τ → S1 decides; else escalate to S2`. **Routing + fails open** (low conf → S2; the router never refuses). **Safety + fails closed** (unparseable denied). **Routing accuracy is not + measured yet** (misroute rate / escalation rate / Brier on + held-out — the numbers this project needs and does not have). + Without a key it runs **degraded keyword mode** — **not an + equivalent S1** (no calibrated confidence; anything not on the + allowlist escalates). Do not copy hooks / Discord / pip. Crossover + metaphor for **business/life**, not only SWE: cheap classify / + route / gate on S1; write / reason / generate on S2. Same split as + jev-reflex-autonomy-lab (S1 keeps control) and jev-hermes (route ≠ + memory), productized as a cascade. Cards: + `mixed-architecture.md`; `toolbox-mapping.md`; `mappings.md` §2; + `agent-self-assessment.md`; `faq.md`. + +### MED (brief) + +5. **[`simonmesmith/jev-arc-agi-v1-experiment`](https://github.com/simonmesmith/jev-arc-agi-v1-experiment)** + (LICENSE **absent** this pass; Python; created 2026-09-18T21:32:19Z; + 0★). Direct Jev on ARC-AGI-1 public eval: **4/400 (1%)** fully + solved; **1.125%** task-weighted (4 + 0.5 / 400); **5/419** exact + grids; **~$2.32**; **10 minutes**; `jev-1.13.0`; two guesses per + grid. **Cell-wise Choice assembly:** height/width 1–30, then one + of ten colours per cell; cells do not see one another; transpose + for the second guess. Dimensions ~**90%** on first attempt; rarely + a complete grid. No partial credit for cells. Teaching: + **combinatorial grid tasks ≠ extractive keep/drop.** A program + library with no competing candidates **never called Jev** (6 + tasks / 1.5%). Combined policy 2.375% includes hand-written + programs — tells you less about Jev. Frozen protocol, traces, + `PROTOCOL.md`. Not a ceiling claim. Cards: `validation.md`; + `mappings.md` §9; `faq.md`. + +6. **[`ikermoel/open-alternative-jev`](https://github.com/ikermoel/open-alternative-jev)** + (Apache-2.0; Python; 3★; Space + [`IkerMoel/open-alternative-jev`](https://huggingface.co/spaces/IkerMoel/open-alternative-jev)). + Packed **one-forward** System One on **any open-weights LLM** + (HF + vLLM). **Not a Jev reproduction** — packages a capability + chat APIs hide; no claim about how Jev works. RACE-H (250 + passages × 4 questions, n=1000) on Qwen3.6-27B 8-bit: packed + **92.9% @ 4.55 q/s** vs one-at-a-time 92.6% @ 1.66; 2.5× fewer + tokens. Interference **6–9%** of answers move vs a 2.7% numeric + noise floor; order rotation 8% MMLU / 2.4% RACE-H. Temperature + scaling ECE 5.4%→2.1% MMLU, 2.8%→1.1% RACE-H. Small 4B packing + costs 2.8 points — use `separate` when accuracy at stake. On + vLLM, prefix-cache `separate` is fastest. Economics/architecture + of **open replicas on the constrained-AR / logprob path**, not + trained decision-only. Cards: `judgment-class.md`. + +7. **[`wfzyx/von`](https://github.com/wfzyx/von)** (Apache-2.0; + Python; 3★). **14 MB** Needle 3 SAN; non-AR local `POST + /v1/systemone` drop-in; sub-15 ms CPU *claim* / ~**38 ms** embed + in their table; ~28 MB RAM. Default needle **52.6%** balanced acc + on OpenJev `authored144` — **not a calibrated Jev replica**. Other + backends (berta-v3 / modern / laya) are optional heavier heads. + **Do not copy the vs-Jev ranking table** (includes a speculative + Jev weight estimate). Distinguish: jev-local **stub until hf**; + kev **trained pointer** on Qwen2.5-0.5B; von **tiny SAN** at the + extreme of the speed/econ class. Cards: `judgment-class.md`; + `faq.md`. + +8. **[`jaredpalmer/kev`](https://github.com/jaredpalmer/kev) delta** + — still active this hour; **100★** this pass (user cited 97; + prior fold 61★). Pushed through 2026-09-18T21:51Z. **Light note + only.** No species rewrite. Hub weights + NOTA training remain + §45. + +### Not this hour + +Archer drop **not landed**. X discourse **blocked** (MCP flap) — no +invented tweets. Do not re-fold §48. + +Cards: `mental-models.md` (boundary map); `validation.md` (DMB + +jevals-data + ARC); `judgment-class.md` (vs constrained LLM; von; +open-alternative-jev); `mixed-architecture.md` (dual-process; +component node; DOM-as-text); `faq.md`; `mappings.md` §2 / §3 / §6 / +§9; `applied-mappings.md`; `question-design.md`; `methods-catalog.md`; +`toolbox-mapping.md`; `agent-self-assessment.md`. No wrapper. diff --git a/research/refresh-log.md b/research/refresh-log.md index c9fe7c0..c9569f6 100644 --- a/research/refresh-log.md +++ b/research/refresh-log.md @@ -464,6 +464,31 @@ ecosystem, CHANGELOG, README. - notes.md §48; sources.json; findings.md batch #32. No wrapper. +## 2026-09-18 21:52 UTC — boundary map / Harbor bake-off / dual-process (~15:52 Boise) + +- Folded into open PR #2 (`cursor/augustus-store-envelope-00b4`). + Docs-only. Not a competing PR. Archer 27B drop still **WATCH**. + X MCP flap blocked discourse this hour — no tweets invented. +- HIGH: jev-capability-atlas (extractable-from-state axis; history + suite A/B/C table; component node; dangerous-high ECE; DOM-as-text); + DMB v2 (frozen protocol vs constrained LLMs; 76.3/93.0/256+; + 264–276 ms; $0.07/1k); jevals-data CC-BY-4.0 feedstock; + dual-process-ai (S1 decide / S2 generate; routing accuracy + unmeasured). +- MED: ARC-AGI Direct Jev 4/400; open-alternative-jev RACE-H 92.9% + @ 4.55 q/s (not a reproduction); von 14 MB SAN (not a replica); + kev 100★ light delta. +- Do not merge Banking77 87% / 76.3% / 79.67%. Do not copy vs-Jev + tables or uv/von-serve how-tos. +- Cards: SKILL.md, mental-models (primary boundary map), + judgment-class (vs constrained LLM; von; open-alternative-jev), + validation (DMB + jevals-data + ARC), mixed-architecture + (dual-process; component node; DOM-as-text), faq, mappings + §2/§3/§6/§9, applied-mappings, question-design, methods-catalog, + toolbox, agent-self-assessment, ecosystem, CHANGELOG, README. +- notes.md §49; sources.json; findings.md batch #33. No wrapper. + + diff --git a/research/sources.json b/research/sources.json index 4f6a5d5..75bc086 100644 --- a/research/sources.json +++ b/research/sources.json @@ -1,6 +1,6 @@ { "refresh_cadence": "hourly", - "retrieved": "2026-09-18T20:52Z", + "retrieved": "2026-09-18T21:52Z", "sources": [ { "kind": "docs", @@ -1423,6 +1423,60 @@ "title": "vinnylarouge/jevlike (star spotcheck)", "url": "https://github.com/vinnylarouge/jevlike", "note": "866 stars this pass. notes.md \u00a748." + }, + { + "kind": "github", + "title": "Zaious/jev-capability-atlas", + "url": "https://github.com/Zaious/jev-capability-atlas", + "note": "Independent when-it-holds map with API receipts, not a leaderboard. README MIT / GitHub SPDX NOASSERTION. History suite A wrong@0.90 / B 0.07 / C right@0.97. notes.md \u00a749." + }, + { + "kind": "github", + "title": "nibzard/decision-model-benchmark", + "url": "https://github.com/nibzard/decision-model-benchmark", + "note": "DMB. Frozen protocol jev vs 8 constrained LLMs vs baselines. v2 report of record. LICENSE absent this pass. notes.md \u00a749." + }, + { + "kind": "github", + "title": "Jevals/jevals-data", + "url": "https://github.com/Jevals/jevals-data", + "note": "CC-BY-4.0 release boards + per-decision JSONL + suites for jevals.com. 2026-09-18 board. Recompute-from-logs. notes.md \u00a749." + }, + { + "kind": "github", + "title": "taro1985/dual-process-ai", + "url": "https://github.com/taro1985/dual-process-ai", + "note": "MIT. Kahneman S1 (Jev) / S2 (Gemini). Routing accuracy unmeasured. Keyword fallback not equivalent S1. notes.md \u00a749." + }, + { + "kind": "github", + "title": "simonmesmith/jev-arc-agi-v1-experiment", + "url": "https://github.com/simonmesmith/jev-arc-agi-v1-experiment", + "note": "Direct Jev ARC-AGI-1 4/400 (1%), ~$2.32. LICENSE absent this pass. Combinatorial \u2260 extractive. notes.md \u00a749." + }, + { + "kind": "github", + "title": "ikermoel/open-alternative-jev", + "url": "https://github.com/ikermoel/open-alternative-jev", + "note": "Apache-2.0. Packed one-forward on open LLMs. RACE-H 92.9% @ 4.55 q/s. Not a Jev reproduction. notes.md \u00a749." + }, + { + "kind": "huggingface", + "title": "IkerMoel/open-alternative-jev (Space)", + "url": "https://huggingface.co/spaces/IkerMoel/open-alternative-jev", + "note": "Live demo for open-alternative-jev. notes.md \u00a749." + }, + { + "kind": "github", + "title": "wfzyx/von", + "url": "https://github.com/wfzyx/von", + "note": "Apache-2.0. 14MB Needle SAN local /v1/systemone. authored144 52.6%. Not a Jev replica. Do not copy vs-Jev table. notes.md \u00a749." + }, + { + "kind": "github", + "title": "jaredpalmer/kev (star delta)", + "url": "https://github.com/jaredpalmer/kev", + "note": "100 stars this pass (user cited 97). Light activity delta only. notes.md \u00a749." } ] } From 900ba5d2625f365160ae9d9075ea2d8e444d62c9 Mon Sep 17 00:00:00 2001 From: Cursor Agent Date: Fri, 18 Sep 2026 22:30:40 +0000 Subject: [PATCH 08/43] Fold GLiNER2.5 extractive compaction as encoder keep/drop Place m-newhauser/gliner25-compaction as architecture notes: pointer compaction vs summarizers; fail-closed keep_full under a mutation envelope; same job as pi/fast-jev-compaction with a GLiNER backend; shadowMode default true. Not Jev, not multimodal. Docs-only. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com> --- .agents/skills/augustus/SKILL.md | 12 +- .../references/agent-self-assessment.md | 16 ++- .../augustus/references/applied-mappings.md | 18 ++- .../references/composition-algebra.md | 4 +- .agents/skills/augustus/references/faq.md | 38 ++++- .../augustus/references/judgment-class.md | 30 +++- .../skills/augustus/references/mappings.md | 17 +++ .../augustus/references/mental-models.md | 6 +- .../augustus/references/methods-catalog.md | 4 +- .../augustus/references/mixed-architecture.md | 18 ++- .../augustus/references/toolbox-mapping.md | 2 +- CHANGELOG.md | 11 ++ README.md | 9 +- docs/ecosystem.md | 10 +- research/archive/findings.md | 33 +++++ research/notes.md | 133 ++++++++++++++++++ research/refresh-log.md | 22 +++ research/sources.json | 12 +- 18 files changed, 359 insertions(+), 36 deletions(-) diff --git a/.agents/skills/augustus/SKILL.md b/.agents/skills/augustus/SKILL.md index cfdb2cd..94e187f 100644 --- a/.agents/skills/augustus/SKILL.md +++ b/.agents/skills/augustus/SKILL.md @@ -1,6 +1,6 @@ --- name: augustus -description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill, open multimodal RLCD / blackwood), announced open decision-model (Watch — still not landed), constrained-AR (TypeAR, pcdServer, decision-token LoRA), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"wait for Archer vs open multimodal RLCD\", \"screenshot/DOM candidates → typed Choice\", \"eval path\", \"jevals\", \"Harbor taskset\", \"shared bake-off ECE/NLL/Brier\", \"LLM-as-judge is not the System One score\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge (allowlist *proves*; fail-open cannot block), \"missing other → confident wrong Choice\", \"lint the Jev request\", \"training confronts Choice other / none-of-the-above\", \"S1 reflex keeps control / optional S2 one-use advice\", \"soft AGENTS.md rules vs linter\" (Abide / jev-pref), \"edit-phase vs turn-phase observation window\", banded confidence fail-open preference lint, \"extractive selection / pointer-not-generator\", \"local /v1/systemone drop-in (stub until hf scorer)\", \"dataframe-native semantic columns\", \"route≠memory\", \"advisory sidecar receipts\", \"structure induction over bags\", \"AST ∩ semantic lint\", \"extractable-from-state / retrieve first\", \"decision-model vs constrained-LLM bake-off\", \"dual-process S1 decide / S2 generate\", \"combinatorial grid ≠ extractive\", TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, or \"formally verify with Jev\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." +description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill, open multimodal RLCD / blackwood), announced open decision-model (Watch — still not landed), constrained-AR (TypeAR, pcdServer, decision-token LoRA), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"wait for Archer vs open multimodal RLCD\", \"screenshot/DOM candidates → typed Choice\", \"eval path\", \"jevals\", \"Harbor taskset\", \"shared bake-off ECE/NLL/Brier\", \"LLM-as-judge is not the System One score\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge (allowlist *proves*; fail-open cannot block), \"missing other → confident wrong Choice\", \"lint the Jev request\", \"training confronts Choice other / none-of-the-above\", \"S1 reflex keeps control / optional S2 one-use advice\", \"soft AGENTS.md rules vs linter\" (Abide / jev-pref), \"edit-phase vs turn-phase observation window\", banded confidence fail-open preference lint, \"extractive selection / pointer-not-generator\", \"encoder GLiNER compaction vs Jev Score compaction (same job; pointer not summarizer)\", \"fail-closed keep_full under mutation envelope\", \"shadow-mode compaction rollout\", \"local /v1/systemone drop-in (stub until hf scorer)\", \"dataframe-native semantic columns\", \"route≠memory\", \"advisory sidecar receipts\", \"structure induction over bags\", \"AST ∩ semantic lint\", \"extractable-from-state / retrieve first\", \"decision-model vs constrained-LLM bake-off\", \"dual-process S1 decide / S2 generate\", \"combinatorial grid ≠ extractive\", TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, or \"formally verify with Jev\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." license: MIT metadata: version: 0.3.0 @@ -54,7 +54,7 @@ classical method you already trust, substitute it, classify the win "paraphrase brittleness", "allowlist then judge", "TOCTOU-of-Noul", "Jev inside the database / sqlite-jev", "Jev picks bitrate / join order / the model", "wait for Archer", "lint the request / missing - other", "training confronts Choice other / none-of-the-above", "soft AGENTS.md rules vs the linter", "screenshot Choice / omni System One", "extractive quotes / pointer not generator", "local drop-in vs stub scorer", "route vs memory", "when does it hold / extractable from state", "decision model vs constrained LLM", "dual-process S1/S2", or "combinatorial grid vs extractive": + other", "training confronts Choice other / none-of-the-above", "soft AGENTS.md rules vs the linter", "screenshot Choice / omni System One", "extractive quotes / pointer not generator", "compaction summarize vs pointer", "encoder vs Jev compaction backend", "shadow-mode compaction rollout", "local drop-in vs stub scorer", "route vs memory", "when does it hold / extractable from state", "decision model vs constrained LLM", "dual-process S1/S2", or "combinatorial grid vs extractive": read `references/faq.md`, then `references/mental-models.md`, then `references/mixed-architecture.md`, then @@ -114,13 +114,13 @@ classical method you already trust, substitute it, classify the win | Familiar method | Judgment shape | Detail | |---|---|---| | Mental models across domains (not SWE-only) | EU, abstention, VOI, MCDA, SDT, search/control, Leveson, NATM/Norman/snap-fit; **extractable-from-state boundary map** (self-contained vs needs outside knowledge) | `references/mental-models.md` | -| Judgment-model class (Jev is exemplar, not monopoly) | Species: decide / locate (GLiNER) / categorize (GLiClass) / rank / perceive; open heads include encoder DeBERTa, LoRA distill, openjev-lm, kev | `references/judgment-class.md` | +| Judgment-model class (Jev is exemplar, not monopoly) | Species: decide / locate (GLiNER) / categorize (GLiClass) / rank / perceive; open heads include encoder DeBERTa, LoRA distill, openjev-lm, kev. Compaction job is backend-agnostic (Jev Score/Noul vs GLiNER2.5 encoder) | `references/judgment-class.md` | | Open weights vs constrained decoding vs proprietary API | Three open paths: encoder open-jev / AR constrained decode (TypeAR + pcdServer; decision-token LoRA; packed one-forward logprob on open LLMs) / trained decision-only (Laya + ONNX port, Nimble, kev, **blackwood-rlcd** multimodal now, Archer Watch still Watch). Local `/v1/systemone` surfaces: jev-local (stub until `hf`), kev (trained pointer), von (tiny SAN — not a replica). Softmax over allowed tokens ≠ Noul | `references/judgment-class.md` (when-to-use table) | | Entropy as allocator (low / medium / high) | Typed low+medium decisions → System One marginals; high-entropy synthesis → frontier decoder. Product rhetoric, not a meter. **Hypothesis** | `references/judgment-class.md` | | Formal / semi-formal (proof vs judgment) | Sensor vs constraint vs searchlight; Alloy vs Apalache; DST trio; TOCTOU-of-Noul, AI×FM | `references/formal-methods.md` (one-screen: `references/formal-semi-formal.md`) | | Mixed architecture (judgment model + LLM) | Provider judges, LLM writes, code owns control; not a stack replacement. Advisory sidecar never changes host routing. Dual-process: S1 decides, S2 generates (routing accuracy unmeasured). Internals ≠ FSM; placement is a **component node**. Browser-use strength = DOM-as-text + speculative fan-out, not vision | `references/mixed-architecture.md` | -| Context sieve | Relevance Noul per block; always-keep set in code; stub + recall key | `references/applied-mappings.md#1-context-sieve` | -| Exact-text keep / drop | Choice include/exclude/mixed over candidates code already holds. Extractive quotes / pointer-not-generator (model never writes the excerpt) | `references/applied-mappings.md#2-exact-text-keep--drop` | +| Context sieve | Relevance Noul per block; always-keep set in code; stub + recall key. Encoder cousin: GLiNER2.5 retention Choice + char-offset spans (gliner25-compaction); fail-closed keep_full; shadowMode default | `references/applied-mappings.md#1-context-sieve` | +| Exact-text keep / drop | Choice include/exclude/mixed over candidates code already holds. Extractive quotes / pointer-not-generator (model never writes the excerpt; char-offset compaction same species) | `references/applied-mappings.md#2-exact-text-keep--drop` | | Environment / harness triage | Scan every step for env failure; LLM autopsy only on flags | `references/applied-mappings.md#3-environment--harness-triage` | | Moderation and ranking | Hold-before-publish vs graded rerank; fail policy per action | `references/applied-mappings.md#4-moderation-and-ranking` | | Skill / tool routing | Choice over a closed catalog + whether-anything-fits; code dispatches. Route ≠ memory: cheap intent gate skips memory tours on easy routes | `references/applied-mappings.md#5-skill--tool-routing` | @@ -133,7 +133,7 @@ classical method you already trust, substitute it, classify the win | Decision tables / circuits / state machines | Judgment predicates, code owns transitions | `references/mappings.md#3-semantic-predicates--decision-circuits` | | Retrieve + expensive relevance fn | Bounded rerank of a retrieved shortlist | `references/mappings.md#4-retrieval--bounded-semantic-reranking` (independent TREC DL2019 benchmark: Jev zero-shot best MAP 0.4748, nDCG@10 0.683 vs tuned monoBERT 0.718 — competitive, not dominant) | | Store as semantic index (SQL / SQLite / zoxide / dataframe) | Cheap exact predicates first; typed questions on the remainder. In-engine extension (sqlite-jev) vs CLI rewrite (jevql) vs path index (joxide) vs dataframe columns (jevpandas / jevframe) | `references/mappings.md#4-retrieval--bounded-semantic-reranking` | -| Soft judgment inside a hard envelope | Model may only match the deterministic policy or be more conservative (bitrate ABR; query-planner override-when-confident) | `references/mappings.md#12-runtime-assurance-sandwich-hypothesis`; `references/mappings.md#18-structural-prove--soft-remainder-hypothesis-as-domain-general-empirical-as-named-shapes` | +| Soft judgment inside a hard envelope | Model may only match the deterministic policy or be more conservative (bitrate ABR; query-planner override-when-confident; compaction mutations/shell operators → keep_full) | `references/mappings.md#12-runtime-assurance-sandwich-hypothesis`; `references/mappings.md#18-structural-prove--soft-remainder-hypothesis-as-domain-general-empirical-as-named-shapes` | | Value of information / gather as an act | Pay for another observation only if EV(decision) improves more than cost; abstain from calling *any* model when a regex already answers (meta-VOI) | `references/mappings.md#6-value-of-information--gather-as-an-enumerated-act` (**Hypothesis** until a labeled act/outcome log; 149-row receipt is Empirical as a shape) | | Signal detection / ROC | Criterion and operating point from costs and base rate, not accuracy | `references/mappings.md#7-signal-detection--criterion-not-accuracy` (**Hypothesis** for non-SWE plots) | | Org / safety control structure | Sensor ≠ constraint (Leveson); STPA if the sensor lies | `references/mappings.md#8-control-structure--sensor--constraint-leveson` | diff --git a/.agents/skills/augustus/references/agent-self-assessment.md b/.agents/skills/augustus/references/agent-self-assessment.md index 3dc94f6..4146c04 100644 --- a/.agents/skills/augustus/references/agent-self-assessment.md +++ b/.agents/skills/augustus/references/agent-self-assessment.md @@ -51,9 +51,13 @@ relying: foreman, pi-jev, pi-warden, winnow, fast-jev-compaction, jev-judgment. tool result with one relevance Noul before it enters context. Hide confident-no blocks behind a stub + recall key; always keep current instruction, recent turns, errors, and opaque blocks. winnow hides at - relevance ≤0.22; fast-jev-compaction asks two nouls per tool call - (should the call stay knowing it was made? should the result stay - verbatim?). + relevance ≤0.22; fast-jev-compaction asks two nouls per tool call + (should the call stay knowing it was made? should the result stay + verbatim?). Encoder-backend cousin: + [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) + — GLiNER2.5 retention Choice + exact character-offset copies; mutating + tools stay `keep_full`; low-confidence fails closed to `keep_full`; + `shadowMode` default true. Not a summarizer. Not Jev (`notes.md` §50). ## Non-negotiable boundaries @@ -66,6 +70,9 @@ relying: foreman, pi-jev, pi-warden, winnow, fast-jev-compaction, jev-judgment. only safe because a hard interlock or sandbox sits underneath. A gate that *selects* or *authorizes* a side effect fails closed instead (`mixed-architecture.md` prefilter table; `mappings.md` §18). + Compaction *drop* is that second kind: + [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) + fails closed to `keep_full` (`notes.md` §50). - Cache identical judgments (~120s) and deduplicate sibling calls into one in-flight request. - pi-warden measured cost makes continuous guarding viable: ~$0.00004 and @@ -80,6 +87,9 @@ relying: foreman, pi-jev, pi-warden, winnow, fast-jev-compaction, jev-judgment. Pointer-not-generator: the model points at line ids; code copies verbatim with place; *not found* is an answer ([jev-reviewer](https://github.com/choxos/jev-reviewer); `notes.md` §48). + Compaction: point at character offsets in the tool result + ([gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction); + `notes.md` §50). - Self-report fidelity: compare the agent's claimed action with its actual trace via decomposed Nouls (right tool? args match schema? result matches call?). Escalate on low confidence; never auto-retry. diff --git a/.agents/skills/augustus/references/applied-mappings.md b/.agents/skills/augustus/references/applied-mappings.md index 183decd..d9ee2ff 100644 --- a/.agents/skills/augustus/references/applied-mappings.md +++ b/.agents/skills/augustus/references/applied-mappings.md @@ -43,6 +43,18 @@ the rest per line; `kevinpita/pi-jev-context` hides (does not delete) older Pi history, always-keep user/system/todos, `/jev off` restores. Pi compaction cousins (`tamaratran/fast-jev-compaction`, `vava-nessa/pi-jev-compaction`) keep verbatim drop, never summarize. +**Encoder backend, same job (Empirical as README behavior, 2026-09-18 +~16:22):** +[gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) +— GLiNER2.5 (`fastino/gliner2.5-base-v1`) chooses +`keep_full` / `keep_evidence` / `keep_call_only` / `drop` and copies +exact character-offset spans. Not a summarizer. Mutating tools / unknown +shell / control operators → `keep_full`. Low-confidence or invalid +evidence **fails closed to `keep_full`** (the reduction is the +irreversible act; from the evidence side this *looks* like keep-on-error). +`shadowMode` defaults true. Characters, not tokens; no published +retention-quality rates (`notes.md` §50). Family: +`judgment-class.md`. Do not copy the plugin. Local teacher-copy for the same hole: [`SargeDev/jev-gate-student-b`](https://huggingface.co/SargeDev/jev-gate-student-b) (Qwen2.5-0.5B LoRA; P(relevant) from yes/no logits; 148,160-row @@ -92,7 +104,11 @@ code numbers sentences; one broadcast (Choice/Noul/Score + per-sentence Nouls); the model never writes; `redecide` retunes thresholds on the log. [jev-reviewer](https://github.com/choxos/jev-reviewer) — the model **points at line ids**; code copies verbatim quotes with place; *not -found* is an answer. Computer-use cousin: +found* is an answer. Compaction cousin +([gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction), +~16:22): the model points at **character offsets** in a tool result; +code copies those bytes; a generator summary is the rejected species +(`notes.md` §50). Computer-use cousin: [solari-reflex](https://github.com/hitakshiA/solari-reflex) — structured observation → typed decision → verified act; **no screenshots**; model output never becomes a selector (`notes.md` §48). diff --git a/.agents/skills/augustus/references/composition-algebra.md b/.agents/skills/augustus/references/composition-algebra.md index 34425bb..cb6f80e 100644 --- a/.agents/skills/augustus/references/composition-algebra.md +++ b/.agents/skills/augustus/references/composition-algebra.md @@ -146,7 +146,9 @@ Reusable shapes when generating applications: batched question set per step → execute via AX actions. 9. **Shadow-mode harness** (jev-harness): policy + confidence gate + shadow mode + offline eval CLI replaying fixtures, asserting on actions; 24-row filter 48.9 s - (Claude CLI) vs 1.3 s Jev at concurrency 8. + (Claude CLI) vs 1.3 s Jev at concurrency 8. Compaction rollout: + gliner25-compaction public default `shadowMode: true` (log proposed + reduction; do not replace history) (`notes.md` §50). Calibration warning (calibre): routing thresholds and ROI do **not** transfer across datasets — every gate is a per-dataset measurement (see validation.md). diff --git a/.agents/skills/augustus/references/faq.md b/.agents/skills/augustus/references/faq.md index 8348704..ddb0ceb 100644 --- a/.agents/skills/augustus/references/faq.md +++ b/.agents/skills/augustus/references/faq.md @@ -145,6 +145,10 @@ Hole first, logo last. These are **species**, not aliases code. Paper: [GLiNER](https://arxiv.org/abs/2311.08526). Local multi-head GLiNER2.5 (fastino-ai) can also classify and extract relations on a laptop — discourse, not a measured 36× (`notes.md` §25). + Compaction receipt: [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) + uses that local multi-head as **categorize** (retention action) plus + **locate** (character-offset spans); code copies; not a summarizer + (`notes.md` §50). - **GLiClass (categorize):** one forward pass over text + *all* labels; sigmoid multi-label or softmax single-label. Use for large or changing tag sets. Scores are class affinities, not automatically a gateable @@ -184,7 +188,9 @@ labels on a GLiNER2 encoder, not Choice / Score / Noul). Empirical open encoder next to GLiClass; not a weight clone. "like jev" is discourse. A GLiGuard score is not a proof. LLM I/O safety is not a coding-agent tool gate (rh-guard for reward-hacking; jevgate shape for allowlist -∩ remainder; Abide for project soft rules on diffs). `judgment-class.md`. +∩ remainder; Abide for project soft rules on diffs; +gliner25-compaction for extractive context compaction — Fastino +sibling class, not GLiGuard). `judgment-class.md`. ## Can I threshold CLIP / SigLIP as a safety gate? @@ -361,8 +367,11 @@ conservative; missing the model returns the policy's answer. mmalisper's JOB planner: Postgres plans first; Jev overrides only when confident (+12% geomean author-reported; join-order Choice alone was 2× slower). routeKit / jev-claw / Higgsfield: classify requirements; -code picks the generator. The envelope is load-bearing -(`mappings.md` §12, §15, §18). +code picks the generator. Compaction cousin (encoder, not Jev): +gliner25-compaction — mutating tools / shell operators prove +`keep_full`; the model may only match that or be more conservative; +uncertain fails closed to `keep_full`. The envelope is load-bearing +(`mappings.md` §12, §15, §18; `notes.md` §50). ## Wait for Archer to ship omni System One? @@ -434,14 +443,31 @@ the stub is not a bake-off. `judgment-class.md`; `notes.md` §48, §49. ## Should the model write the quote / the citation / the click? No. Extractive keep/drop: code already holds the sentences, line ids, -or numbered controls; the model **selects**; code **copies or clicks**. +character offsets, or numbered controls; the model **selects**; code +**copies or clicks**. [testimonial-miner](https://github.com/AppitStudio/testimonial-miner) assembles quotes from per-sentence Nouls and `redecide`s without new calls. [jev-reviewer](https://github.com/choxos/jev-reviewer) points at ids; *not found* is an answer. [solari-reflex](https://github.com/hitakshiA/solari-reflex) -never lets model output become a selector. Generation is only for +never lets model output become a selector. Compaction is the same +species: [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) +copies exact source spans; a prose summary of the tool result is +generation, not keep/drop (`notes.md` §50). Generation is only for TYPE/prose when something must be written. `applied-mappings.md` §2; -`notes.md` §48. +`notes.md` §48, §50. + +## Should compaction summarize? + +No. Pointer/extractive compaction and generator summarizers are +different species. The former is auditable (every kept byte occurs in +the input). The latter can invent. Same job as +fast-jev-compaction / pi-jev-compaction (Jev Noul/Score backends); +GLiNER2.5 is an encoder backend. Mutating tools and shell operators +stay `keep_full` in **code**. Low-confidence / invalid evidence fail +**closed to `keep_full`** — the *reduction* is the irreversible act, +unlike Abide / jevgate fail-open. Ship `shadowMode` first (default +true: log, do not replace history). Not Jev. Not multimodal. +`judgment-class.md`; `notes.md` §50. ## Is routing the same as memory? diff --git a/.agents/skills/augustus/references/judgment-class.md b/.agents/skills/augustus/references/judgment-class.md index 12f9b08..12ca816 100644 --- a/.agents/skills/augustus/references/judgment-class.md +++ b/.agents/skills/augustus/references/judgment-class.md @@ -81,6 +81,24 @@ below, next to the when-to-use table. records + relations in one schema. Discourse (2026-09-18): same *agentic decision* jobs as Jev, local / free / laptop; a 36× Browser Use cost claim is a tweet, not a re-run (`notes.md` §25). + **Extractive compaction (Empirical as README behavior, 2026-09-18 + ~16:22):** + [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) + (Apache-2.0) is the named *job* on this family, not a new species. + Default checkpoint `fastino/gliner2.5-base-v1`. One encoder does + **categorize** (retention Choice: `keep_full` / `keep_evidence` / + `keep_call_only` / `drop`) and **locate** (character-offset spans); + code copies verbatim. **Not Jev, not a Noul, not a prose + summarizer, not multimodal.** Same compaction hole as + fast-jev-compaction / pi-jev-compaction (Jev Noul/Score backends); + Augustus stays backend-agnostic. Mutating tools and shell operators + are a hard `keep_full` envelope; low-confidence / invalid evidence + fail closed to `keep_full` (the *reduction* is the irreversible + act — contrast many fail-open Jev preference gates). Public default + `shadowMode: true`. Reduction is measured in characters, not + tokens; no published retention-quality rates (`notes.md` §50). + Fastino/GLiGuard sibling *class*, not a GLiGuard safety-schema + clone. Do not copy the plugin install. - **Decide.** Typed Choice/Score/Noul with a decision/proper-scoring objective. That is Jev's product claim. Open heads copy the *shape*; distillation copies the *teacher* (openjev-lm, jev-gate-student-b). @@ -121,7 +139,11 @@ below, next to the when-to-use table. reward-hack gate (README fetched this pass); jevgate is the allowlist-then-judge shape (`mappings.md` §18); [Abide](https://github.com/coldteadotai/abide) is project-instruction - soft rules on diffs (`mixed-architecture.md`). Different holes. + soft rules on diffs (`mixed-architecture.md`). + [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) + is extractive context compaction (locate + categorize on tool + transcripts; Fastino sibling class, not a GLiGuard clone) + (`notes.md` §50). Different holes. Do not point one model at both, and do not copy a hook install here. **Aggregation is policy-in-code, already taught.** The README's @@ -171,7 +193,11 @@ multi-head (GLiNER2.5) can do both plus relations. Classification sigmoid/softmax *can* be a cheap multi-label sieve. It is not, without your calibration plot, a decision API. Great for "which of these 80 tags fire" or "which spans are the allergy / the amount / the verb"; -not a silent fail-closed authorize. The 255-option Choice limit is +not a silent fail-closed authorize. Named compaction job +([gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction)): +locate + categorize, then code copies exact offsets; fail-closed +`keep_full` is the *reduction* policy, not a Noul (`notes.md` §50). +The 255-option Choice limit is Jev's, not the class's — this family is why. Rule of composition (`applied-mappings.md` §4): diff --git a/.agents/skills/augustus/references/mappings.md b/.agents/skills/augustus/references/mappings.md index 118f001..4cc3e0a 100644 --- a/.agents/skills/augustus/references/mappings.md +++ b/.agents/skills/augustus/references/mappings.md @@ -505,6 +505,16 @@ the sandwich must refuse. **Hypothesis** as domain-general; bitrate is Empirical as the named envelope. Links: `mental-models.md` conformal; `formal-methods.md` help list. +**Named compaction envelope (Empirical as README behavior, 2026-09-18 +~16:22):** [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) +— the mutation monitor is **code** (mutating tools, unknown shell, +control operators / pipelines / substitutions / redirections → +`keep_full`). GLiNER2.5 may only propose a reduction on the remainder +or be more conservative; low-confidence / invalid evidence fail closed +to `keep_full`. Soft judgment inside a hard envelope, encoder backend +— not a Jev Score and not a summarizer (`notes.md` §50). Same sandwich +shape as bitrate-advisor; different family. + ## 13. DST multiverse triage (Hypothesis) **Method**: Antithesis / Resonate DST artifacts → failure taxonomy → @@ -665,6 +675,13 @@ same *family* on project instructions: the **linter proves** lintable rules; Jev Scores only residual soft AGENTS.md rules; fail-open, banded (`notes.md` §47). Different remainder from jevgate's unlisted verbs and from rh-guard's eval-integrity hole — do not merge products. +Compaction polarity is the other way: +[gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) +— code proves mutating / dangerous shell → `keep_full`; the encoder +judges only the remainder; uncertain **fails closed to `keep_full`** +(`notes.md` §50). Same sandwich, opposite fail policy from jevgate +(cannot block) and Abide (fail-open on diffs): the authorized act is +a destructive reduction of memory. **Beyond SWE (Hypothesis):** recipe book ∩ "does this leftover look done?"; labor-law allowlist ∩ hiring-fit Noul; SPF/DKIM pass ∩ phishing Noul on the body. **Counterexample:** diff --git a/.agents/skills/augustus/references/mental-models.md b/.agents/skills/augustus/references/mental-models.md index 82f7ea9..84fb16b 100644 --- a/.agents/skills/augustus/references/mental-models.md +++ b/.agents/skills/augustus/references/mental-models.md @@ -319,7 +319,10 @@ picks its next tool in a loop (`boundary-audit.md`); PufferLib Ocean scores as a capability claim (`formal-methods.md` DST trio). Mapping card for the cross-domain loop: `mappings.md` §9. Soft judgment inside a hard envelope: bitrate-advisor (ABR) and mmalisper's JOB -hybrid (Postgres plans first) — `notes.md` §44. +hybrid (Postgres plans first) — `notes.md` §44. Compaction envelope +(encoder, not Jev): gliner25-compaction — mutating tools / shell +operators prove `keep_full`; the model may only match that or be more +conservative (`notes.md` §50). ## Signal detection @@ -460,6 +463,7 @@ Use these as *existence proofs of a position*. Write your own card. | Hiring | interview / reject / hold | evidence Nouls; veto rules in policy | labor law, scorecards you wrote | | Inbox | reply / snooze / archive | urgency Noul + aboutness Choice | send, calendar | | Knowledge work | extract a quote / a cited fact | per-sentence or per-line-id Noul/Choice (**Empirical**: testimonial-miner, jev-reviewer) | verbatim join; place; human publish permission | +| Agent context | compact completed tool results without inventing prose | retention Choice + char-offset locate (**Empirical**: gliner25-compaction; same *job* as fast-jev-compaction / pi-jev-compaction) | mutation/shell envelope → keep_full; fail-closed keep_full; shadowMode before replace; copy exact bytes | | Dataframe labeling | classify / score rows | Noul/Choice/Score + full `p__` (**Empirical** as jevframe / jevpandas *shape*) | pandas/Polars, thresholds in code | | Computer-use speed | one verified act per step | operation + target Choice on numbered controls (**Empirical**: solari-reflex) | Guard check; deny-list absence; no screenshots | | Agent turn | skip memory tour on easy intent | intent Choice (**Empirical**: jev-hermes) | Memory still writes; complex still searches | diff --git a/.agents/skills/augustus/references/methods-catalog.md b/.agents/skills/augustus/references/methods-catalog.md index 6e970de..f1ff0b7 100644 --- a/.agents/skills/augustus/references/methods-catalog.md +++ b/.agents/skills/augustus/references/methods-catalog.md @@ -57,7 +57,7 @@ judgment component is new). | Method | Judgment-shaped component | Jev substitution | Stays in code | Status | |---|---|---|---|---| | Entropy as uncertainty signal | Measuring "how spread is this belief" | Entropy of returned distributions across repeats or options — computed in code from returned probabilities | All arithmetic | **Empirical recipe** (cookbook pattern) | -| Detector / Neyman filter (context) | Is this artifact relevant to the current task? | One relevance Noul per block before it enters context; stub + recall key | Cache, recall, safety keeps | **Empirical recipe** (winnow ≤0.22 hide; compaction 2-noul rule; pi-jev-context hide-not-delete) | +| Detector / Neyman filter (context) | Is this artifact relevant to the current task? | One relevance Noul per block before it enters context; stub + recall key. Encoder cousin: retention Choice + span locate, copy verbatim | Cache, recall, safety keeps; mutation/shell envelope | **Empirical recipe** (winnow ≤0.22 hide; compaction 2-noul rule; pi-jev-context hide-not-delete; gliner25-compaction GLiNER2.5 keep_full/keep_evidence/keep_call_only/drop, `notes.md` §50) | | Anomaly detection | Does this deviate from expected shape? | Guard nouls + harm Score over {input, output, tool trace} | Baselines, alert thresholds | **Empirical recipe** (guardrails cookbook; pi-jev output judge) | | Allowlist ∩ remainder (code-then-model) | Unlisted / unstructured leftovers after a **proof** | Typed questions only on the unknown tier; admit iff every p < τ | Proven/refused in code; cannot block unless a sandbox sits under | **Empirical recipe** (jevgate 0/59 unsafe unasked held-out; allowlist *proves* read-only verbs; doc-router 1.74× $). Domain-general: `mappings.md` §18 | | Decision-token LoRA (constrained-AR) | Specialize a generator for parallel constrained fields | Loss only on the single decision token; KV broadcast across fields | Schema, candidate tokens, policy | **Empirical recipe** as Foodoo1 200-case / 4-field receipt (fraud_risk 64→95%, overall 85.2→98.8%, ~234 ms); **Hypothesis** as a general recipe. Synthetic; not a financial product. Softmax ≠ Noul | @@ -68,7 +68,7 @@ judgment component is new). | Method | Judgment-shaped component | Jev substitution | Stays in code | Status | |---|---|---|---|---| | Claim–evidence entailment (NLI) | supports / contradicts / not-established per claim–source pair | One Choice per pair + review flag; judge against the cited source text only | Quote extraction, citation graph, audit log. When the answer *is* a span you already hold, **point** at line ids and copy verbatim — the model never writes the excerpt (jev-reviewer) | **Empirical recipe** (citation_check cookbook; jev-reviewer sample study, `notes.md` §48). Atlas: paraphrase_support and reversed_meaning_high_overlap correctly judged when both texts are in `state` (`notes.md` §49) | -| Extractive selection + offline re-threshold | Keep/drop over sentences or ids code already numbered | Per-item Noul/Choice + one broadcast; join in order; `redecide` on the log with no new calls | Numbering, header skip, thresholds, publish permission | **Empirical recipe** (testimonial-miner 8-request fixture; jev-reviewer; `notes.md` §48). Cousin of applied-mappings §2 | +| Extractive selection + offline re-threshold | Keep/drop over sentences, ids, or character offsets code already holds | Per-item Noul/Choice + one broadcast; join in order; `redecide` on the log with no new calls. Compaction: retention Choice + span locate | Numbering, header skip, thresholds, publish permission; mutation envelope; copy exact bytes | **Empirical recipe** (testimonial-miner 8-request fixture; jev-reviewer; gliner25-compaction char-offset copies, `notes.md` §48, §50). Cousin of applied-mappings §2 | | Combinatorial grid / program synthesis | Consistent whole-object from many cells | **Rejected as extractive.** Cell-wise Choice does not assemble ARC grids (4/400 Direct Jev) | Search, a program, a simulator | **Empirical as a negative** (`notes.md` §49) | | Spec vs artifact conformance (model checking *mindset*) | Property holds / violated / unverifiable for a named requirement | One Noul/Score per requirement, batched; violated → named rule back into context (pi-warden / Abide shape). This is **not** TLC/Apalache/GNATprove | Requirement enumeration, enforcement, logging; the **linter** if the rule is lintable; the real checker if you have one | **Empirical recipe** (pi-warden: 6→0 rule breaks, 150 paired runs; jev-pref: YOU define the rule; Abide: productized compile/calibrate/tune/replay, `notes.md` §47). Ownership split: `formal-methods.md` | | AST ∩ semantic lint | Semantic remainder after a parser already extracted units | Typed questions on Tree-sitter targets; do not execute scanned code | Parser, selection, fail-on; `tenbin` owns the lint *skill* | **Empirical as a shape** (jevscan 0.2.0rc4; not a calibration claim; `notes.md` §48) | diff --git a/.agents/skills/augustus/references/mixed-architecture.md b/.agents/skills/augustus/references/mixed-architecture.md index 0ed0f25..02aaf43 100644 --- a/.agents/skills/augustus/references/mixed-architecture.md +++ b/.agents/skills/augustus/references/mixed-architecture.md @@ -178,6 +178,7 @@ not a global virtue: | Action | Typical failure policy | Why | |---|---|---| | Drop a RAG chunk or log line | **Fail open** (keep on error) | A false drop loses evidence; a false keep costs tokens | +| Compact / drop a completed tool result | **Fail closed** to keep-full (`gliner25-compaction`) | Compaction is a destructive edit of memory. Uncertain *looks* like keep-on-error from the evidence side; name the *reduction* as the act. Contrast Abide / jevgate fail-open | | Route to a tool / start a side effect | **Fail closed** (don't call) | A wrong tool is an action | | Rerank a retrieved list | Fail open: keep retrieval order (`WiktorB2004/llama-index-jev`, **Empirical recipe** on BEIR nfcorpus: MiniLM 0.340 nDCG@5 → MiniLM+Jev 0.396; rerank fails open, *select* fails closed). Listwise/cross-encoder scores belong here, not on the row above. | Ranking errors are quality; selection errors are control-flow | @@ -189,6 +190,10 @@ Worked placements (2026-09-18 topic:jev hour + prior archive): relevance against the task; last-N lines and error signatures kept in *code* before Jev sees anything; full output recoverable by id. Same shape as winnow/fast-jev-compaction (`agent-self-assessment.md`). + Encoder-backend cousin: + [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) + — GLiNER2.5 retention Choice + exact spans; fail closed to + `keep_full`; `shadowMode` default true (`notes.md` §50). Not Jev. - **Diff hunks before `git add`** — `ibrahemid/git-jev-stage`: one Choice per hunk (`include` / `exclude` / `mixed`); mixed and low-confidence stay unstaged; lines never split; staging is an exact patch after confirm. @@ -373,7 +378,11 @@ Related placements: the *practice* is first-class: eval CLI asserts on the **action**, not on prose; recipes span alerts / RTB / sports-bet / prediction markets (`notes.md` §44). LLM-as-judge is not the primary System One - score (`faq.md`). Do not copy the client. + score (`faq.md`). Compaction rollout of the same instinct: + [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) + public default `shadowMode: true` (analyze + log; do not replace + history until explicitly enabled) (`notes.md` §50). Do not copy the + client. - **Hybrid countable + judgment rules** — `DanRWilloughby/snifftest`: deterministic tells score 1.00; judgment rules flag only outside the unsure band. Explicit: a reading near 0.5 is *no judgment*, never a pass. @@ -408,7 +417,7 @@ decision-design card. Do not clone APIs from READMEs. | Hold-before-publish moderation | Hazard Nouls + harm Score | Block/review/pass policy | Near Here / firehose family | | Tool / engine / skill select | Choice + fits-Noul | Dispatch, auth, reject-all | skillranker, LlamaIndex selectors, Toolrouter | | Preference lint | Per-rule Score/Noul on a diff | Rule text, linter for hard rules, bands + fail-open | jev-pref (contract), Abide (productized), JevLint | -| Context / log prune | Per-line or per-block relevance | Always-keep set, recall keys | jevprune, winnow | +| Context / log prune | Per-line or per-block relevance; or a retention Choice + spans | Always-keep set, recall keys; mutation envelope in code; shadow before replace | jevprune, winnow; fast-jev-compaction / pi-jev-compaction (Jev); gliner25-compaction (GLiNER2.5) | | Exact hunk staging | Per-hunk include/exclude/mixed | `git diff`, atomic apply | git-jev-stage | | Semantic `WHERE` | Noul/`jev_prob` over a row | SQL, indexes, LIMIT | jevql (CLI; DB sees ordinary SQL); sqlite-jev (in-engine extension) | | Formula / query embedding | JUDGE as a function | Spreadsheet/SQL engine | judge-sheets, jevql, sqlite-jev | @@ -424,7 +433,7 @@ decision-design card. Do not clone APIs from READMEs. | S1 reflex + optional S2 advice | Typed action Choice; planner one-use on low p | Collision, legality, the stick stays with S1 | jev-reflex-autonomy-lab (experimental) | | Decision-as-business-tool | Named judgment; gate is part of the result | Registry, arithmetic, hard guards | jev-decision-layer (unofficial) | | NL cases → checked e2e | Jev selects observed controls | Playwright expectations; PASS/FAIL/BLOCKED | jev-e2e (alpha) | -| Extractive quotes / pointer evidence | Per-sentence or per-line-id Noul/Choice | Verbatim join; place; `redecide` / CSV; model never writes the excerpt | testimonial-miner; jev-reviewer | +| Extractive quotes / pointer evidence | Per-sentence, per-line-id, or char-offset Noul/Choice | Verbatim join; place; `redecide` / CSV; model never writes the excerpt | testimonial-miner; jev-reviewer; gliner25-compaction | | Structured observe → decide → act | Operation + target Choice on numbered controls | Guard check; deny-list absence; no screenshots; TYPE is the only generation | solari-reflex | | Dataframe semantic columns | Noul / Choice / Score per row; full `p__` | pandas/Polars, indexes, never silent renormalize | jevpandas; jevframe (PyPI + Polars) | | Route ≠ memory | Intent Choice before a turn | Config + flat tools on easy routes; memory stays on for hard ones | jev-hermes | @@ -474,7 +483,8 @@ not the class monopoly. **Local contract drop-in this hour:** (`notes.md` §48). **ONNX replica path:** [`Mattepiu/laya-onnx`](https://huggingface.co/Mattepiu/laya-onnx) — do not copy the inherited vs-Jev table. GLiNER (locate) / GLiClass (categorize) / -GLiNER2.5 (local multi-head), listwise, and vision families: +GLiNER2.5 (local multi-head; extractive compaction is a named *job* on +that family, `notes.md` §50), listwise, and vision families: `judgment-class.md`. ## Design-card extras for mixed systems diff --git a/.agents/skills/augustus/references/toolbox-mapping.md b/.agents/skills/augustus/references/toolbox-mapping.md index 042bbc1..5c45774 100644 --- a/.agents/skills/augustus/references/toolbox-mapping.md +++ b/.agents/skills/augustus/references/toolbox-mapping.md @@ -77,7 +77,7 @@ component; keep the rest of the method in code. | Discrete math: width vs depth | Fan out in width (parallel ≈ free), pay depth linearly; two-stage only when next options depend on an earlier answer | **Empirical recipe** (fan-out: 12.2× cheaper, 10× faster) | | Psychology: Kahneman | System 2 generates/proposes (LLM), System 1 discriminates (Jev); never the reverse. S1 keeps control; optional S2 is one-use advice and does not fly. Productized cascade: `conf ≥ τ` → S1 decides else S2 writes; routing fails open / safety fails closed; **routing accuracy unmeasured**; keyword fallback ≠ S1 | **Empirical recipe** (mcts-agent role split; 2026-09-18 mixed-architecture discourse; jev-reflex-autonomy-lab as a control-loop shape, `notes.md` §46; [dual-process-ai](https://github.com/taro1985/dual-process-ai) as a business/life cascade, `notes.md` §49). **Route ≠ memory:** a cheap intent gate skips memory/tool *tours* on easy routes; memory still writes; complex still searches (jev-hermes, `notes.md` §48) | | IR / cascades | Cheap relevance / irrelevance before an expensive ranker or generator | **Empirical recipe** (RAG cookbook; jevprune; git-jev-stage; LlamaIndex Jev rerank) | -| IR / extractive keep-drop | Number candidates in code; model selects; copy verbatim; `redecide` thresholds on the log with no new calls | **Empirical recipe** (testimonial-miner; jev-reviewer pointer-not-generator; `notes.md` §48) | +| IR / extractive keep-drop | Number candidates in code; model selects; copy verbatim; `redecide` thresholds on the log with no new calls. Compaction: same pointer job on tool results (Jev Noul/Score *or* GLiNER encoder) | **Empirical recipe** (testimonial-miner; jev-reviewer pointer-not-generator, `notes.md` §48; gliner25-compaction char-offset + fail-closed keep_full, `notes.md` §50) | | Spec / lint | Project-defined semantic rules as predicates over a diff; linter owns hard rules. AST remainder: Tree-sitter units, then typed questions; do not execute scanned code | **Empirical recipe** (jev-pref contract; Abide productized path — replay 93 sessions, edit precision ~26% / turn ~73% before tune, `notes.md` §47; JevLint file-level Noul; pi-warden; snifftest unsure-band; jevscan AST∩semantic, `tenbin` owns the lint skill, `notes.md` §48) | | Formal methods / DST / safety | Judgment triages counterexamples, failing seeds, and named-rule conformance; proof/MC/DST stay with their tools. Alloy finder ≠ Apalache BMC ≠ Quint run. DST trio: Antithesis hypervisor / Resonate HQ Lean+oracle+SDK (durable async) / PufferLib env+seed. Noul is a sensor, not a discharged PO. Semi-formal diagrams are vocabularies, not enforcers | **Hypothesis as product**, **Contract** as ownership (matching `mappings.md` §8 and `methods-catalog.md`; worked shape pi-warden — `formal-methods.md`, `formal-semi-formal.md`) | | Decision analysis: VOI | Gather as an enumerated act; pay iff expected decision-loss drop > cost | **Hypothesis** as calculator (`mappings.md` §6; `mental-models.md`) | diff --git a/CHANGELOG.md b/CHANGELOG.md index fcc0eb3..35aaa48 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -199,6 +199,17 @@ Pegged against [`typesafe-ai/skills` v0.5.7](https://github.com/typesafe-ai/skil surface ([von](https://github.com/wfzyx/von) 14 MB; not a replica). kev light delta **100★**. Do not merge Banking77 87% / 76.3% / 79.67%. No wrapper. +- GLiNER2.5 extractive compaction (`research/notes.md` §50, + [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction), + Apache-2.0): architecture notes, not a plugin how-to. Pointer + keep-drop (character-offset copies) vs generator summarizers; + family with testimonial-miner / jev-reviewer. Soft retention Choice + under a hard mutation envelope (mutating tools / shell operators → + `keep_full`); low-confidence / invalid evidence fail closed to + `keep_full` — contrast many fail-open Jev gates. Same compaction + *job* as fast-jev-compaction / pi-jev-compaction; GLiNER encoder + backend; Fastino/GLiGuard sibling class. `shadowMode` default true. + Not Jev. Not multimodal. No invented metrics. - Effect-oriented loops (`notes.md` §28, `mappings.md` §19): Ward's ZIO client keeps Jev as the outer Choice and the handler as the effect. Not Effect.ts. GLiNER author: GLiNER2 "like jev" is GLiGuard diff --git a/README.md b/README.md index d65adee..f4e4283 100644 --- a/README.md +++ b/README.md @@ -34,7 +34,8 @@ never launder a Noul as a proof. distill), constrained-AR (TypeAR, pcdServer), announced decision-model (Watch), open multimodal RLCD (blackwood-rlcd; not Archer), Laya ONNX port, contract-compatible local `/v1/systemone` (stub until hf scorer; also kev pointer / von tiny SAN — not replicas), - GLiNER/GLiClass species (locate vs categorize vs local multi-head), + GLiNER/GLiClass species (locate vs categorize vs local multi-head; + GLiNER2.5 extractive compaction as a named job, not a new species), listwise vs decision objectives, vision scoring, when-to-use axes (including decision-model vs constrained LLM), agent-architecture portents @@ -50,12 +51,12 @@ never launder a Noul as a proof. - `.agents/skills/augustus/references/mixed-architecture.md` — default placement: judgment-class model + LLM + code; preference lint; provider (Jev default / other family with self-eval); dual-process S1 decide / S2 - generate; component node; DOM-as-text + fan-out + generate; component node; DOM-as-text + fan-out; shadow-mode compaction rollout - `.agents/skills/augustus/references/applied-mappings.md` — context sieve, - exact-text keep/drop (extractive / pointer-not-generator), env triage, moderation/ranking, skill routing (route ≠ memory) + exact-text keep/drop (extractive / pointer-not-generator; char-offset compaction), env triage, moderation/ranking, skill routing (route ≠ memory) - `.agents/skills/augustus/references/faq.md` — "just classification", stack replacement, Jev vs open head vs encoder vs LoRA vs constrained AR vs kev vs blackwood, - wait-for-Archer, missing-other confident-wrong, soft project rules vs linter (Abide), extractive/pointer-not-generator, local drop-in vs stub scorer, route ≠ memory, when-it-holds / extractable-from-state, decision-model vs constrained LLM, dual-process S1/S2, combinatorial grid ≠ extractive, GLiNER vs GLiClass vs CLIP, LLM-as-judge, in-engine vs CLI store, + wait-for-Archer, missing-other confident-wrong, soft project rules vs linter (Abide), extractive/pointer-not-generator, compaction summarize vs pointer, encoder vs Jev compaction, fail-closed keep_full, shadow-mode rollout, local drop-in vs stub scorer, route ≠ memory, when-it-holds / extractable-from-state, decision-model vs constrained LLM, dual-process S1/S2, combinatorial grid ≠ extractive, GLiNER vs GLiClass vs CLIP, LLM-as-judge, in-engine vs CLI store, hard envelope (bitrate / planner), not-another-how-to - `.agents/skills/augustus/references/mappings.md` — classical-method mappings with boundaries, counterexamples, acceptance tests (including diff --git a/docs/ecosystem.md b/docs/ecosystem.md index 93bcc53..dd6719b 100644 --- a/docs/ecosystem.md +++ b/docs/ecosystem.md @@ -50,7 +50,7 @@ Per-keystroke launchers (104ms median, sequence-tagged staleness), firehose mode - **jeiel85/jevscope** — local-first visual debugger + JSONL regression for Choice/Score/Noul; compare two definitions; policy buckets are JevScope-derived. Sits next to jevals. Pointer: `research/notes.md` §25. ### Local / open heads & GLi\* species -- **GLiNER / GLiNER2.5 / GLiClass** — species map: locate spans vs categorize the sequence vs local multi-head (fastino-ai GLiNER2.5 CPU-first). Peer of Jev, not a footnote. `references/judgment-class.md`. Author primary source: GLiNER2 "like jev" is schema-conditioned categorize (GLiGuard), not a Noul (`notes.md` §28). 36× Browser Use claim is a tweet (`notes.md` §25). +- **GLiNER / GLiNER2.5 / GLiClass** — species map: locate spans vs categorize the sequence vs local multi-head (fastino-ai GLiNER2.5 CPU-first). Peer of Jev, not a footnote. `references/judgment-class.md`. Author primary source: GLiNER2 "like jev" is schema-conditioned categorize (GLiGuard), not a Noul (`notes.md` §28). 36× Browser Use claim is a tweet (`notes.md` §25). Named job: [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) — extractive retention Choice + char-offset copies; not a summarizer; not Jev (`notes.md` §50). - **GLiGuard** (fastino-ai) — 0.3B GLiNER2 encoder, checkpoint `fastino/gliguard-LLMGuardrails-300M`. One bidirectional pass over a safety schema. Same interface shape as batched questions; different objective. Not a Jev weight clone. `judgment-class.md`; `notes.md` §30. - **DECRUX9812/openjev-lm** — Qwen2.5-0.5B+LoRA distilled from hosted Jev answers; 65/70 = 92.9% on 70 hand-labelled rows (one annotator, one domain, one seed) overnight on 6 vCPU, $0/call. Its 98.1% on fresh rows is teacher *agreement*, not gold. Receipts pattern: `notes.md` §25, §44. - **convaiinnovations/laya** — open Choice/Score/Noul head, text-only, 512 tok. Companion packaging this hour: [`laya-typed-decisions`](https://huggingface.co/convaiinnovations/laya-typed-decisions) (421.3M, acc 0.766 / Brier 0.066 unverified). Shared bake-off: [`pngwn/open-jev-laya-bench`](https://huggingface.co/datasets/pngwn/open-jev-laya-bench) (26+9 tasks, 11959 items; ECE/NLL/Brier; not TypeSafe Jev vs Laya). ONNX replica: [`Mattepiu/laya-onnx`](https://huggingface.co/Mattepiu/laya-onnx) (~15 ms CPU; do not copy vs-Jev table). `notes.md` §18, §42, §46, §48. @@ -77,7 +77,7 @@ Per-keystroke launchers (104ms median, sequence-tagged staleness), firehose mode ### Agent harnesses extras (this hour) - **kevinpita/pi-jev-context** — reversible Pi context sieve: hide, do not delete; `/jev off` restores. Cousin of winnow/jevprune. -- **vava-nessa/pi-jev-compaction** (and `tamaratran/fast-jev-compaction`) — verbatim drop, never summarize. Pair with jev-gate-student-b for local memory-gating. +- **vava-nessa/pi-jev-compaction** (and `tamaratran/fast-jev-compaction`) — verbatim drop, never summarize. Pair with jev-gate-student-b for local memory-gating. Same *job* as [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) (GLiNER2.5 encoder backend; `notes.md` §50). - **reachjalil/jev-tree** — authored taxonomy so each Choice stays under 255; truncate silently drops the tail (`jev-tree-choice-cap`). - **HacksonClark / SREGym-Lite** — Jev ranks next tests/evidence; does not diagnose; 20/50→24/50 with 2 regressions. `notes.md` §33. - **ddfeyes/jev-mode** — bulk triage/tag/route off frontier context; synthetic 1,000: −77.8% tokens; accuracy is parity. `notes.md` §42. @@ -153,6 +153,12 @@ Patterns, not a catalog. `notes.md` §49. TypeSafe Jev is the exemplar, not a mo - **wfzyx/von** — 14 MB SAN local drop-in. Distinguishes from jev-local stub / kev pointer. Do not copy vs-Jev table. - **jaredpalmer/kev** — light delta: **100★** this pass. No species rewrite. +### Hourly ~16:22 Boise (GLiNER2.5 extractive compaction) + +Architecture notes, not a plugin catalog. `notes.md` §50. TypeSafe Jev is the exemplar, not a monopoly. **Not Jev. Not multimodal.** Archer still Watch. + +- **m-newhauser/gliner25-compaction** — local GLiNER2.5 (`fastino/gliner2.5-base-v1`) retention Choice (`keep_full` / `keep_evidence` / `keep_call_only` / `drop`) + exact character-offset copies. Pointer family with testimonial-miner / jev-reviewer. Mutating tools / shell operators → `keep_full` in code. Fail-closed to `keep_full` (contrast many fail-open Jev gates). Same compaction *job* as fast-jev-compaction / pi-jev-compaction; encoder backend; Fastino/GLiGuard sibling class. `shadowMode` default true. Experimental; characters not tokens; no published retention-quality rates. Apache-2.0. + See `references/mixed-architecture.md` in the skill. Class-level family choice: `references/judgment-class.md`. Proof vs judgment (Alloy vs Apalache; DST trio Antithesis / Resonate HQ / PufferLib): diff --git a/research/archive/findings.md b/research/archive/findings.md index 931ca31..784772e 100644 --- a/research/archive/findings.md +++ b/research/archive/findings.md @@ -884,6 +884,39 @@ routing accuracy; (bs) combinatorial assembly is not extractive keep/drop; (bt) local `/v1/systemone` has three surfaces (stub / pointer / tiny SAN). +## Batch #34 (2026-09-18 ~16:22 Boise) — GLiNER2.5 extractive compaction + +Note: `research/notes.md` §50. Docs-only. Folded into PR #2. Not a +second PR. Archer still Watch. **Not Jev. Not multimodal.** No +invented metrics. Do not re-fold §48 extractive recipes, §49 +bake-off, Abide, kev, GLiGuard species. + +- **m-newhauser/gliner25-compaction (Empirical as README + behavior).** Apache-2.0. Created 2026-09-18T17:22:34Z; 1★ at + capture. Local GLiNER2.5 (`fastino/gliner2.5-base-v1`) chooses + `keep_full` | `keep_evidence` | `keep_call_only` | `drop` for + completed tool pairs and copies **exact character-offset** spans. + Not a prose summarizer. Mutating tools / unknown shell / control + operators → `keep_full`. Low-confidence / invalid evidence fail + closed to `keep_full`. `shadowMode` default true (log, do not + replace history). Reduction in characters, not tokens. No + published retention-quality rates. +- **Mental models:** (1) pointer/extractive vs generator summarizers + (family with testimonial-miner / jev-reviewer); (2) soft retention + Choice under a hard mutation envelope — fail-closed contrast vs + many fail-open Jev gates; (3) same compaction *job* as + fast-jev-compaction / pi-jev-compaction, GLiNER encoder backend, + Fastino/GLiGuard sibling class; (4) shadow mode as safe rollout. +- **rh-guard:** sibling note only (fail-closed retention, hard shell + mutation, shadow). Not reward-hack detection. + +Cross-repo addition: (bu) compaction that writes prose is a +different species from pointer keep/drop; (bv) fail-closed +`keep_full` names the *reduction* as the irreversible act; (bw) +compaction job is backend-agnostic (Jev Score/Noul vs GLiNER +encoder); (bx) shadow-mode default is the rollout for memory +mutation. + diff --git a/research/notes.md b/research/notes.md index 5381ddc..20b1da8 100644 --- a/research/notes.md +++ b/research/notes.md @@ -3270,3 +3270,136 @@ open-alternative-jev); `mixed-architecture.md` (dual-process; component node; DOM-as-text); `faq.md`; `mappings.md` §2 / §3 / §6 / §9; `applied-mappings.md`; `question-design.md`; `methods-catalog.md`; `toolbox-mapping.md`; `agent-self-assessment.md`. No wrapper. + +## 50. GLiNER2.5 extractive compaction — encoder backend, same keep/drop job (2026-09-18 ~16:22 Boise) + +America/Boise ~16:22 = 22:22 UTC. Docs-only fold into open PR #2 +(`cursor/augustus-store-envelope-00b4`). Not a competing PR. Archer +27B drop still **WATCH**. Identity lock vs `typesafe-ai` / `tenbin` / +`decision-first` holds. No wrapper, no `--plugin-dir` / `uv` how-to, +no copied timeouts or Hub download scripts. No invented metrics. +**Not Jev. Not multimodal.** Do not re-fold §48 extractive recipes, +§49 bake-off, Abide, kev, GLiGuard as a species rewrite, or +pi-jev-compaction as a new product. + +Backend-agnostic: this is a **compaction / context-sieve placement** +(pointer keep-drop, soft Choice under a hard mutation envelope, +shadow-mode rollout). TypeSafe Jev is one backend for that *job* +(`tamaratran/fast-jev-compaction`, `vava-nessa/pi-jev-compaction`); +GLiNER2.5 is another. Augustus stays family-first. + +### HIGH + +1. **[`m-newhauser/gliner25-compaction`](https://github.com/m-newhauser/gliner25-compaction)** + (Apache-2.0; Python + Claude Code plugin; created + 2026-09-18T17:22:34Z; 1★ at capture). Local, **evidence-first** + context compaction. Default checkpoint + [`fastino/gliner2.5-base-v1`](https://huggingface.co/fastino/gliner2.5-base-v1) + (checkpoint model card Apache-2.0 per their README) chooses a + retention action for **completed** eligible tool interactions and + extracts **exact source spans** when the full result is unnecessary. + **Not a prose summarizer.** User and assistant text unchanged. + Retained evidence is copied from original **character offsets**. + Mutating tool interactions are preserved in full. Inference runs in + a local Python worker after Hub download; transcript analysis does + not require a remote inference API. + + Per completed pair, one retention action (closed set): + + - `keep_full` — complete call and result + - `keep_evidence` — exact excerpts from the result + - `keep_call_only` — call stays; result replaced with a rerun notice + - `drop` — paired call and result removed + + Inputs: current goal, nearby conversation, tool name and input, + original tool result. Deterministic safeguards override uncertain + predictions, protect validated diagnostic spans, preserve recent + messages and mutations, validate exact offsets, and reject orphaned + tool results. + + **Four load-bearing mental models (architecture, not a plugin + catalog):** + + 1. **Pointer / extractive, not generator.** Same family as + [testimonial-miner](https://github.com/AppitStudio/testimonial-miner) + and [jev-reviewer](https://github.com/choxos/jev-reviewer) + (`notes.md` §48): the model **selects**; code **copies + verbatim**. Compaction that invents a prose summary is a + different (worse) species for auditability. GLiNER locate is not + a footnote here — character offsets *are* the keep/drop + candidates. One local multi-head does **categorize** (which + retention action) and **locate** (which spans) in the same job + (`judgment-class.md` species map). + + 2. **Soft retention Choice under a hard envelope.** Code owns the + mutation monitor: mutating tools, unknown shell, and shell + control operators / pipelines / substitutions / redirections are + treated as mutating → `keep_full`. Missing or invalid evidence + **fails closed to `keep_full`**. Low-confidence retention + predictions likewise fail closed to `keep_full`. Contrast: many + Jev *preference / remainder* gates fail-open (Abide; jevgate + cannot block). Compaction *drop* (and lossy `keep_evidence`) is + the irreversible act, so the authorized reduction fails closed. + From the evidence-preservation view the outcome looks like + context-sieve "keep on error" (`applied-mappings.md` §1) — name + the *act*, not the slogan. Conservative shell over-retention is + their documented limit, not a bug to "fix" by failing open. + + 3. **Same compaction job, encoder backend.** + [`tamaratran/fast-jev-compaction`](https://github.com/tamaratran/fast-jev-compaction) + asks two Nouls (should the *call* stay? should the *result* stay + verbatim?). [`vava-nessa/pi-jev-compaction`](https://github.com/vava-nessa/pi-jev-compaction) + is the Pi cousin: verbatim drop, never summarize. Here GLiNER2.5 + chooses discrete retention actions + evidence spans. Fastino / + GLiGuard sibling *class* (schema-in-encoder, local) — not a + GLiGuard safety-schema clone, not a Jev Score, not a Noul. + Augustus does not pick a vendor for the hole. + + 4. **Shadow mode as safe rollout.** Public default `shadowMode: + true`: local analysis logs the proposed reduction **without + replacing session history** until explicitly set false. Same + rollout instinct as jev-harness / is-malicious (log would-do + first). Compaction mutates memory; shadow is the default because + a bad drop is not a reversible token cost. + + **Limits (theirs, README; experimental).** Reduction measured in + **characters, not tokens**. Conservative shell policy may retain + commands that are actually read-only. Only completed + tool-call/result pairs are candidates. Domain-specific tuning and + broad production evaluation remain future work. No published + token-reduction or retention-quality rates this pass — do not + invent them. Their config defaults (`minimumConfidence` 0.7, + `minimumEvidenceConfidence` 0.5, `minReductionRatio` 0.25, + `preserveRecentMessages` 6) are *their* knobs, not class constants. + + **Siblings — complementary, do not merge.** + + - **`24601/rh-guard`:** reward-hack / eval-integrity on tool use. + Shared notes only: fail-closed retention, hard shell mutation + policy, shadow-mode rollout. Different hole. This is not + reward-hack detection. + - **GLiGuard:** Fastino encoder sibling (safety-schema classify). + Compaction is locate+categorize on tool transcripts, not LLM I/O + moderation. + - **Abide:** same Claude Code hook-host surface; Abide is fail-open + Score on diffs; this is fail-closed `keep_full` on compaction. + + **Placement.** Context sieve + exact-text keep/drop + (`applied-mappings.md` §1–§2). Pillar: selective classification / + SDT criterion (false drop >> false keep) + runtime-assurance + sandwich (mutation monitor in code). Hole: sieve / keep-drop. + Family: GLi\* encoder (GLiNER2.5 local multi-head). Fail-closed on + the reduction. Eval path: none published this pass (experimental); + characters≠tokens is the honesty constraint. **Empirical** as + README behavior. **Hypothesis** that the same envelope transfers to + *your* transcript domain. Cards: `judgment-class.md` (primary); + `applied-mappings.md` §1–§2; `mappings.md` §12 / §18; + `mixed-architecture.md`; `faq.md`; `mental-models.md`; + `methods-catalog.md`; `toolbox-mapping.md`; + `agent-self-assessment.md`. No wrapper. + +### Omni / Jev-omni + +Not multimodal. Usage: extractive context management as a +perception/memory hygiene stage **before** decide. Archive + +landscape pointer. diff --git a/research/refresh-log.md b/research/refresh-log.md index c9569f6..9124a09 100644 --- a/research/refresh-log.md +++ b/research/refresh-log.md @@ -488,6 +488,28 @@ toolbox, agent-self-assessment, ecosystem, CHANGELOG, README. - notes.md §49; sources.json; findings.md batch #33. No wrapper. +## 2026-09-18 22:22 UTC — GLiNER2.5 extractive compaction (~16:22 Boise) + +- Folded into open PR #2 (`cursor/augustus-store-envelope-00b4`). + Docs-only. Not a competing PR. Archer 27B drop still **WATCH**. + **Not Jev. Not multimodal.** No invented metrics. +- HIGH: [m-newhauser/gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) + (Apache-2.0, created 2026-09-18). GLiNER2.5 + `fastino/gliner2.5-base-v1` retention Choice + exact char-offset + copies. Mutating tools / shell operators → `keep_full`. Fail-closed + to `keep_full`. `shadowMode` default true. Same *job* as + fast-jev-compaction / pi-jev-compaction; encoder backend. + Fastino/GLiGuard sibling class. Pointer family with + testimonial-miner / jev-reviewer. +- Limits (theirs): experimental; characters not tokens; conservative + shell over-retains; completed pairs only. No published + retention-quality rates. +- Cards: SKILL.md, judgment-class (primary), applied-mappings §1–§2, + mappings §12/§18, mixed-architecture, faq, mental-models, + methods-catalog, toolbox, agent-self-assessment, ecosystem, + CHANGELOG, README. +- notes.md §50; sources.json; findings.md batch #34. No wrapper. + diff --git a/research/sources.json b/research/sources.json index 75bc086..9c5e553 100644 --- a/research/sources.json +++ b/research/sources.json @@ -1,6 +1,6 @@ { "refresh_cadence": "hourly", - "retrieved": "2026-09-18T21:52Z", + "retrieved": "2026-09-18T22:22Z", "sources": [ { "kind": "docs", @@ -654,7 +654,7 @@ "kind": "docs", "title": "fastino/gliner2.5-base-v1", "url": "https://huggingface.co/fastino/gliner2.5-base-v1", - "note": "English 194M boundary checkpoint. Do not copy AutoExtractor into Augustus." + "note": "English 194M boundary checkpoint. Default for m-newhauser/gliner25-compaction (notes.md \u00a750). Do not copy AutoExtractor into Augustus." }, { "kind": "repo", @@ -912,7 +912,7 @@ "kind": "repo", "title": "vava-nessa/pi-jev-compaction", "url": "https://github.com/vava-nessa/pi-jev-compaction", - "note": "Pi cousin of fast-jev-compaction: verbatim drop, never summarize. Pointer family, not a how-to." + "note": "Pi cousin of fast-jev-compaction: verbatim drop, never summarize. Pointer family, not a how-to. Same *job* as gliner25-compaction (GLiNER encoder backend); notes.md \u00a750." }, { "kind": "social", @@ -1477,6 +1477,12 @@ "title": "jaredpalmer/kev (star delta)", "url": "https://github.com/jaredpalmer/kev", "note": "100 stars this pass (user cited 97). Light activity delta only. notes.md \u00a749." + }, + { + "kind": "github", + "title": "m-newhauser/gliner25-compaction", + "url": "https://github.com/m-newhauser/gliner25-compaction", + "note": "Apache-2.0. Local GLiNER2.5 (fastino/gliner2.5-base-v1) extractive compaction for Claude Code. Retention Choice keep_full|keep_evidence|keep_call_only|drop; exact char-offset copies; mutating tools keep_full; fail-closed to keep_full; shadowMode default true. Not Jev, not a summarizer, not multimodal. notes.md \u00a750." } ] } From d220485362e9e381c329ad2f27b32beab3de7017 Mon Sep 17 00:00:00 2001 From: Cursor Agent Date: Fri, 18 Sep 2026 23:03:34 +0000 Subject: [PATCH 09/43] Fold 16:48 Boise watch: CI merge-gate, fail-open wake, Harbor on/off Docs-only architecture notes for latch (cluster-then-policy PASS/BLOCK), wakegate (fail-open VOI resume), GLiNER S1 indexer (10-50x unfilled), clear-head claim/evidence Stop, Harbor on/off routing, and cousins. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com> --- .agents/skills/augustus/SKILL.md | 16 +- .../references/agent-self-assessment.md | 30 +- .../augustus/references/applied-mappings.md | 23 +- .../references/composition-algebra.md | 3 + .agents/skills/augustus/references/faq.md | 30 +- .../augustus/references/judgment-class.md | 12 + .../skills/augustus/references/mappings.md | 35 +- .../augustus/references/mental-models.md | 4 + .../augustus/references/methods-catalog.md | 9 +- .../augustus/references/mixed-architecture.md | 26 +- .../augustus/references/toolbox-mapping.md | 7 +- .../skills/augustus/references/validation.md | 17 + CHANGELOG.md | 16 + README.md | 13 +- docs/ecosystem.md | 14 + research/archive/findings.md | 50 +++ research/notes.md | 306 ++++++++++++++++++ research/refresh-log.md | 28 ++ research/sources.json | 92 +++++- 19 files changed, 699 insertions(+), 32 deletions(-) diff --git a/.agents/skills/augustus/SKILL.md b/.agents/skills/augustus/SKILL.md index 94e187f..a086110 100644 --- a/.agents/skills/augustus/SKILL.md +++ b/.agents/skills/augustus/SKILL.md @@ -1,6 +1,6 @@ --- name: augustus -description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill, open multimodal RLCD / blackwood), announced open decision-model (Watch — still not landed), constrained-AR (TypeAR, pcdServer, decision-token LoRA), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"wait for Archer vs open multimodal RLCD\", \"screenshot/DOM candidates → typed Choice\", \"eval path\", \"jevals\", \"Harbor taskset\", \"shared bake-off ECE/NLL/Brier\", \"LLM-as-judge is not the System One score\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge (allowlist *proves*; fail-open cannot block), \"missing other → confident wrong Choice\", \"lint the Jev request\", \"training confronts Choice other / none-of-the-above\", \"S1 reflex keeps control / optional S2 one-use advice\", \"soft AGENTS.md rules vs linter\" (Abide / jev-pref), \"edit-phase vs turn-phase observation window\", banded confidence fail-open preference lint, \"extractive selection / pointer-not-generator\", \"encoder GLiNER compaction vs Jev Score compaction (same job; pointer not summarizer)\", \"fail-closed keep_full under mutation envelope\", \"shadow-mode compaction rollout\", \"local /v1/systemone drop-in (stub until hf scorer)\", \"dataframe-native semantic columns\", \"route≠memory\", \"advisory sidecar receipts\", \"structure induction over bags\", \"AST ∩ semantic lint\", \"extractable-from-state / retrieve first\", \"decision-model vs constrained-LLM bake-off\", \"dual-process S1 decide / S2 generate\", \"combinatorial grid ≠ extractive\", TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, or \"formally verify with Jev\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." +description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill, open multimodal RLCD / blackwood), announced open decision-model (Watch — still not landed), constrained-AR (TypeAR, pcdServer, decision-token LoRA), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"wait for Archer vs open multimodal RLCD\", \"screenshot/DOM candidates → typed Choice\", \"eval path\", \"jevals\", \"Harbor taskset\", \"shared bake-off ECE/NLL/Brier\", \"LLM-as-judge is not the System One score\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge (allowlist *proves*; fail-open cannot block), \"missing other → confident wrong Choice\", \"lint the Jev request\", \"training confronts Choice other / none-of-the-above\", \"S1 reflex keeps control / optional S2 one-use advice\", \"soft AGENTS.md rules vs linter\" (Abide / jev-pref), \"edit-phase vs turn-phase observation window\", banded confidence fail-open preference lint, \"extractive selection / pointer-not-generator\", \"encoder GLiNER compaction vs Jev Score compaction (same job; pointer not summarizer)\", \"fail-closed keep_full under mutation envelope\", \"CI flaky-vs-real merge gate\", \"fail-open VOI wake/resume (Horvitz)\", \"claim/evidence Stop integrity\", \"S1 extract + escalate-S2 indexer\", \"Harbor on/off routing\", \"policy-as-judgment PR marshal\", \"shadow-mode compaction rollout\", \"local /v1/systemone drop-in (stub until hf scorer)\", \"dataframe-native semantic columns\", \"route≠memory\", \"advisory sidecar receipts\", \"structure induction over bags\", \"AST ∩ semantic lint\", \"extractable-from-state / retrieve first\", \"decision-model vs constrained-LLM bake-off\", \"dual-process S1 decide / S2 generate\", \"combinatorial grid ≠ extractive\", TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, or \"formally verify with Jev\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." license: MIT metadata: version: 0.3.0 @@ -54,7 +54,7 @@ classical method you already trust, substitute it, classify the win "paraphrase brittleness", "allowlist then judge", "TOCTOU-of-Noul", "Jev inside the database / sqlite-jev", "Jev picks bitrate / join order / the model", "wait for Archer", "lint the request / missing - other", "training confronts Choice other / none-of-the-above", "soft AGENTS.md rules vs the linter", "screenshot Choice / omni System One", "extractive quotes / pointer not generator", "compaction summarize vs pointer", "encoder vs Jev compaction backend", "shadow-mode compaction rollout", "local drop-in vs stub scorer", "route vs memory", "when does it hold / extractable from state", "decision model vs constrained LLM", "dual-process S1/S2", or "combinatorial grid vs extractive": + other", "training confronts Choice other / none-of-the-above", "soft AGENTS.md rules vs the linter", "screenshot Choice / omni System One", "extractive quotes / pointer not generator", "compaction summarize vs pointer", "encoder vs Jev compaction backend", "shadow-mode compaction rollout", "CI flaky-vs-real merge gate", "fail-open VOI wake/resume", "claim vs session evidence", "S1 indexer escalate-S2", "Harbor on/off routing", "fail-open vs fail-closed wake vs CI gate", "local drop-in vs stub scorer", "route vs memory", "when does it hold / extractable from state", "decision model vs constrained LLM", "dual-process S1/S2", or "combinatorial grid vs extractive": read `references/faq.md`, then `references/mental-models.md`, then `references/mixed-architecture.md`, then @@ -114,14 +114,14 @@ classical method you already trust, substitute it, classify the win | Familiar method | Judgment shape | Detail | |---|---|---| | Mental models across domains (not SWE-only) | EU, abstention, VOI, MCDA, SDT, search/control, Leveson, NATM/Norman/snap-fit; **extractable-from-state boundary map** (self-contained vs needs outside knowledge) | `references/mental-models.md` | -| Judgment-model class (Jev is exemplar, not monopoly) | Species: decide / locate (GLiNER) / categorize (GLiClass) / rank / perceive; open heads include encoder DeBERTa, LoRA distill, openjev-lm, kev. Compaction job is backend-agnostic (Jev Score/Noul vs GLiNER2.5 encoder) | `references/judgment-class.md` | +| Judgment-model class (Jev is exemplar, not monopoly) | Species: decide / locate (GLiNER) / categorize (GLiClass) / rank / perceive; open heads include encoder DeBERTa, LoRA distill, openjev-lm, kev. Compaction job is backend-agnostic (Jev Score/Noul vs GLiNER2.5 encoder). Indexer cousin: GLiNER extract + escalate-S2 (10–50× unfilled) | `references/judgment-class.md` | | Open weights vs constrained decoding vs proprietary API | Three open paths: encoder open-jev / AR constrained decode (TypeAR + pcdServer; decision-token LoRA; packed one-forward logprob on open LLMs) / trained decision-only (Laya + ONNX port, Nimble, kev, **blackwood-rlcd** multimodal now, Archer Watch still Watch). Local `/v1/systemone` surfaces: jev-local (stub until `hf`), kev (trained pointer), von (tiny SAN — not a replica). Softmax over allowed tokens ≠ Noul | `references/judgment-class.md` (when-to-use table) | | Entropy as allocator (low / medium / high) | Typed low+medium decisions → System One marginals; high-entropy synthesis → frontier decoder. Product rhetoric, not a meter. **Hypothesis** | `references/judgment-class.md` | | Formal / semi-formal (proof vs judgment) | Sensor vs constraint vs searchlight; Alloy vs Apalache; DST trio; TOCTOU-of-Noul, AI×FM | `references/formal-methods.md` (one-screen: `references/formal-semi-formal.md`) | -| Mixed architecture (judgment model + LLM) | Provider judges, LLM writes, code owns control; not a stack replacement. Advisory sidecar never changes host routing. Dual-process: S1 decides, S2 generates (routing accuracy unmeasured). Internals ≠ FSM; placement is a **component node**. Browser-use strength = DOM-as-text + speculative fan-out, not vision | `references/mixed-architecture.md` | +| Mixed architecture (judgment model + LLM) | Provider judges, LLM writes, code owns control; not a stack replacement. Advisory sidecar never changes host routing. Dual-process: S1 decides, S2 generates (routing accuracy unmeasured); S1 specialists + S2 coordinator is the same split (description-only greenfield this hour). Internals ≠ FSM; placement is a **component node**. Browser-use strength = DOM-as-text + speculative fan-out, not vision. Fail polarity is per act: skip-wake fail-open vs merge-gate BLOCK fail-closed | `references/mixed-architecture.md` | | Context sieve | Relevance Noul per block; always-keep set in code; stub + recall key. Encoder cousin: GLiNER2.5 retention Choice + char-offset spans (gliner25-compaction); fail-closed keep_full; shadowMode default | `references/applied-mappings.md#1-context-sieve` | | Exact-text keep / drop | Choice include/exclude/mixed over candidates code already holds. Extractive quotes / pointer-not-generator (model never writes the excerpt; char-offset compaction same species) | `references/applied-mappings.md#2-exact-text-keep--drop` | -| Environment / harness triage | Scan every step for env failure; LLM autopsy only on flags | `references/applied-mappings.md#3-environment--harness-triage` | +| Environment / harness triage | Scan every step for env failure; LLM autopsy only on flags. Merge-gate cousin: cluster in code, judge labels cause, policy owns PASS/BLOCK (latch; judge never says ignore alone) | `references/applied-mappings.md#3-environment--harness-triage` | | Moderation and ranking | Hold-before-publish vs graded rerank; fail policy per action | `references/applied-mappings.md#4-moderation-and-ranking` | | Skill / tool routing | Choice over a closed catalog + whether-anything-fits; code dispatches. Route ≠ memory: cheap intent gate skips memory tours on easy routes | `references/applied-mappings.md#5-skill--tool-routing` | | Expensive observation router | Structural prove (text layer) ∩ remainder Noul (needs OCR?) | `references/applied-mappings.md#6-expensive-observation-router` | @@ -134,7 +134,7 @@ classical method you already trust, substitute it, classify the win | Retrieve + expensive relevance fn | Bounded rerank of a retrieved shortlist | `references/mappings.md#4-retrieval--bounded-semantic-reranking` (independent TREC DL2019 benchmark: Jev zero-shot best MAP 0.4748, nDCG@10 0.683 vs tuned monoBERT 0.718 — competitive, not dominant) | | Store as semantic index (SQL / SQLite / zoxide / dataframe) | Cheap exact predicates first; typed questions on the remainder. In-engine extension (sqlite-jev) vs CLI rewrite (jevql) vs path index (joxide) vs dataframe columns (jevpandas / jevframe) | `references/mappings.md#4-retrieval--bounded-semantic-reranking` | | Soft judgment inside a hard envelope | Model may only match the deterministic policy or be more conservative (bitrate ABR; query-planner override-when-confident; compaction mutations/shell operators → keep_full) | `references/mappings.md#12-runtime-assurance-sandwich-hypothesis`; `references/mappings.md#18-structural-prove--soft-remainder-hypothesis-as-domain-general-empirical-as-named-shapes` | -| Value of information / gather as an act | Pay for another observation only if EV(decision) improves more than cost; abstain from calling *any* model when a regex already answers (meta-VOI) | `references/mappings.md#6-value-of-information--gather-as-an-enumerated-act` (**Hypothesis** until a labeled act/outcome log; 149-row receipt is Empirical as a shape) | +| Value of information / gather as an act | Pay for another observation only if EV(decision) improves more than cost; abstain from calling *any* model when a regex already answers (meta-VOI). Fail-open wake/resume: skip the LLM turn only if the judge answers and p(wake) is low (Horvitz) | `references/mappings.md#6-value-of-information--gather-as-an-enumerated-act` (**Hypothesis** until a labeled act/outcome log; 149-row receipt is Empirical as a shape; wakegate 21/21 is smoke) | | Signal detection / ROC | Criterion and operating point from costs and base rate, not accuracy | `references/mappings.md#7-signal-detection--criterion-not-accuracy` (**Hypothesis** for non-SWE plots) | | Org / safety control structure | Sensor ≠ constraint (Leveson); STPA if the sensor lies | `references/mappings.md#8-control-structure--sensor--constraint-leveson` | | Search / control loops (any domain) | Algorithm stays yours; judgment substitutes one classifier step | `references/mappings.md#9-search--control-loops--one-substituted-classifier-step` | @@ -148,10 +148,10 @@ classical method you already trust, substitute it, classify the win | Input brittleness / paraphrase stability | Synonymous wording that swings p → abstain or rewrite | `references/mappings.md#17-input-brittleness--sensitivity-calibration-selective-abstention-hypothesis` (**Hypothesis**) | | Structural prove ∩ soft remainder | Allowlist *proves* the easy verbs; judge only unlisted leftovers; fail-open (cannot block) | `references/mappings.md#18-structural-prove--soft-remainder-hypothesis-as-domain-general-empirical-as-named-shapes` (**Hypothesis**; jevgate/OCR shapes Empirical) | | Effect-oriented state-machine loops | Soft predicates on transitions; code owns the transition | `references/mappings.md#19-effect-oriented-state-machine-loops-hypothesis` (**Hypothesis**; ZIO client, not Effect.ts) | -| Agent self-supervision / on-track detection | Pre-gate → output judge → done-check → supervisor nouls. S1 reflex keeps control; optional S2 is one-use advice | `references/agent-self-assessment.md` | +| Agent self-supervision / on-track detection | Pre-gate → output judge → done-check → supervisor nouls. S1 reflex keeps control; optional S2 is one-use advice. Claim/evidence Stop (anti-hallucinated-done); bounded Pi supervisor (shadow recovery, never generates commands) | `references/agent-self-assessment.md` | | Optimizer/program frameworks (Ax, DSPy) | Typed fields → one provider request; judge metrics; threshold discipline. Ax and DSPy climb LM-program knobs only | `references/optimizer-integration.md` | | Perception → decision pipeline / measure / hill-climb | Stages with a versioned state contract; frozen taskset; DSPy/Ax only on the LM-program slice. Specialist composition stays **Hypothesis**; open multimodal decide (blackwood-rlcd) is a named receipt. Structured observe→decide→verified-act (no screenshots) is a computer-use speed-layer receipt. Same section as the row below | `references/validation.md#eval--hill-climb` | -| Eval & hill-climb | Decision-stage jevals hygiene; Harbor taskset × harness × runtime; one score-composition table. Shared bake-off exemplar: open-jev-laya-bench (ECE/NLL/Brier; LLM-as-judge is not the score). Harbor-style frozen protocol vs constrained LLMs: DMB (accuracy/calibration/latency/cost; raw logs). Feedstock: jevals-data CC-BY-4.0 boards + JSONL (recompute-from-logs). Collab-arm curriculum: llm_autonomous vs scripted_plus_jev vs llm_plus_jev (Wilson / McNemar). Negative: combinatorial grid assembly ≠ extractive (ARC-AGI Direct Jev 4/400) | `references/validation.md#eval--hill-climb` | +| Eval & hill-climb | Decision-stage jevals hygiene; Harbor taskset × harness × runtime; one score-composition table. Shared bake-off exemplar: open-jev-laya-bench (ECE/NLL/Brier; LLM-as-judge is not the score). Harbor-style frozen protocol vs constrained LLMs: DMB (accuracy/calibration/latency/cost; raw logs). Feedstock: jevals-data CC-BY-4.0 boards + JSONL (recompute-from-logs). Collab-arm curriculum: llm_autonomous vs scripted_plus_jev vs llm_plus_jev (Wilson / McNemar). Negative: combinatorial grid assembly ≠ extractive (ARC-AGI Direct Jev 4/400). Harbor on/off routing: chess-engine tasks, hidden perft verifier, one-run preliminary (jev-gateway-bench). Pair CI merge-gate with Harbor + rh-guard | `references/validation.md#eval--hill-climb` | | (meta) Finding new mappings & applications | Toolbox sweep: judgment-shaped component of a known method, substituted + falsified | `references/toolbox-mapping.md` | | Named methods / operators / theorems | Substitution tiers: operand-judgments, preconditioned theorems, non-substitutable | `references/methods-catalog.md` | | (meta) Where a judgment model sits relative to any construct | 11 positions + logical-operator rules + position×construct traversal as the application generator | `references/composition-algebra.md` | diff --git a/.agents/skills/augustus/references/agent-self-assessment.md b/.agents/skills/augustus/references/agent-self-assessment.md index 4146c04..dddeaa3 100644 --- a/.agents/skills/augustus/references/agent-self-assessment.md +++ b/.agents/skills/augustus/references/agent-self-assessment.md @@ -22,7 +22,12 @@ relying: foreman, pi-jev, pi-warden, winnow, fast-jev-compaction, jev-judgment. answers it — one Noul only for the semantic remainder ("does this reply claim the work is finished?"), one threshold. Spending the model on the countable half is the `/bin/ls`-as-first-tier pattern - (`mappings.md` §18). + (`mappings.md` §18). Claim/evidence cousin + ([clear-head](https://github.com/VladyslavHontar/clear-head), + ~16:48): check factual claims against **what was actually read this + session**; keyword retriever, not semantic; below `JEV_FIRM` 0.6 + never blocks; true-but-unread still flags unsupported + (`notes.md` §51). Anti-hallucinated-done, not a test runner. 4. **Stuck-detector**: three failures with the same strategy → ask for a new hypothesis, not another retry. 5. **Supervision during long runs** (foreman): separate concurrent loop @@ -46,6 +51,18 @@ relying: foreman, pi-jev, pi-warden, winnow, fast-jev-compaction, jev-judgment. `conf ≥ τ` S1 decides else S2 generates; routing fails open; safety fails closed; **routing accuracy not measured**; keyword fallback is not S1 (`notes.md` §49). Tune τ on your escalation log. + **Distinguish names:** the existing "foreman" *shape* here is the + Kevthetech143/super-jev loop (named probabilities → hysteresis + table). [`reification-labs/foreman`](https://github.com/reification-labs/foreman) + (~16:48) is a **description-only** Elixir/Phoenix scaffold claiming + parallel S1 specialists + one S2 coordinator with typed + `{value, probability}` — README is stock Phoenix; `mix.exs` has no + Jev dep. Do not invent an Elixir API (`notes.md` §51). + Bounded Pi supervisor of the same lifecycle: + [jevons](https://github.com/LilDojd/jevons) — not a second agent; + default recovery **shadow**; steering never generates commands. + Distinguish from pi-jev-approver (fail-closed remainder) and + pi-jev-context (sieve). 6. **Context economy**: the context-sieve card (`references/applied-mappings.md#1-context-sieve`). Judge every large tool result with one relevance Noul before it enters context. Hide @@ -89,7 +106,10 @@ relying: foreman, pi-jev, pi-warden, winnow, fast-jev-compaction, jev-judgment. ([jev-reviewer](https://github.com/choxos/jev-reviewer); `notes.md` §48). Compaction: point at character offsets in the tool result ([gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction); - `notes.md` §50). + `notes.md` §50). Session-evidence Stop + ([clear-head](https://github.com/VladyslavHontar/clear-head)): claims + vs **what the assistant actually read**; keyword retriever; below + firm-confidence never blocks (`notes.md` §51). - Self-report fidelity: compare the agent's claimed action with its actual trace via decomposed Nouls (right tool? args match schema? result matches call?). Escalate on low confidence; never auto-retry. @@ -109,7 +129,11 @@ fuller productized path of that contract (compile / calibrate / tune / replay; one Score per rule on the diff; bands + fail-open). Soft rules → soft judgment; the linter owns hard rules. Shadow-mode the gate first; permit remains a separate axis from confidence. -`notes.md` §47. +`notes.md` §47. Plain-English PR-check cousin +([if-ai](https://github.com/Victor-Casado/if-ai)): one condition + +required min-confidence; fail-closed on error / empty / low +confidence. [jev-marshal](https://github.com/LightningK0ala/jev-marshal) +is Watch / empty repo this pass (`notes.md` §51). ## Using Jev to test and optimize the skill suite itself diff --git a/.agents/skills/augustus/references/applied-mappings.md b/.agents/skills/augustus/references/applied-mappings.md index d9ee2ff..6f36cb7 100644 --- a/.agents/skills/augustus/references/applied-mappings.md +++ b/.agents/skills/augustus/references/applied-mappings.md @@ -104,7 +104,11 @@ code numbers sentences; one broadcast (Choice/Noul/Score + per-sentence Nouls); the model never writes; `redecide` retunes thresholds on the log. [jev-reviewer](https://github.com/choxos/jev-reviewer) — the model **points at line ids**; code copies verbatim quotes with place; *not -found* is an answer. Compaction cousin +found* is an answer. Claim/evidence Stop cousin +([clear-head](https://github.com/VladyslavHontar/clear-head), ~16:48): +the model judges claims against **keyword-retrieved session lines**, +not against another model's prose; `JEV_FIRM` below 0.6 never blocks +(`notes.md` §51). Compaction cousin ([gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction), ~16:22): the model points at **character offsets** in a tool result; code copies those bytes; a generator summary is the rejected species @@ -159,7 +163,22 @@ confidence goes back to the agent, not into silent mitigation. Noul workaround; run status silent / disclosed / recovered / clean. Heuristic fixture 12 traces: P=R=0.86; Jev on that fixture not yet measured. Pre-mortem: scan a new sandbox *before* users meet it, fail -the build on *silent* env-breaks — a shape, not a CLI. **Counterexample**: sampling 2% of production with an LLM judge — +the build on *silent* env-breaks — a shape, not a CLI. +**Merge-gate cousin (Empirical as README behavior / offline demo, +2026-09-18 ~16:48):** +[latch](https://github.com/CaseReed/latch) — cluster a finished red +run (Playwright / Jest / pytest / JUnit) **in code**; Jev labels each +cause (≤8 calls; cached free); **policy** returns Gate: PASS (infra +noise) vs Gate: BLOCK (real failure). The judge never says "ignore" +alone; `ignore_as_infra` needs `env_cascade` + an infra fingerprint. +Reporter never fails Playwright (missing key → `needs_human`); +`--gate` is a separate CI step. Demo: 8 connection errors → PASS; 5 +assertions → BLOCK. Message-based grouping fragments logic +regressions (`pallets/click`: 13 failures → 10 clusters). Pair with +Harbor (frozen CI artifacts × PASS/BLOCK) and rh-guard +(eval-integrity). Their policy thresholds are not class constants +(`notes.md` §51). Do not copy the reporter. +**Counterexample**: sampling 2% of production with an LLM judge — the economics inversion is the point. **Test**: planted harness bugs recovered; false-flag rate on known-clean runs; LLM never runs on the clean majority. High-stakes cousin: `luantak/is-malicious` is a *pre-run* diff --git a/.agents/skills/augustus/references/composition-algebra.md b/.agents/skills/augustus/references/composition-algebra.md index cb6f80e..864cd12 100644 --- a/.agents/skills/augustus/references/composition-algebra.md +++ b/.agents/skills/augustus/references/composition-algebra.md @@ -149,6 +149,9 @@ Reusable shapes when generating applications: (Claude CLI) vs 1.3 s Jev at concurrency 8. Compaction rollout: gliner25-compaction public default `shadowMode: true` (log proposed reduction; do not replace history) (`notes.md` §50). + Recovery cousin: [jevons](https://github.com/LilDojd/jevons) default + recovery **shadow** (record, do not interrupt); steering never + generates commands (`notes.md` §51). Calibration warning (calibre): routing thresholds and ROI do **not** transfer across datasets — every gate is a per-dataset measurement (see validation.md). diff --git a/.agents/skills/augustus/references/faq.md b/.agents/skills/augustus/references/faq.md index ddb0ceb..aac8986 100644 --- a/.agents/skills/augustus/references/faq.md +++ b/.agents/skills/augustus/references/faq.md @@ -149,6 +149,11 @@ Hole first, logo last. These are **species**, not aliases uses that local multi-head as **categorize** (retention action) plus **locate** (character-offset spans); code copies; not a summarizer (`notes.md` §50). + Indexer cousin: [s1-graphify-indexer](https://github.com/GreyssonEnterprises/s1-graphify-indexer) + uses GLiNER2 on the bulk of a repo graph and escalates an LLM only + on the ambiguous tail **if the backend loaded**. GitHub one-liner + 10–50× is a **target, not a measured speedup** — table TBD + (`notes.md` §51). Not a Noul. - **GLiClass (categorize):** one forward pass over text + *all* labels; sigmoid multi-label or softmax single-label. Use for large or changing tag sets. Scores are class affinities, not automatically a gateable @@ -321,9 +326,12 @@ window (edit vs turn). Fix false positives in the rubric, not the model. Productized path: [Abide](https://github.com/coldteadotai/abide); earlier contract pointer: jev-pref. Complementary, not the same product: [rh-guard](https://github.com/24601/rh-guard) (reward-hacking / -eval integrity). Request-shape lint still sits upstream (wellposed / +eval integrity). Plain-English PR check: +[if-ai](https://github.com/Victor-Casado/if-ai) (one condition + +required min-confidence; fail-closed on error). [jev-marshal](https://github.com/LightningK0ala/jev-marshal) +is Watch / empty this pass. Request-shape lint still sits upstream (wellposed / `tenbin`). `mixed-architecture.md`; `question-design.md`; `notes.md` -§47. +§47, §51. ## Can confidence gating catch a forced wrong Choice? @@ -452,7 +460,10 @@ ids; *not found* is an answer. [solari-reflex](https://github.com/hitakshiA/sola never lets model output become a selector. Compaction is the same species: [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) copies exact source spans; a prose summary of the tool result is -generation, not keep/drop (`notes.md` §50). Generation is only for +generation, not keep/drop (`notes.md` §50). Claim/evidence Stop: +[clear-head](https://github.com/VladyslavHontar/clear-head) judges +against retrieved session lines, not generated prose (`notes.md` +§51). Generation is only for TYPE/prose when something must be written. `applied-mappings.md` §2; `notes.md` §48, §50. @@ -469,6 +480,19 @@ unlike Abide / jevgate fail-open. Ship `shadowMode` first (default true: log, do not replace history). Not Jev. Not multimodal. `judgment-class.md`; `notes.md` §50. +## Fail-open or fail-closed — which? + +Name the **irreversible act**, then pick polarity. Compaction *drop* +fails closed to `keep_full`. Wake *skip* fails open (wake on error / +unsure): [wakegate](https://github.com/shitianfang/wakegate) skips +only if Jev answers and p(wake) < 0.2. Merge *PASS* on a red run +fails closed at the gate: [latch](https://github.com/CaseReed/latch) +`--gate` BLOCKs unless infra is confirmed; the Playwright reporter +stays fail-open. [if-ai](https://github.com/Victor-Casado/if-ai) +fails the Action on error / empty / low confidence. jevgate cannot +block; pi-jev-approver fails closed without a key; Abide is fail-open +on diffs. Same sandwich, opposite authorized act. `notes.md` §50, §51. + ## Is routing the same as memory? No. A cheap intent gate can skip a memory/tool *tour* on diff --git a/.agents/skills/augustus/references/judgment-class.md b/.agents/skills/augustus/references/judgment-class.md index 12ca816..36a4e67 100644 --- a/.agents/skills/augustus/references/judgment-class.md +++ b/.agents/skills/augustus/references/judgment-class.md @@ -99,6 +99,18 @@ below, next to the when-to-use table. tokens; no published retention-quality rates (`notes.md` §50). Fastino/GLiGuard sibling *class*, not a GLiGuard safety-schema clone. Do not copy the plugin install. + **Code-graph indexer (Empirical as README behavior; 10–50× is a + target, 2026-09-18 ~16:48):** + [s1-graphify-indexer](https://github.com/GreyssonEnterprises/s1-graphify-indexer) + (+ sibling `s1-indexer`) uses GLiNER2 as the default System-1 + backend to build a semantic code graph and escalates an LLM only + on the ambiguous tail **and only if the backend loaded**. Jev / + Needle stubs `available=False` until implemented. If GLiNER2 + cannot load: degraded file-node graph, exit nonzero, do not dump + the repo to `S1_LLM_CMD`. Query does not invent edges. GitHub + one-liner 10–50× is **unfilled** — not Empirical (`notes.md` §51). + Same locate family as compaction; different hole (index vs + keep/drop). License not on GitHub this pass. - **Decide.** Typed Choice/Score/Noul with a decision/proper-scoring objective. That is Jev's product claim. Open heads copy the *shape*; distillation copies the *teacher* (openjev-lm, jev-gate-student-b). diff --git a/.agents/skills/augustus/references/mappings.md b/.agents/skills/augustus/references/mappings.md index 4cc3e0a..cc642f3 100644 --- a/.agents/skills/augustus/references/mappings.md +++ b/.agents/skills/augustus/references/mappings.md @@ -128,7 +128,12 @@ your state machine" is the architectural instinct — `notes.md` §49). **Example**: game director — Jev judges whether player dialogue is conciliatory or threatening; code enforces inventory, prerequisites, chronology, reachable scenes (cf. HEIST//ONE: six guards batched, simulation -validates every proposal). **Counterexample**: decomposing tool-trace +validates every proposal). **Merge-gate circuit (Empirical as README +behavior, 2026-09-18 ~16:48):** +[latch](https://github.com/CaseReed/latch) — Jev labels a clustered +cause; a **table** maps cause × confidence × fingerprint → PASS / +BLOCK / needs_human. The judge is a sensor, not the merge act +(`notes.md` §51). **Counterexample**: decomposing tool-trace verification into per-call schema nouls works; asking "is the trace correct" as one Noul hides nine judgments. **Test**: full truth table / transition cases incl. contradictory outputs, stale observations, invalid combos. @@ -291,6 +296,15 @@ then ask. Atlas history suite: wrong @ 0.90 without context → right @ 0.97 with the passage (`notes.md` §49; `mental-models.md` §boundary). That observation is VOI with a named receipt. Do not rely on bare recall. +**Fail-open wake/resume (Empirical as README safety table; 21/21 is +smoke, 2026-09-18 ~16:48):** +[wakegate](https://github.com/shitianfang/wakegate) — skip a sleeping +agent's LLM turn only if Jev answers **and** p(wake) < 0.2; user +message / nothing-to-judge / skip-limit / error / unsure all **wake**. +Horvitz mixed-initiative: pay for the turn iff EV(decision) beats +the token cost. Savings unmeasured. Same-author scenarios+question; +not a benchmark (`notes.md` §51). Contrast pi-jev-approver +fail-closed without a key and jevgate cannot-block. **Beyond SWE (Hypothesis until you log act/outcome pairs):** full PDF vs abstract; customer call vs CRM fields that already fail a hard rule (credit limit is exact); blood test vs @@ -322,6 +336,10 @@ report hits / false alarms at the operating point, not accuracy **Example (Empirical as family shape):** firehose / Near Here moderation — judge once, re-policy in code. +**CI merge-gate (Empirical as README / demo, 2026-09-18 ~16:48):** +[latch](https://github.com/CaseReed/latch) — false PASS on a real bug +>> false BLOCK on infra; criterion lives in the policy table, not in +the cause label (`notes.md` §51). [`jp-sns-jev7-estimator`](https://huggingface.co/kokuren/jp-sns-jev7-estimator) is the rare-class warning in one table: seven distilled teacher scores that the card says are **not** calibrated probabilities, and `threat` @@ -376,6 +394,11 @@ mapping §5 is the taxonomy-beam special case). Economics inversion: per-node judgments were known and too expensive; they are now default. Control: hysteresis, continue / stop / retry / verify — the model estimates named probabilities; the controller is a table with memory. +**Bounded Pi supervisor (Empirical as README policy, 2026-09-18 +~16:48):** [jevons](https://github.com/LilDojd/jevons) — not a second +agent; Jev interprets evidence; code owns freshness/limits; default +recovery **shadow**; steering never generates commands +(`notes.md` §51). Distinguish from pi-jev-approver / pi-jev-context. **Does not transfer**: Jev as the planner that picks its next tool in a loop; bandits without observed rewards; speculative depth without a simulator; PufferLib Ocean scores as a capability claim @@ -682,6 +705,16 @@ judges only the remainder; uncertain **fails closed to `keep_full`** (`notes.md` §50). Same sandwich, opposite fail policy from jevgate (cannot block) and Abide (fail-open on diffs): the authorized act is a destructive reduction of memory. +**Name the irreversible act (2026-09-18 ~16:48).** Wake *skip* is +irreversible (the agent stays asleep) → +[wakegate](https://github.com/shitianfang/wakegate) authorizes skip +only at p < 0.2 and otherwise **wakes** (fail-open on the skip). +Merge *PASS* is irreversible if the bug was real → +[latch](https://github.com/CaseReed/latch) `--gate` BLOCKs unless +infra is confirmed; the Playwright reporter stays fail-open. +[if-ai](https://github.com/Victor-Casado/if-ai) fails the Action on +error / empty / low confidence (fail-closed on the check). +`notes.md` §51. **Beyond SWE (Hypothesis):** recipe book ∩ "does this leftover look done?"; labor-law allowlist ∩ hiring-fit Noul; SPF/DKIM pass ∩ phishing Noul on the body. **Counterexample:** diff --git a/.agents/skills/augustus/references/mental-models.md b/.agents/skills/augustus/references/mental-models.md index 84fb16b..d865c4b 100644 --- a/.agents/skills/augustus/references/mental-models.md +++ b/.agents/skills/augustus/references/mental-models.md @@ -474,6 +474,10 @@ Use these as *existence proofs of a position*. Write your own card. | Browser / DOM candidates → act | numbered elements from a **text** snapshot | Choice / Nouls over those ids (**Empirical** as atlas browser-use *shape*: DOM-as-text + fan-out, not vision) | Click in code; no screenshots | | Knowledge / recall | fact that is not in the document | **Do not ask.** Retrieve the passage first; then a self-contained Choice (**Empirical**: history suite A wrong@0.90 → C right@0.97) | Index, citation, the passage in `state` | | Dual-process cascade | cheap classify / route vs write | S1 typed decision + τ; S2 generates only on low conf (**Empirical as a productized metaphor**; routing accuracy **unmeasured** — dual-process-ai) | Safety still fail-closed in code | +| CI merge-gate | ignore infra noise without merging a real bug | cause Choice per cluster (**Empirical**: latch demo PASS vs BLOCK) | Cluster + fingerprint + `--gate` table; reporter never fails the runner | +| Sleeping-agent resume | skip a worthless LLM turn | p(wake) (**Empirical** as safety table; 21/21 smoke — wakegate) | User-message / skip-limit / error always wake | +| Claim integrity at Stop | do not ship hallucinated-done | supports/contradicts vs session evidence (**Empirical**: clear-head) | Keyword retrieve; firm-confidence floor never blocks | +| Code-graph index | cheap S1 extract, S2 only on the tail | GLiNER locate + confidence escalate (**Hypothesis** as 10–50×; **Empirical** as degraded-load / no-invent-edges) | Graph in code; do not dump repo if S1 failed to load | | Combinatorial puzzle | whole grid / program that must be consistent | **Rejected as extractive.** Cell-wise Choice assembly is not keep/drop (ARC-AGI Direct Jev 4/400) | Search, a program, a simulator | | Moderation | hold before publish | hazard Nouls (**Empirical** as family) | block/review policy | | Phishing / fraud screen | hold vs deliver | SDT criterion on a Noul | blocklist, SPF/DKIM exact (**Hypothesis**) | diff --git a/.agents/skills/augustus/references/methods-catalog.md b/.agents/skills/augustus/references/methods-catalog.md index f1ff0b7..b7367c6 100644 --- a/.agents/skills/augustus/references/methods-catalog.md +++ b/.agents/skills/augustus/references/methods-catalog.md @@ -31,10 +31,11 @@ judgment component is new). | Self-consistency / ensembling of judges | Repeated independent ratings of the same object | N repeats over one state (output tokens free); entropy/disagreement across repeats as the review signal | Aggregation, escalation policy | **Empirical recipe** (self-consistency: nouls cookbook) | | Judge qualification (interrater reliability) | A judge worth gating must be repeatable | Repeated judgments over frozen outputs before trusting either Jev or LLM as judge | Variance stats, agreement metrics | **Empirical recipe** (jev-as-a-judge: 224–279× tighter than GPT judge) | | Neyman–Pearson / selective classification | Decision threshold under error costs | One threshold per action, set on split A, reported on split B; abstention path | Loss model, ROC analysis | **Contract + empirical** (confidence-routing; evaluator script) | -| Value of information (EVPI / EVSI) | Whether another observation is worth its cost | Gather as an enumerated act; pay iff expected decision-loss drop > cost | Cost of the observation; the loss table | **Hypothesis** as a numeric calculator; **Contract** as the placement (`mappings.md` §6) | +| Value of information (EVPI / EVSI) | Whether another observation is worth its cost | Gather as an enumerated act; pay iff expected decision-loss drop > cost. Wake/resume: skip the LLM turn only if the judge answers and p is low | Cost of the observation; the loss table; skip-limit | **Hypothesis** as a numeric calculator; **Contract** as the placement (`mappings.md` §6). wakegate 21/21 is smoke (`notes.md` §51) | | Signal detection (Green & Swets) | Evidence variable + criterion | Noul as noisy evidence; t from costs and base rate; ROC/PR on your labels | Operating point, base-rate tracking | **Hypothesis** for non-SWE plots; **Empirical** as moderation *shape* (`mappings.md` §7) | | Reliability calibration (Platt/temperature) | Raw scores → calibrated probabilities | Noul is natively calibrated **in-distribution only**; verify with reliability bins on your own population; re-fit a correction out-of-distribution | Calibration fitting, binning | **Empirical recipe** (ECE 0.0313 in-distribution; 32% OOD collapse — Archer Hume). Atlas: DAIR Emotion dangerous-high (48% / 0.819); DMB S5 ECE 0.246 (`notes.md` §49) | | Frozen-protocol bake-off vs constrained LLMs | Same items, accuracy + ECE + latency + cost + honesty | Decision-model as one contender class, not the score | Protocol, raw logs, baselines | **Empirical as Harbor/jevals practice** (DMB v2; jevals-data CC-BY-4.0 recompute-from-logs; `notes.md` §49) | +| Harbor on/off routing | Same task, routing on vs off, hidden verifier | Tool Choice per turn; cheaper unsolved is not a saving | Fresh gateway; perft / checks the agent never sees | **Empirical as a *shape* and one-run signal** (jev-gateway-bench chess-bugfix 36/36 both; 4 vs 6 LLM req; `notes.md` §51). Not a measurement until reps ≥5 | | Survey scoring / psychometrics | Rubric level judgment with defined anchors | Score with concrete level descriptions; probabilities read beside every score | Weighted aggregation, reliability analysis | **Contract** (score docs: split composite judgments) | ## Search, planning & operations research @@ -57,7 +58,7 @@ judgment component is new). | Method | Judgment-shaped component | Jev substitution | Stays in code | Status | |---|---|---|---|---| | Entropy as uncertainty signal | Measuring "how spread is this belief" | Entropy of returned distributions across repeats or options — computed in code from returned probabilities | All arithmetic | **Empirical recipe** (cookbook pattern) | -| Detector / Neyman filter (context) | Is this artifact relevant to the current task? | One relevance Noul per block before it enters context; stub + recall key. Encoder cousin: retention Choice + span locate, copy verbatim | Cache, recall, safety keeps; mutation/shell envelope | **Empirical recipe** (winnow ≤0.22 hide; compaction 2-noul rule; pi-jev-context hide-not-delete; gliner25-compaction GLiNER2.5 keep_full/keep_evidence/keep_call_only/drop, `notes.md` §50) | +| Detector / Neyman filter (context) | Is this artifact relevant to the current task? | One relevance Noul per block before it enters context; stub + recall key. Encoder cousin: retention Choice + span locate, copy verbatim. Indexer cousin: GLiNER extract on the bulk, escalate LLM on the tail | Cache, recall, safety keeps; mutation/shell envelope; do not dump the repo if S1 failed to load | **Empirical recipe** (winnow ≤0.22 hide; compaction 2-noul rule; pi-jev-context hide-not-delete; gliner25-compaction GLiNER2.5 keep_full/keep_evidence/keep_call_only/drop, `notes.md` §50; s1-graphify-indexer degraded-load / no-invent-edges, 10–50× unfilled, `notes.md` §51) | | Anomaly detection | Does this deviate from expected shape? | Guard nouls + harm Score over {input, output, tool trace} | Baselines, alert thresholds | **Empirical recipe** (guardrails cookbook; pi-jev output judge) | | Allowlist ∩ remainder (code-then-model) | Unlisted / unstructured leftovers after a **proof** | Typed questions only on the unknown tier; admit iff every p < τ | Proven/refused in code; cannot block unless a sandbox sits under | **Empirical recipe** (jevgate 0/59 unsafe unasked held-out; allowlist *proves* read-only verbs; doc-router 1.74× $). Domain-general: `mappings.md` §18 | | Decision-token LoRA (constrained-AR) | Specialize a generator for parallel constrained fields | Loss only on the single decision token; KV broadcast across fields | Schema, candidate tokens, policy | **Empirical recipe** as Foodoo1 200-case / 4-field receipt (fraud_risk 64→95%, overall 85.2→98.8%, ~234 ms); **Hypothesis** as a general recipe. Synthetic; not a financial product. Softmax ≠ Noul | @@ -67,10 +68,10 @@ judgment component is new). | Method | Judgment-shaped component | Jev substitution | Stays in code | Status | |---|---|---|---|---| -| Claim–evidence entailment (NLI) | supports / contradicts / not-established per claim–source pair | One Choice per pair + review flag; judge against the cited source text only | Quote extraction, citation graph, audit log. When the answer *is* a span you already hold, **point** at line ids and copy verbatim — the model never writes the excerpt (jev-reviewer) | **Empirical recipe** (citation_check cookbook; jev-reviewer sample study, `notes.md` §48). Atlas: paraphrase_support and reversed_meaning_high_overlap correctly judged when both texts are in `state` (`notes.md` §49) | +| Claim–evidence entailment (NLI) | supports / contradicts / not-established per claim–source pair | One Choice per pair + review flag; judge against the cited source text only. Stop-hook cousin: claims vs **session** evidence | Quote extraction, citation graph, audit log. When the answer *is* a span you already hold, **point** at line ids and copy verbatim — the model never writes the excerpt (jev-reviewer). Keyword retrieve is not semantic (clear-head) | **Empirical recipe** (citation_check cookbook; jev-reviewer sample study, `notes.md` §48; clear-head Stop, `notes.md` §51). Atlas: paraphrase_support and reversed_meaning_high_overlap correctly judged when both texts are in `state` (`notes.md` §49) | | Extractive selection + offline re-threshold | Keep/drop over sentences, ids, or character offsets code already holds | Per-item Noul/Choice + one broadcast; join in order; `redecide` on the log with no new calls. Compaction: retention Choice + span locate | Numbering, header skip, thresholds, publish permission; mutation envelope; copy exact bytes | **Empirical recipe** (testimonial-miner 8-request fixture; jev-reviewer; gliner25-compaction char-offset copies, `notes.md` §48, §50). Cousin of applied-mappings §2 | | Combinatorial grid / program synthesis | Consistent whole-object from many cells | **Rejected as extractive.** Cell-wise Choice does not assemble ARC grids (4/400 Direct Jev) | Search, a program, a simulator | **Empirical as a negative** (`notes.md` §49) | -| Spec vs artifact conformance (model checking *mindset*) | Property holds / violated / unverifiable for a named requirement | One Noul/Score per requirement, batched; violated → named rule back into context (pi-warden / Abide shape). This is **not** TLC/Apalache/GNATprove | Requirement enumeration, enforcement, logging; the **linter** if the rule is lintable; the real checker if you have one | **Empirical recipe** (pi-warden: 6→0 rule breaks, 150 paired runs; jev-pref: YOU define the rule; Abide: productized compile/calibrate/tune/replay, `notes.md` §47). Ownership split: `formal-methods.md` | +| Spec vs artifact conformance (model checking *mindset*) | Property holds / violated / unverifiable for a named requirement | One Noul/Score per requirement, batched; violated → named rule back into context (pi-warden / Abide shape). This is **not** TLC/Apalache/GNATprove | Requirement enumeration, enforcement, logging; the **linter** if the rule is lintable; the real checker if you have one | **Empirical recipe** (pi-warden: 6→0 rule breaks, 150 paired runs; jev-pref: YOU define the rule; Abide: productized compile/calibrate/tune/replay, `notes.md` §47; if-ai: plain-English PR check, fail-closed on error, `notes.md` §51). Ownership split: `formal-methods.md` | | AST ∩ semantic lint | Semantic remainder after a parser already extracted units | Typed questions on Tree-sitter targets; do not execute scanned code | Parser, selection, fail-on; `tenbin` owns the lint *skill* | **Empirical as a shape** (jevscan 0.2.0rc4; not a calibration claim; `notes.md` §48) | | Alloy finder vs Apalache / TLC | Which bound, which counterexample, is the property tautological? | Triage instances/CEs; never "this spec looks right" | Analyzer / SMT / explicit-state engine | **Hypothesis** as product; **Contract** as ownership (`formal-methods.md` §2) | | Type-checking analog | Does this planned call match the schema/operation/target? | Decomposed nouls over {request, schema, trace}; never trust a Jev pass as authorization | Real validation of operation+target in code | **Empirical recipe** (validation.md self-monitoring) | diff --git a/.agents/skills/augustus/references/mixed-architecture.md b/.agents/skills/augustus/references/mixed-architecture.md index 02aaf43..1e9fa6d 100644 --- a/.agents/skills/augustus/references/mixed-architecture.md +++ b/.agents/skills/augustus/references/mixed-architecture.md @@ -179,6 +179,9 @@ not a global virtue: |---|---|---| | Drop a RAG chunk or log line | **Fail open** (keep on error) | A false drop loses evidence; a false keep costs tokens | | Compact / drop a completed tool result | **Fail closed** to keep-full (`gliner25-compaction`) | Compaction is a destructive edit of memory. Uncertain *looks* like keep-on-error from the evidence side; name the *reduction* as the act. Contrast Abide / jevgate fail-open | +| Skip waking a sleeping agent | **Fail open** (wake on error / unsure / no key) (`wakegate`) | Skip is the irreversible act. User-message, skip-limit, nothing-to-judge, and p in 0.2–0.5 all wake. Contrast pi-jev-approver fail-closed without a key | +| Merge a red CI run | **Fail closed** on `--gate` (`latch`); reporter stays fail-open | False PASS merges a real bug. Missing key never fails Playwright; the gate is a separate step. Judge never says ignore alone | +| Plain-English PR check | **Fail closed** on error / empty / low confidence (`if-ai`) | A skipped or timed-out check is not a pass. Threshold is policy, not measured correctness | | Route to a tool / start a side effect | **Fail closed** (don't call) | A wrong tool is an action | | Rerank a retrieved list | Fail open: keep retrieval order (`WiktorB2004/llama-index-jev`, **Empirical recipe** on BEIR nfcorpus: MiniLM 0.340 nDCG@5 → MiniLM+Jev 0.396; rerank fails open, *select* fails closed). Listwise/cross-encoder scores belong here, not on the row above. | Ranking errors are quality; selection errors are control-flow | @@ -203,7 +206,9 @@ Worked placements (2026-09-18 topic:jev hour + prior archive): runs for root cause. Prefilter of *analyst attention*. The named cut this hour: Noul `env_broken` *as opposed to* the agent's own bug; silent vs disclosed vs recovered vs clean. Pre-mortem of a new - sandbox *before* users meet it (`notes.md` §42). + sandbox *before* users meet it (`notes.md` §42). Merge-gate of the + same split: [latch](https://github.com/CaseReed/latch) clusters in + code, labels with Jev, policy owns PASS/BLOCK (`notes.md` §51). - **Realtime hold-before-publish** — community moderation claims ~200ms (**Hypothesis** as a number; **Empirical** as a family via Near Here / jev-experiments firehose in the archive). Thresholds stay yours. @@ -258,6 +263,11 @@ Live ecosystem (examples of the *shape*, not SDKs to copy): receipt. - Function-calling cookbook (**Contract**): function *names* and closed-set args as questions; code still validates the call. +- [jev-gateway](https://github.com/vinilana/jev-gateway) (~16:48) — host + adapter: Jev picks the tool (and closed-set args); LLM fills open + args or is skipped (`direct`). Fail-open passthrough if Jev is down. + Harbor on/off measurement: [jev-gateway-bench](https://github.com/vinilana/jev-gateway-bench) + (one-run signal, not a measurement; `notes.md` §51). Do not copy ports. Routing ROI does not transfer across datasets (`validation.md`, calibre). `FirasSX914/Janus` exists to *measure* when Jev vs another model wins on @@ -370,6 +380,10 @@ Related placements: - **AGENTS.md / project prefs as criteria** — jev-pref states the contract; Abide productizes it; pi-warden rule breaks 6→0 on 150 paired runs (`agent-self-assessment.md`; `notes.md` §47). + [if-ai](https://github.com/Victor-Casado/if-ai): one plain-English + condition + required min-confidence; fail-closed on error + (`notes.md` §51). [jev-marshal](https://github.com/LightningK0ala/jev-marshal) + is Watch / empty. - **Confidence gates + shadow mode** — `AntonioCoppe/jev-harness` (48.9s Claude CLI vs 1.3s Jev on a 24-row filter). Log would-do until evals pass. Selective abstention (`mappings.md` §2): low confidence is @@ -416,7 +430,7 @@ decision-design card. Do not clone APIs from READMEs. |---|---|---|---| | Hold-before-publish moderation | Hazard Nouls + harm Score | Block/review/pass policy | Near Here / firehose family | | Tool / engine / skill select | Choice + fits-Noul | Dispatch, auth, reject-all | skillranker, LlamaIndex selectors, Toolrouter | -| Preference lint | Per-rule Score/Noul on a diff | Rule text, linter for hard rules, bands + fail-open | jev-pref (contract), Abide (productized), JevLint | +| Preference lint | Per-rule Score/Noul on a diff | Rule text, linter for hard rules, bands + fail-open | jev-pref (contract), Abide (productized), JevLint; if-ai (plain-English PR check, fail-closed on error); jev-marshal (Watch / empty repo) | | Context / log prune | Per-line or per-block relevance; or a retention Choice + spans | Always-keep set, recall keys; mutation envelope in code; shadow before replace | jevprune, winnow; fast-jev-compaction / pi-jev-compaction (Jev); gliner25-compaction (GLiNER2.5) | | Exact hunk staging | Per-hunk include/exclude/mixed | `git diff`, atomic apply | git-jev-stage | | Semantic `WHERE` | Noul/`jev_prob` over a row | SQL, indexes, LIMIT | jevql (CLI; DB sees ordinary SQL); sqlite-jev (in-engine extension) | @@ -447,6 +461,14 @@ decision-design card. Do not clone APIs from READMEs. | Jump-by-description | Noul relevance on a local shortlist | zoxide index, local paths only | joxide | | Game move | Choice over legal actions | Rules, legality, win check | jev-plays-games | | Analyst attention cascade | Step-level silent-failure Nouls | Grouping, LLM autopsy | OpenSmoke | +| CI merge-gate (flaky vs real) | Cause Choice per clustered signature | Cluster + fingerprint + PASS/BLOCK table; reporter never fails the runner | latch (`notes.md` §51) | +| Fail-open wake / resume | p(wake) on waitingFor × event | Sleep duration, skip-limit, user-message always wakes | wakegate | +| Claim/evidence Stop | supports / contradicts / not-addressed per claim | Keyword retrieve session lines; firm-confidence floor never blocks | clear-head | +| Harbor on/off routing | Tool Choice per turn | Hidden verifier; fail-open if Jev down | jev-gateway + jev-gateway-bench (one-run signal) | +| Device-loop Choice | Folder among a closed catalog | Never invent folders; extension-map fallback | jev-downloads-sorter | +| S1 extract + escalate-S2 index | GLiNER spans / relations on the bulk | Graph in code; LLM only if backend loaded and low conf; query does not invent edges | s1-graphify-indexer (10–50× unfilled) | +| S1 specialists + S2 coordinator | Typed {value, probability} | Coordinator / hysteresis in code | reification-labs/foreman (**description-only** Phoenix scaffold; not the super-jev loop) | +| Bounded Pi supervisor | Skills / recovery / review / verify | Shadow default; never generates commands | jevons | On-device / Home Assistant / mobile are newly-feasible via the economics inversion, not proven ports of every app. Named placements this hour diff --git a/.agents/skills/augustus/references/toolbox-mapping.md b/.agents/skills/augustus/references/toolbox-mapping.md index 5c45774..09134ea 100644 --- a/.agents/skills/augustus/references/toolbox-mapping.md +++ b/.agents/skills/augustus/references/toolbox-mapping.md @@ -74,13 +74,14 @@ component; keep the rest of the method in code. | Experimental design: extractable-from-state axis | Same question with vs without a supporting passage; citation paraphrase vs reversed-meaning | **Empirical as a boundary map** (jev-capability-atlas history suite N=3; `notes.md` §49). Qualitative, not a knowledge-breadth estimate | | Experimental design: combinatorial negative | Cell-wise Choice assembly of a grid vs extractive keep/drop | **Empirical as a negative** (ARC-AGI Direct Jev 4/400; `notes.md` §49) | | Experimental design: collab arms | `llm_autonomous` vs `scripted_plus_jev` vs `llm_plus_jev`; Wilson + McNemar; the decision model is **not** a peer arm | **Empirical as a harness shape** (jev-testbench; bake into jevals/Harbor, `notes.md` §48) | +| Experimental design: Harbor on/off routing | Same coding-agent task with routing on vs off; hidden verifier; cheaper unsolved is not a saving | **Empirical as a *shape* and one-run signal** (jev-gateway-bench; `notes.md` §51). Pair CI merge-gate with Harbor + rh-guard | | Discrete math: width vs depth | Fan out in width (parallel ≈ free), pay depth linearly; two-stage only when next options depend on an earlier answer | **Empirical recipe** (fan-out: 12.2× cheaper, 10× faster) | -| Psychology: Kahneman | System 2 generates/proposes (LLM), System 1 discriminates (Jev); never the reverse. S1 keeps control; optional S2 is one-use advice and does not fly. Productized cascade: `conf ≥ τ` → S1 decides else S2 writes; routing fails open / safety fails closed; **routing accuracy unmeasured**; keyword fallback ≠ S1 | **Empirical recipe** (mcts-agent role split; 2026-09-18 mixed-architecture discourse; jev-reflex-autonomy-lab as a control-loop shape, `notes.md` §46; [dual-process-ai](https://github.com/taro1985/dual-process-ai) as a business/life cascade, `notes.md` §49). **Route ≠ memory:** a cheap intent gate skips memory/tool *tours* on easy routes; memory still writes; complex still searches (jev-hermes, `notes.md` §48) | +| Psychology: Kahneman | System 2 generates/proposes (LLM), System 1 discriminates (Jev); never the reverse. S1 keeps control; optional S2 is one-use advice and does not fly. Productized cascade: `conf ≥ τ` → S1 decides else S2 writes; routing fails open / safety fails closed; **routing accuracy unmeasured**; keyword fallback ≠ S1. S1 specialists + S2 coordinator is the same split (reification-labs/foreman is description-only Phoenix scaffold this pass — do not invent an Elixir API). Indexer: S1 GLiNER extract on the bulk, escalate LLM on the tail (10–50× unfilled) | **Empirical recipe** (mcts-agent role split; 2026-09-18 mixed-architecture discourse; jev-reflex-autonomy-lab as a control-loop shape, `notes.md` §46; [dual-process-ai](https://github.com/taro1985/dual-process-ai) as a business/life cascade, `notes.md` §49; s1-graphify-indexer, `notes.md` §51). **Route ≠ memory:** a cheap intent gate skips memory/tool *tours* on easy routes; memory still writes; complex still searches (jev-hermes, `notes.md` §48) | | IR / cascades | Cheap relevance / irrelevance before an expensive ranker or generator | **Empirical recipe** (RAG cookbook; jevprune; git-jev-stage; LlamaIndex Jev rerank) | | IR / extractive keep-drop | Number candidates in code; model selects; copy verbatim; `redecide` thresholds on the log with no new calls. Compaction: same pointer job on tool results (Jev Noul/Score *or* GLiNER encoder) | **Empirical recipe** (testimonial-miner; jev-reviewer pointer-not-generator, `notes.md` §48; gliner25-compaction char-offset + fail-closed keep_full, `notes.md` §50) | -| Spec / lint | Project-defined semantic rules as predicates over a diff; linter owns hard rules. AST remainder: Tree-sitter units, then typed questions; do not execute scanned code | **Empirical recipe** (jev-pref contract; Abide productized path — replay 93 sessions, edit precision ~26% / turn ~73% before tune, `notes.md` §47; JevLint file-level Noul; pi-warden; snifftest unsure-band; jevscan AST∩semantic, `tenbin` owns the lint skill, `notes.md` §48) | +| Spec / lint | Project-defined semantic rules as predicates over a diff; linter owns hard rules. AST remainder: Tree-sitter units, then typed questions; do not execute scanned code. Plain-English PR check: one condition + min-confidence; fail-closed on error | **Empirical recipe** (jev-pref contract; Abide productized path — replay 93 sessions, edit precision ~26% / turn ~73% before tune, `notes.md` §47; JevLint file-level Noul; pi-warden; snifftest unsure-band; jevscan AST∩semantic, `tenbin` owns the lint skill, `notes.md` §48; if-ai, `notes.md` §51). jev-marshal is Watch / empty this pass | | Formal methods / DST / safety | Judgment triages counterexamples, failing seeds, and named-rule conformance; proof/MC/DST stay with their tools. Alloy finder ≠ Apalache BMC ≠ Quint run. DST trio: Antithesis hypervisor / Resonate HQ Lean+oracle+SDK (durable async) / PufferLib env+seed. Noul is a sensor, not a discharged PO. Semi-formal diagrams are vocabularies, not enforcers | **Hypothesis as product**, **Contract** as ownership (matching `mappings.md` §8 and `methods-catalog.md`; worked shape pi-warden — `formal-methods.md`, `formal-semi-formal.md`) | -| Decision analysis: VOI | Gather as an enumerated act; pay iff expected decision-loss drop > cost | **Hypothesis** as calculator (`mappings.md` §6; `mental-models.md`) | +| Decision analysis: VOI | Gather as an enumerated act; pay iff expected decision-loss drop > cost. Wake/resume: skip the LLM turn only if the judge answers and p is low (Horvitz); user-message / skip-limit already answer without a model | **Hypothesis** as calculator (`mappings.md` §6; `mental-models.md`). wakegate 21/21 is smoke (`notes.md` §51) | | Signal detection | Noul as evidence variable; criterion from costs and base rate; ROC/PR on your labels | **Hypothesis** for non-SWE plots (`mappings.md` §7) | | Safety engineering: STPA | Sensor ≠ constraint; table of unsafe control actions if the sensor lies | **Contract** as ownership (`mappings.md` §8; Leveson) | | Bandits / RL | Value from observed rewards only — Jev provides none; rejected without an environment. PufferLib Ocean is a trainer contract, not a baseline | **Rejected** (standing boundary) | diff --git a/.agents/skills/augustus/references/validation.md b/.agents/skills/augustus/references/validation.md index 6a706f6..fa5d573 100644 --- a/.agents/skills/augustus/references/validation.md +++ b/.agents/skills/augustus/references/validation.md @@ -325,6 +325,7 @@ Rules: | Reward-hack / eval gaming | [rh-guard](https://github.com/24601/rh-guard) | structural deny + System One sidecar | | Project soft-rule lint | [Abide](https://github.com/coldteadotai/abide) | Score per rule on the diff; bands; fail-open; replay + independent review | | Collab / computer-use product loop | [jev-testbench](https://github.com/ufx7/jev-testbench); [solari-reflex](https://github.com/hitakshiA/solari-reflex) | Wilson/McNemar arms; independently checked task time | +| Agent routing on vs off | [jev-gateway-bench](https://github.com/vinilana/jev-gateway-bench) | Hidden perft; cost/quality; one-run signal this pass | rh-guard is a reward-hack hook, a different surface from jevgate and from Abide (eval-integrity vs allowlist-remainder vs project soft @@ -348,6 +349,22 @@ vs 98.4 s (`notes.md` §48). **Collab-arm curriculum:** McNemar; Jev is not a peer arm. Bake into jevals/Harbor hygiene, do not copy the harness. +**Harbor on/off routing (Empirical as a *shape* and as a one-run +signal, not a measurement; 2026-09-18 ~16:48).** +[jev-gateway-bench](https://github.com/vinilana/jev-gateway-bench) +(MIT): real coding agents on chess-engine tasks; Jev routing on vs +off; hidden perft verifier the agent never sees; fresh gateway per +run. Author: first signal, not a measurement; <5 runs/mode the +summary says so. Preliminary `chess-bugfix` Codex: both 36/36 +checks; on 4 LLM req / 76,678 in / 35 s vs off 6 / 118,709 / 88 s; +Jev 4 calls ~$0.0008. A cheaper unsolved run is not a saving. A +wrongly forced tool can derail a turn. Product sibling +[jev-gateway](https://github.com/vinilana/jev-gateway) fails open if +Jev is down. Pair CI merge-gate +([latch](https://github.com/CaseReed/latch)) with this substrate +(frozen JUnit artifacts × PASS/BLOCK) and rh-guard (eval-integrity). +Do not copy npm/ports (`notes.md` §51). + **Harbor-style frozen protocol vs constrained LLMs (Empirical as that named receipt, not a ranking).** [`nibzard/decision-model-benchmark`](https://github.com/nibzard/decision-model-benchmark) diff --git a/CHANGELOG.md b/CHANGELOG.md index 35aaa48..755ec74 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -210,6 +210,22 @@ Pegged against [`typesafe-ai/skills` v0.5.7](https://github.com/typesafe-ai/skil *job* as fast-jev-compaction / pi-jev-compaction; GLiNER encoder backend; Fastino/GLiGuard sibling class. `shadowMode` default true. Not Jev. Not multimodal. No invented metrics. +- CI merge-gate / fail-open wake VOI / S1 indexer / claim-evidence + (`research/notes.md` §51): architecture notes, not a plugin how-to. + [latch](https://github.com/CaseReed/latch) cluster-then-policy + PASS/BLOCK (pair Harbor + rh-guard). + [wakegate](https://github.com/shitianfang/wakegate) skip only if + p(wake)<0.2 (21/21 smoke). s1-graphify-indexer GLiNER extract + + escalate-S2 (10–50× unfilled). + [clear-head](https://github.com/VladyslavHontar/clear-head) + claims vs session evidence. + reification-labs/foreman description-only Phoenix scaffold (not the + super-jev loop). + [jev-gateway-bench](https://github.com/vinilana/jev-gateway-bench) + Harbor on/off one-run signal. + jev-marshal Watch/empty; jevons bounded Pi supervisor (shadow + recovery). MED: if-ai, omp-auto-mode, downloads-sorter, label-desk, + herdr-jev. Archer still Watch. No invented metrics. No wrapper. - Effect-oriented loops (`notes.md` §28, `mappings.md` §19): Ward's ZIO client keeps Jev as the outer Choice and the handler as the effect. Not Effect.ts. GLiNER author: GLiNER2 "like jev" is GLiGuard diff --git a/README.md b/README.md index f4e4283..44ce414 100644 --- a/README.md +++ b/README.md @@ -35,7 +35,8 @@ never launder a Noul as a proof. open multimodal RLCD (blackwood-rlcd; not Archer), Laya ONNX port, contract-compatible local `/v1/systemone` (stub until hf scorer; also kev pointer / von tiny SAN — not replicas), GLiNER/GLiClass species (locate vs categorize vs local multi-head; - GLiNER2.5 extractive compaction as a named job, not a new species), + GLiNER2.5 extractive compaction as a named job, not a new species; + GLiNER code-graph indexer + escalate-S2, 10–50× unfilled), listwise vs decision objectives, vision scoring, when-to-use axes (including decision-model vs constrained LLM), agent-architecture portents @@ -51,12 +52,13 @@ never launder a Noul as a proof. - `.agents/skills/augustus/references/mixed-architecture.md` — default placement: judgment-class model + LLM + code; preference lint; provider (Jev default / other family with self-eval); dual-process S1 decide / S2 - generate; component node; DOM-as-text + fan-out; shadow-mode compaction rollout + generate; component node; DOM-as-text + fan-out; shadow-mode compaction rollout; + fail-open wake vs fail-closed merge-gate; Harbor on/off routing - `.agents/skills/augustus/references/applied-mappings.md` — context sieve, - exact-text keep/drop (extractive / pointer-not-generator; char-offset compaction), env triage, moderation/ranking, skill routing (route ≠ memory) + exact-text keep/drop (extractive / pointer-not-generator; char-offset compaction), env triage (OpenSmoke + latch merge-gate), moderation/ranking, skill routing (route ≠ memory) - `.agents/skills/augustus/references/faq.md` — "just classification", stack replacement, Jev vs open head vs encoder vs LoRA vs constrained AR vs kev vs blackwood, - wait-for-Archer, missing-other confident-wrong, soft project rules vs linter (Abide), extractive/pointer-not-generator, compaction summarize vs pointer, encoder vs Jev compaction, fail-closed keep_full, shadow-mode rollout, local drop-in vs stub scorer, route ≠ memory, when-it-holds / extractable-from-state, decision-model vs constrained LLM, dual-process S1/S2, combinatorial grid ≠ extractive, GLiNER vs GLiClass vs CLIP, LLM-as-judge, in-engine vs CLI store, + wait-for-Archer, missing-other confident-wrong, soft project rules vs linter (Abide), extractive/pointer-not-generator, compaction summarize vs pointer, encoder vs Jev compaction, fail-closed keep_full, shadow-mode rollout, fail-open vs fail-closed wake vs CI gate, local drop-in vs stub scorer, route ≠ memory, when-it-holds / extractable-from-state, decision-model vs constrained LLM, dual-process S1/S2, combinatorial grid ≠ extractive, GLiNER vs GLiClass vs CLIP, LLM-as-judge, in-engine vs CLI store, hard envelope (bitrate / planner), not-another-how-to - `.agents/skills/augustus/references/mappings.md` — classical-method mappings with boundaries, counterexamples, acceptance tests (including @@ -69,7 +71,8 @@ never launder a Noul as a proof. CC-BY-4.0 recompute-from-logs feedstock; Abide replay as Harbor-adjacent soft-rule measurement; solari-reflex Harbor-style computer-use; jev-testbench collab arms; ARC-AGI Direct Jev as - combinatorial-≠-extractive negative) + combinatorial-≠-extractive negative; jev-gateway-bench Harbor on/off + routing one-run signal) - `.agents/skills/augustus/references/boundary-audit.md` — existing-system insertion: fit test, opportunity map, smallest boundary, red flags - `.agents/skills/augustus/scripts/evaluate_decisions.py` — offline evaluator diff --git a/docs/ecosystem.md b/docs/ecosystem.md index dd6719b..9ebe5eb 100644 --- a/docs/ecosystem.md +++ b/docs/ecosystem.md @@ -159,6 +159,20 @@ Architecture notes, not a plugin catalog. `notes.md` §50. TypeSafe Jev is the e - **m-newhauser/gliner25-compaction** — local GLiNER2.5 (`fastino/gliner2.5-base-v1`) retention Choice (`keep_full` / `keep_evidence` / `keep_call_only` / `drop`) + exact character-offset copies. Pointer family with testimonial-miner / jev-reviewer. Mutating tools / shell operators → `keep_full` in code. Fail-closed to `keep_full` (contrast many fail-open Jev gates). Same compaction *job* as fast-jev-compaction / pi-jev-compaction; encoder backend; Fastino/GLiGuard sibling class. `shadowMode` default true. Experimental; characters not tokens; no published retention-quality rates. Apache-2.0. +### Hourly ~16:48 Boise (CI merge-gate / fail-open wake / S1 indexer / claim-evidence) + +Architecture notes, not a plugin catalog. `notes.md` §51. TypeSafe Jev is the exemplar, not a monopoly. Archer still Watch. + +- **CaseReed/latch** — merge-gate: cluster in code, Jev labels cause, policy owns Gate PASS (infra) vs BLOCK (real). Judge never says ignore alone. Playwright reporter fail-open; `--gate` is a separate step. Pair Harbor + rh-guard. MIT. +- **shitianfang/wakegate** — fail-open VOI wake/resume (Horvitz). Skip only if Jev answers and p(wake)<0.2. 21/21 smoke (same author wrote scenarios+question). Contrast pi-jev-approver fail-closed / jevgate cannot-block. MIT. +- **GreyssonEnterprises/s1-graphify-indexer** (+ `s1-indexer`) — GLiNER default code-graph; escalate LLM only if backend loaded and low conf. 10–50× **unfilled**. Query does not invent edges. License not on GitHub this pass. +- **VladyslavHontar/clear-head** — Stop hook: claims vs session evidence. Keyword retriever; `JEV_FIRM` 0.6 never blocks below. MIT. 1★. +- **reification-labs/foreman** — description-only Phoenix scaffold (S1 specialists + S2 coordinator). Not the super-jev "foreman" loop. No Jev dep. Do not invent an Elixir API. +- **vinilana/jev-gateway-bench** — Harbor on/off routing; hidden chess perft; one-run signal (36/36 both; 4 vs 6 LLM req). Product sibling `jev-gateway` fail-open if Jev down. MIT. +- **LightningK0ala/jev-marshal** — Watch / empty repo. Policy-as-judgment PR cousin of Abide / if-ai. +- **LilDojd/jevons** — bounded Pi supervisor; shadow recovery; never generates commands. MIT. +- MED: if-ai (plain-English PR checks, fail-closed on error); omp-auto-mode (safe/unsafe/ask); jev-downloads-sorter (device-loop Choice); jev-label-desk (description-only); herdr-jev (~260 ms triage + triad; no-key heuristic). + See `references/mixed-architecture.md` in the skill. Class-level family choice: `references/judgment-class.md`. Proof vs judgment (Alloy vs Apalache; DST trio Antithesis / Resonate HQ / PufferLib): diff --git a/research/archive/findings.md b/research/archive/findings.md index 784772e..6c5685c 100644 --- a/research/archive/findings.md +++ b/research/archive/findings.md @@ -919,6 +919,56 @@ mutation. +## Batch #35 (2026-09-18 ~16:48 Boise) — CI merge-gate, fail-open wake, S1 indexer, claim-evidence + +Note: `research/notes.md` §51. Docs-only. Folded into PR #2. Not a +second PR. Archer still Watch. No invented metrics. Do not re-fold +§50 compaction, §49 bake-off, Abide, kev, pi-jev-approver, jevgate, +rh-guard species. + +- **CaseReed/latch (Empirical as README / offline demo).** MIT. + Created 2026-09-18T22:09:12Z; 0★. Cluster in code; Jev labels + cause; policy owns Gate PASS vs BLOCK. Judge never says ignore + alone. Reporter never fails Playwright; `--gate` is a separate + step. Demo: 8 connection errors → PASS; 5 assertions → BLOCK. + Message-based grouping fragments (`pallets/click` 13→10). Pair + Harbor + rh-guard. +- **shitianfang/wakegate (Empirical as safety table; 21/21 smoke).** + MIT. Skip only if Jev answers and p(wake)<0.2. Same author wrote + scenarios+question. Savings unmeasured. Contrast fail-closed + pi-jev-approver / cannot-block jevgate / fail-closed compaction. +- **GreyssonEnterprises/s1-graphify-indexer (+ s1-indexer).** GLiNER2 + default; escalate LLM only if backend loaded. 10–50× **unfilled**. + Degraded file-node graph if GLiNER cannot load; query does not + invent edges. License not on GitHub this pass. +- **VladyslavHontar/clear-head (Empirical as README).** MIT. 1★. + Stop hook: claims vs session lines. Keyword retriever; JEV_FIRM + 0.6 never blocks below. +- **reification-labs/foreman.** Description-only Phoenix scaffold. + No Jev dep. Distinguish from super-jev "foreman" loop. +- **vinilana/jev-gateway-bench (Empirical as *shape* + one-run + signal).** MIT. Chess perft hidden verifier; on vs off. Both + 36/36; 4 vs 6 LLM req. Author: not a measurement. Sibling + jev-gateway fail-open if Jev down. +- **LightningK0ala/jev-marshal.** Empty repo. Watch / description. +- **LilDojd/jevons (Empirical as README policy).** MIT. Bounded Pi + supervisor; shadow recovery; never generates commands. +- **MED:** Victor-Casado/if-ai (plain-English PR checks, fail-closed + on error); alexsatch/omp-auto-mode (one-line README); jolehuit/ + jev-downloads-sorter (device-loop Choice); LakshyaChaudhry/ + jev-label-desk (empty README); flaviomartil/herdr-jev (~260 ms + triage + triad; no-key heuristic). + +Cross-repo addition: (by) cluster in code / judge labels / policy +decides (CI flaky-vs-real); (bz) name the irreversible act before +picking fail polarity (wake skip vs merge PASS vs compaction drop); +(ca) Harbor on/off routing with a hidden verifier; (cb) S1 extract ++ escalate-S2 is not a measured 10–50× until the table is filled; +(cc) claim/evidence Stop is anti-hallucinated-done, not a test +runner; (cd) description-only greenfield is not a product receipt. + + + diff --git a/research/notes.md b/research/notes.md index 20b1da8..14dfdfe 100644 --- a/research/notes.md +++ b/research/notes.md @@ -3403,3 +3403,309 @@ GLiNER2.5 is another. Augustus stays family-first. Not multimodal. Usage: extractive context management as a perception/memory hygiene stage **before** decide. Archive + landscape pointer. + +## 51. CI merge-gate, fail-open wake VOI, S1 indexer, claim-evidence Stop (2026-09-18 ~16:48 Boise) + +America/Boise ~16:48 = 22:48 UTC. Docs-only fold into open PR #2 +(`cursor/augustus-store-envelope-00b4`). Not a competing PR. Archer +27B drop still **WATCH**. Identity lock vs `typesafe-ai` / `tenbin` / +`decision-first` holds. No wrapper, no npm/wrangler/install.sh/devenv +how-to, no copied ports, thresholds as class constants, or invented +metrics. Do not re-fold §50 GLiNER2.5 compaction, §49 bake-off, Abide, +kev, pi-jev-approver, jevgate, or rh-guard as a species rewrite. + +Backend-agnostic: this hour is **decision-model as classifier / +gate / coordinator** (flaky-vs-real CI, VOI resume, S1 extract + +escalate-S2, claim/evidence integrity, Harbor on/off routing, +policy-as-judgment). TypeSafe Jev is the documented exemplar, not +the monopoly. Augustus stays family-first. + +### HIGH + +1. **[`CaseReed/latch`](https://github.com/CaseReed/latch)** + (MIT; TypeScript; created 2026-09-18T22:09:12Z; 0★ at capture). + Merge-gate triage for a **finished** red test run (Playwright / + Jest / pytest / JUnit XML): cluster failures by signature in + **code**, TypeSafe Jev labels each cause (one call per cluster, at + most 8; cached free), **code owns** Gate: PASS (infra noise) vs + Gate: BLOCK (real failure). The judge never gets to say "ignore" + alone. `ignore_as_infra` needs `env_cascade` + an infra + fingerprint. Missing key still prints clusters (`needs_human` / + `no_key`) and **never fails Playwright** — the reporter is + fail-open; `--gate` is a **separate** CI step. + + Demo (offline, README): 8 identical connection errors → 1 cause → + Gate PASS; 5 failing assertions → Gate BLOCK. Policy (theirs, not + class constants): no key / API error / `cause.confidence < 0.55` / + `same_root < 0.5` → `needs_human`; `env_cascade` + `same_root >= + 0.7` + fingerprint → `ignore_as_infra`; flake / locator_drift → + `fix_test`; `assertion_bug` and (`blocks_merge >= 0.55` or Jev + `action = fix_product`) → `fix_product`. `blocks_merge` sits in a + noise band ~±0.03; 0.55 is in the empty gap (their calibrate + claim). + + **Limits (theirs).** Grouping is message-based; a logic regression + fragments. Measured on `pallets/click`: 2 real regressions → 13 + failures → **10 clusters**. The "70 → 1" figure is an infra-cascade + property, not a general one. Signature = `apiName` + first 80 chars + of the normalized error. Error text is redacted in every output. + + **Four load-bearing mental models:** + + 1. **Cluster in code, judge labels, policy decides.** Same split as + OpenSmoke (scan every step; LLM autopsy only on flags): the + model estimates a *cause class*; the merge act is a table. + Internals are not a state machine; placement is a component + node (`notes.md` §49). + 2. **Decision-model as CI flaky-vs-real classifier.** The hole is + SDT criterion (false PASS on a real bug >> false BLOCK on + infra), not "make CI smarter." Pair with Harbor (frozen + artifacts × PASS/BLOCK labels; independent verifier) and + rh-guard (eval-integrity — do not let the agent game the tests + latch is classifying). Different holes; shared notes only. + 3. **Fail polarity is per surface.** Reporter never fails + Playwright (false block of the *run* loses evidence). `--gate` + fails closed on merge when a cluster is not confirmed infra. + Name the *act*. + 4. **Do not copy the reporter.** Thresholds, ledger path, and npm + wiring stay theirs. + + **Placement.** Environment / harness triage + (`applied-mappings.md` §3) + decision circuits (`mappings.md` §3) + + SDT (`mappings.md` §7). Pillar: SDT + Leveson sensor≠constraint. + Hole: triage / gate. Family: closed decision API. Eval path: none + published as Harbor this pass (demo + unit goldens). **Empirical** + as README behavior / demo. **Hypothesis** that the same cluster → + label → policy transfers to *your* runner. Cards: + `applied-mappings.md` §3; `mappings.md` §3 / §7 / §18; + `mixed-architecture.md`; `validation.md`; `faq.md`. No wrapper. + +2. **[`shitianfang/wakegate`](https://github.com/shitianfang/wakegate)** + (MIT; TypeScript; created 2026-09-18T22:13:11Z; 0★). Fail-open + **wake gate** before resuming a sleeping agent (Workers / Durable + Objects / Node). One typed question: given `waitingFor` plus an + optional event/observation, is this worth a full LLM turn? Code + owns sleep duration, skip counters, and force-wake. Skip only if + Jev answers **and** p(wake) < 0.2. Every other path wakes: + `fromUser`, nothing-to-judge, skip-limit (default 10), error / no + key / timeout (5 s), unsure band 0.2–0.5, p ≥ 0.5. + + **Eval (theirs; discount).** 21/21 on 21 hand-written scenarios + (11 wake / 10 sleep); p50 253 ms, p95 519 ms (n=21). Two of the + 11 correct wakes came only from the unsure band (calendar invite + 0.38; sold out 0.48). Same person wrote the scenarios and the + question. First yes/no wording scored 16/21 on this set. Current + three-way Choice picked on a separate 16-scenario **dev** set + (not in repo). **Smoke, not a benchmark.** Savings unmeasured. + + **Contrast (fail polarity, not products).** + [`phin-tech/pi-jev-approver`](https://github.com/phin-tech/pi-jev-approver) + fails **closed** without a key (tool remainder). jevgate **cannot + block** (allowlist proves; remainder fail-open). Compaction + (`notes.md` §50) fails closed to `keep_full` because *drop* is + irreversible. Wake *skip* is the irreversible act here (the agent + stays asleep), so the authorized skip is the rare, high-confidence + path; everything else wakes. Horvitz mixed-initiative / VOI: pay + for the LLM turn only if EV(decision) beats the token cost; a + regex / user-message / skip-limit already answers without a model + (meta-VOI). Event/observation are untrusted; `maxSkips` bounds + sleep, "neither is a security boundary" (their README). + + **Placement.** VOI / gather (`mappings.md` §6) + durable-agent + control (`mappings.md` §14, Hypothesis) + mixed-architecture + per-action fail table. Hole: gate / abstain. Family: closed + decision API. **Empirical** as README safety table. **Hypothesis** + as production savings. Do not copy wrangler/npm. Cards: + `mappings.md` §6 / §18; `mixed-architecture.md`; `faq.md`. + +3. **[`GreyssonEnterprises/s1-graphify-indexer`](https://github.com/GreyssonEnterprises/s1-graphify-indexer)** + (+ sibling [`s1-indexer`](https://github.com/GreyssonEnterprises/s1-indexer); + created 2026-09-18T22:23:04Z / 22:21:15Z; 0★; **license not on + GitHub this pass — do not invent**). Local-first semantic + indexer: small zero-shot System-1 models build a knowledge graph + of a repo; an LLM is used only on the ambiguous tail, and **only + when the System-1 backend actually loaded**. Default backend + `gliner2`. Stubs `jev` and `needle` ship `available=False` until + implemented. GitHub one-liner "10–50× faster than LLM-based + indexing" is a **target, not a measured speedup**. README: do not + treat it as a benchmark until the table is filled from one repo, + both backends, same machine. Table is TBD. + + If GLiNER2 cannot load: artifacts still written (file nodes from + the chunker, `run_status: degraded`), process exits nonzero, repo + is **not** dumped to `S1_LLM_CMD`. `query` tokenizes, matches node + names, walks two hops; no match prints `Insufficient evidence`; + **it does not invent edges**. `gliner2[local]` extra is heavier + than declared deps and is not pulled by default. + + **Mental model.** S1 zero-shot **locate/extract** (GLiNER spans / + relations) on the bulk; escalate-to-S2 only on low-confidence + remainder — same sandwich as allowlist ∩ remainder and as + dual-process `conf ≥ τ` → S1 else S2, here the expensive act is + *indexing tokens* not a chat reply. GLiNER-as-Jev-class cousin + (`judgment-class.md` locate vs decide): the indexer is not a Noul. + Same *family* as gliner25-compaction (encoder on code/text) with a + different hole (graph construction vs memory keep/drop). Do not + copy pip extras. + + **Placement.** Locate species + escalate-S2. Hole: perceive / + gather. **Hypothesis** as 10–50×. **Empirical** as degraded-load + and no-invent-edges README behavior. Cards: `judgment-class.md`; + `mixed-architecture.md`; `faq.md`. + +4. **[`VladyslavHontar/clear-head`](https://github.com/VladyslavHontar/clear-head)** + (MIT; Python; created 2026-09-18T22:15:08Z; 1★). Claude Code + **Stop** hook: Jev checks factual claims in the assistant's answer + against **what it actually read this session**. Anti-hallucinated- + done. Per claim, keyword-and-frequency retriever (not semantic) + sends matching tool-output *lines* — not whole files. Classifies + sentences as factual / proposal / recap / neither; then + supports / contradicts / doesn't-address. Blocks on contradicted, + or unsupported with **no** relevant evidence. Coverage below + `JEV_EVIDENCE_FLOOR` (default 0.3) is "nothing relevant found"; + above the floor, "Jev can't confirm a specific derived fact" is + usually their documented limit, not a block. `JEV_FIRM` default + 0.6: below that, logged, **never blocks**. + + **Limits (theirs).** Keyword retriever misses paraphrases and can + match stale session evidence on generic overlap (no recency / + topic-boundary). A true claim unread this session still flags + unsupported. Excerpts leave the machine. Do not copy `install.sh`. + + **Placement.** Done-check / claim–evidence entailment + (`agent-self-assessment.md`; `methods-catalog.md` NLI row). + Pointer family with jev-reviewer: the model judges against + retrieved lines, never against another model's prose. Hole: gate. + **Empirical** as README behavior. **Hypothesis** as transfer to + *your* transcript domain. Cards: `agent-self-assessment.md`; + `applied-mappings.md` §2; `faq.md`. + +5. **[`reification-labs/foreman`](https://github.com/reification-labs/foreman)** + (created 2026-09-18T22:46:57Z; 0★; **no license field in + `mix.exs` — do not invent**). GitHub description: "Parallel + specialist agents returning typed, calibrated answers behind a + single System Two foreman. Elixir/Phoenix, Jev/TypeSafe-native — + `{value, probability}` everywhere." **The checkout is a stock + Phoenix 1.8 scaffold.** README is the Phoenix generator starter; + `AGENTS.md` is Phoenix guidelines; `mix.exs` has Phoenix/Ecto/ + Bandit/Req — **no typesafe / jev dependency**. Fold as + **description-only greenfield**, not a measured product. Do not + invent an Elixir Jev API. + + **Distinguish** from the existing "foreman" *shape* in + `agent-self-assessment.md` (Kevthetech143/super-jev loop: + progress/stuck/complete → continue/stop/retry/verify; the model + estimates named probabilities; code owns hysteresis). Dual-process + cousin of [dual-process-ai](https://github.com/taro1985/dual-process-ai) + (`notes.md` §49): S1 specialists decide; S2 coordinates / writes. + Typed probability everywhere is the *class* claim, not a receipt + this pass. + + **Placement.** Mixed architecture / dual-process. **Watch / + description-only.** Cards: `mixed-architecture.md`; + `agent-self-assessment.md`; `mental-models.md`. + +6. **[`vinilana/jev-gateway-bench`](https://github.com/vinilana/jev-gateway-bench)** + (MIT; created 2026-09-18T22:29:58Z; 0★). Harbor-shaped bench for + sibling [`vinilana/jev-gateway`](https://github.com/vinilana/jev-gateway) + (MIT; product; fail-open if Jev down/slow/wrong key; never fails + the LLM request). Real coding agents (Codex / Claude Code) on + chess-engine tasks with Jev routing **on vs off**. Hidden verifier + (perft + targeted checks) the agent never sees. Chess chosen + because perft counts are published and one wrong rule changes + them. Tasks: `chess-engine` (build), `chess-bugfix` (five injected + bugs), `chess-san` (notation feature). Modes alternate; order + swaps between reps. + + **Preliminary one-run (author: first signal, not a measurement; + 2026-09-18; Codex 0.154 / `gpt-6-astra` / `jev-latest`; + `chess-bugfix`):** both 36/36 hidden checks; routing on 4 LLM req / + 76,678 in / 1,313 out / 35 s vs off 6 / 118,709 / 3,231 / 88 s; + Jev 4 calls ~$0.0008. Jev forced `exec` three times (p 0.91 / 0.99 + / 0.94) then switched tools off to answer (0.98). An earlier pair + the same day (pre token-metering fix) showed the same *shape* (4 + vs 6 requests; 40 s vs 69 s). With fewer than five runs per mode + the summary says so. A cheaper unsolved run is not a saving. A + wrongly forced tool can derail a turn. Do not copy npm/ports. + Public repo caveat: an agent with web access could find the + reference; tasks give no reason to look. + + **Placement.** Harbor/jevals practice (`validation.md`): taskset + (chess + hidden verifier) × harness (Codex/Claude) × runtime + (fresh gateway per run) × on/off treatment. Sibling of DMB + (class bake-off) and jev-testbench (collab arms). **Empirical** as + a *shape* and as one-run signal. **Hypothesis** as a cost/quality + claim. Cards: `validation.md`; `methods-catalog.md`; + `toolbox-mapping.md`. + +7. **[`LightningK0ala/jev-marshal`](https://github.com/LightningK0ala/jev-marshal)** + (created 2026-09-18T22:42:24Z). Description: "Repository rules for + pull requests, enforced by Jev." **Empty git repo** this pass + (default-branch 409). Fold as **Watch / description-only**. Cousin + of Abide / jev-pref / if-ai (policy-as-judgment on a PR). No + metrics, no fail polarity, no license to invent. + +8. **[`LilDojd/jevons`](https://github.com/LilDojd/jevons)** + (MIT; TypeScript; created 2026-09-18T22:43:34Z; 0★). Bounded **Pi** + execution supervisor. README 100% slop badge. **Not a second + coding agent:** Jev interprets evidence; ordinary code controls + freshness, limits, and permitted responses. Skill selection (≤3 + or none), failure recovery (actual completed tool outcomes; exact + repeats in code), review (chunk × rule), optional investigation, + verification (select among configured commands; never generates + them). Default recovery **shadow**. Steering (opt-in) delivers + fixed replan/ask-user guidance, **never generated commands**. + Optional pre-tool feedback is off by default. Distinguish from + `phin-tech/pi-jev-approver` (fail-closed remainder) and + `kevinpita/pi-jev-context` (sieve). Do not copy devenv/bun. + + **Placement.** Agent self-supervision lifecycle + (`agent-self-assessment.md`) as a bounded supervisor, not a + planner. Hole: gate / route. **Empirical** as README policy + (shadow default). **Hypothesis** as improved task completion — + "small fixture experiments are useful smoke tests, not evidence" + (theirs). Cards: `agent-self-assessment.md`; `mappings.md` §9; + `mixed-architecture.md`. + +### MED (brief) + +- **[`Victor-Casado/if-ai`](https://github.com/Victor-Casado/if-ai)** + (MIT; created 2026-09-18T22:09:14Z; 0★). Plain-English PR checks: + one condition, a required `min-confidence`, one GitHub Action. + Modes `pr-body` / `diff` / `per-file`. Pass only when true **and** + confidence ≥ threshold. Empty body/diff, timeout, API error → + **fail** (fail-closed on the Action). Two-option Choice (Noul has + no native confidence). PR text is data, not a security boundary. + Cousin of Abide / jev-pref / jev-marshal. Do not copy the Action + YAML. +- **[`alexsatch/omp-auto-mode`](https://github.com/alexsatch/omp-auto-mode)** + (MIT; created 2026-09-18T22:05:41Z; 0★). oh-my-pi plugin. README + is one line: classify tool calls `safe` / `unsafe` / `ask`. + Pre-action gate cousin. Description-thin. +- **[`jolehuit/jev-downloads-sorter`](https://github.com/jolehuit/jev-downloads-sorter)** + (MIT; created 2026-09-18T22:25:30Z; 0★). Device-loop Choice: + launchd `WatchPaths` on `~/Downloads`; one decision per file; + ~400 ms README examples. Closed folder catalog; never invents + folders; never overwrites; OpenRouter unreachable → extension-map + fallback. Do not copy `install.sh`. +- **[`LakshyaChaudhry/jev-label-desk`](https://github.com/LakshyaChaudhry/jev-label-desk)** + (created 2026-09-18T22:49:49Z; 0★). Description: weekend project + using Jev to automate trace/data labeling for a provided taxonomy. + README empty this pass. **Watch / description-only.** +- **[`flaviomartil/herdr-jev`](https://github.com/flaviomartil/herdr-jev)** + (created 2026-09-18T22:23:53Z; 0★; license not stated this pass). + Jev triage ~260 ms + triad orchestration (advisor / implementer / + reviewer). No key → local heuristic, 0 ms. Do not copy + `install.sh` or the model matrix. +- **jev-gateway siblings.** Product + [`vinilana/jev-gateway`](https://github.com/vinilana/jev-gateway) + (fail-open passthrough) + bench above. Claude Code is `hint` mode + (cannot force `tool_choice` with thinking / cache). Do not copy + ports. + +### Omni / Jev-omni + +Not multimodal. Usage: merge-gate / wake / claim-evidence as +text-state decisions **before** a generative turn. Archive + +landscape pointer. Archer still Watch. + diff --git a/research/refresh-log.md b/research/refresh-log.md index 9124a09..ad59f8b 100644 --- a/research/refresh-log.md +++ b/research/refresh-log.md @@ -510,6 +510,34 @@ CHANGELOG, README. - notes.md §50; sources.json; findings.md batch #34. No wrapper. +## 2026-09-18 22:48 UTC — CI merge-gate / fail-open wake / S1 indexer / claim-evidence (~16:48 Boise) + +- Folded into open PR #2 (`cursor/augustus-store-envelope-00b4`). + Docs-only. Not a competing PR. Archer 27B drop still **WATCH**. + No invented metrics. No wrapper. +- HIGH: [CaseReed/latch](https://github.com/CaseReed/latch) (MIT) — + cluster in code, Jev labels, policy PASS/BLOCK; pair Harbor + + rh-guard. [shitianfang/wakegate](https://github.com/shitianfang/wakegate) + (MIT) — fail-open VOI wake; 21/21 smoke. + [s1-graphify-indexer](https://github.com/GreyssonEnterprises/s1-graphify-indexer) + (+ s1-indexer) — GLiNER extract + escalate-S2; 10–50× unfilled. + [clear-head](https://github.com/VladyslavHontar/clear-head) (MIT, + 1★) — claims vs session evidence. + [reification-labs/foreman](https://github.com/reification-labs/foreman) + — description-only Phoenix scaffold. + [jev-gateway-bench](https://github.com/vinilana/jev-gateway-bench) + (MIT) — Harbor on/off one-run signal. + [jev-marshal](https://github.com/LightningK0ala/jev-marshal) — + empty / Watch. [jevons](https://github.com/LilDojd/jevons) (MIT) + — bounded Pi supervisor, shadow recovery. +- MED: if-ai, omp-auto-mode, jev-downloads-sorter, jev-label-desk, + herdr-jev, jev-gateway sibling. +- Cards: SKILL.md, applied-mappings §3, mappings §3/§6/§9/§18, + mixed-architecture, validation, faq, judgment-class, mental-models, + methods-catalog, toolbox, agent-self-assessment, composition-algebra, + ecosystem, CHANGELOG, README. +- notes.md §51; sources.json; findings.md batch #35. No wrapper. + diff --git a/research/sources.json b/research/sources.json index 9c5e553..7087f88 100644 --- a/research/sources.json +++ b/research/sources.json @@ -1,6 +1,6 @@ { "refresh_cadence": "hourly", - "retrieved": "2026-09-18T22:22Z", + "retrieved": "2026-09-18T22:48Z", "sources": [ { "kind": "docs", @@ -1483,6 +1483,96 @@ "title": "m-newhauser/gliner25-compaction", "url": "https://github.com/m-newhauser/gliner25-compaction", "note": "Apache-2.0. Local GLiNER2.5 (fastino/gliner2.5-base-v1) extractive compaction for Claude Code. Retention Choice keep_full|keep_evidence|keep_call_only|drop; exact char-offset copies; mutating tools keep_full; fail-closed to keep_full; shadowMode default true. Not Jev, not a summarizer, not multimodal. notes.md \u00a750." + }, + { + "kind": "github", + "title": "CaseReed/latch", + "url": "https://github.com/CaseReed/latch", + "note": "MIT. Merge-gate: cluster in code, Jev labels cause, policy Gate PASS vs BLOCK. Judge never says ignore alone. Pair Harbor + rh-guard. notes.md \u00a751." + }, + { + "kind": "github", + "title": "shitianfang/wakegate", + "url": "https://github.com/shitianfang/wakegate", + "note": "MIT. Fail-open VOI wake/resume. Skip only if Jev answers and p(wake)<0.2. 21/21 smoke (same author wrote scenarios+question). notes.md \u00a751." + }, + { + "kind": "github", + "title": "GreyssonEnterprises/s1-graphify-indexer", + "url": "https://github.com/GreyssonEnterprises/s1-graphify-indexer", + "note": "GLiNER2 default code-graph; escalate LLM only if backend loaded. 10-50x unfilled. License not on GitHub this pass. notes.md \u00a751." + }, + { + "kind": "github", + "title": "GreyssonEnterprises/s1-indexer", + "url": "https://github.com/GreyssonEnterprises/s1-indexer", + "note": "Sibling of s1-graphify-indexer; same README claims. notes.md \u00a751." + }, + { + "kind": "github", + "title": "VladyslavHontar/clear-head", + "url": "https://github.com/VladyslavHontar/clear-head", + "note": "MIT. 1 star. Claude Code Stop hook: claims vs session evidence. Keyword retriever; JEV_FIRM 0.6 never blocks below. notes.md \u00a751." + }, + { + "kind": "github", + "title": "reification-labs/foreman", + "url": "https://github.com/reification-labs/foreman", + "note": "Description-only Phoenix scaffold claiming S1 specialists + S2 coordinator. mix.exs has no Jev dep. No license field. Distinguish from super-jev foreman loop. notes.md \u00a751." + }, + { + "kind": "github", + "title": "vinilana/jev-gateway-bench", + "url": "https://github.com/vinilana/jev-gateway-bench", + "note": "MIT. Harbor-shaped on/off routing bench; hidden chess perft. One-run signal (36/36 both; 4 vs 6 LLM req). Not a measurement. notes.md \u00a751." + }, + { + "kind": "github", + "title": "vinilana/jev-gateway", + "url": "https://github.com/vinilana/jev-gateway", + "note": "MIT. Product sibling of jev-gateway-bench. Fail-open if Jev down. Do not copy ports. notes.md \u00a751." + }, + { + "kind": "github", + "title": "LightningK0ala/jev-marshal", + "url": "https://github.com/LightningK0ala/jev-marshal", + "note": "Empty git repo this pass. Description: PR rules enforced by Jev. Watch. notes.md \u00a751." + }, + { + "kind": "github", + "title": "LilDojd/jevons", + "url": "https://github.com/LilDojd/jevons", + "note": "MIT. Bounded Pi execution supervisor. Shadow recovery default; never generates commands. Not a second agent. notes.md \u00a751." + }, + { + "kind": "github", + "title": "Victor-Casado/if-ai", + "url": "https://github.com/Victor-Casado/if-ai", + "note": "MIT. Plain-English PR checks: one condition, required min-confidence, one Action. Fail-closed on error/empty/low conf. notes.md \u00a751." + }, + { + "kind": "github", + "title": "alexsatch/omp-auto-mode", + "url": "https://github.com/alexsatch/omp-auto-mode", + "note": "MIT. oh-my-pi plugin. README one line: classify tool calls safe/unsafe/ask. notes.md \u00a751." + }, + { + "kind": "github", + "title": "jolehuit/jev-downloads-sorter", + "url": "https://github.com/jolehuit/jev-downloads-sorter", + "note": "MIT. Device-loop Choice on ~/Downloads. Never invents folders; extension-map fallback. notes.md \u00a751." + }, + { + "kind": "github", + "title": "LakshyaChaudhry/jev-label-desk", + "url": "https://github.com/LakshyaChaudhry/jev-label-desk", + "note": "Description-only weekend trace labeling. README empty this pass. notes.md \u00a751." + }, + { + "kind": "github", + "title": "flaviomartil/herdr-jev", + "url": "https://github.com/flaviomartil/herdr-jev", + "note": "Jev triage ~260ms + triad orchestration. No-key local heuristic. Do not copy install.sh. License not stated this pass. notes.md \u00a751." } ] } From 5225ab883245594e1230c8921bfae48833ffb7fe Mon Sep 17 00:00:00 2001 From: Cursor Agent Date: Fri, 18 Sep 2026 23:07:57 +0000 Subject: [PATCH 10/43] Fold GLiNER2 Ultrafast as encoder observe-score-act backend MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Docs-only: sahibzada-allahyar/gliner2-ultrafast proves observe→score-among-candidates→code-acts is backend-agnostic (Jev Ultrafast ↔ local GLiNER2). Hybrid local decide + remote fill; DONE ≠ verified success. Numbered §52 after the 16:48 Boise watch. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com> --- .agents/skills/augustus/SKILL.md | 12 +- .../references/agent-self-assessment.md | 8 + .../augustus/references/applied-mappings.md | 11 +- .../references/composition-algebra.md | 4 +- .agents/skills/augustus/references/faq.md | 38 +++- .../augustus/references/judgment-class.md | 31 +++- .../skills/augustus/references/mappings.md | 20 +++ .../augustus/references/mental-models.md | 10 +- .../augustus/references/methods-catalog.md | 3 +- .../augustus/references/mixed-architecture.md | 14 +- .../augustus/references/toolbox-mapping.md | 2 +- .../skills/augustus/references/validation.md | 9 +- CHANGELOG.md | 13 ++ README.md | 14 +- docs/ecosystem.md | 10 +- research/archive/findings.md | 35 +++- research/notes.md | 165 ++++++++++++++++++ research/refresh-log.md | 24 +++ research/sources.json | 23 ++- 19 files changed, 409 insertions(+), 37 deletions(-) diff --git a/.agents/skills/augustus/SKILL.md b/.agents/skills/augustus/SKILL.md index a086110..145e213 100644 --- a/.agents/skills/augustus/SKILL.md +++ b/.agents/skills/augustus/SKILL.md @@ -1,6 +1,6 @@ --- name: augustus -description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill, open multimodal RLCD / blackwood), announced open decision-model (Watch — still not landed), constrained-AR (TypeAR, pcdServer, decision-token LoRA), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"wait for Archer vs open multimodal RLCD\", \"screenshot/DOM candidates → typed Choice\", \"eval path\", \"jevals\", \"Harbor taskset\", \"shared bake-off ECE/NLL/Brier\", \"LLM-as-judge is not the System One score\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge (allowlist *proves*; fail-open cannot block), \"missing other → confident wrong Choice\", \"lint the Jev request\", \"training confronts Choice other / none-of-the-above\", \"S1 reflex keeps control / optional S2 one-use advice\", \"soft AGENTS.md rules vs linter\" (Abide / jev-pref), \"edit-phase vs turn-phase observation window\", banded confidence fail-open preference lint, \"extractive selection / pointer-not-generator\", \"encoder GLiNER compaction vs Jev Score compaction (same job; pointer not summarizer)\", \"fail-closed keep_full under mutation envelope\", \"CI flaky-vs-real merge gate\", \"fail-open VOI wake/resume (Horvitz)\", \"claim/evidence Stop integrity\", \"S1 extract + escalate-S2 indexer\", \"Harbor on/off routing\", \"policy-as-judgment PR marshal\", \"shadow-mode compaction rollout\", \"local /v1/systemone drop-in (stub until hf scorer)\", \"dataframe-native semantic columns\", \"route≠memory\", \"advisory sidecar receipts\", \"structure induction over bags\", \"AST ∩ semantic lint\", \"extractable-from-state / retrieve first\", \"decision-model vs constrained-LLM bake-off\", \"dual-process S1 decide / S2 generate\", \"combinatorial grid ≠ extractive\", TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, or \"formally verify with Jev\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." +description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill, open multimodal RLCD / blackwood), announced open decision-model (Watch — still not landed), constrained-AR (TypeAR, pcdServer, decision-token LoRA), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"wait for Archer vs open multimodal RLCD\", \"screenshot/DOM candidates → typed Choice\", \"eval path\", \"jevals\", \"Harbor taskset\", \"shared bake-off ECE/NLL/Brier\", \"LLM-as-judge is not the System One score\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge (allowlist *proves*; fail-open cannot block), \"missing other → confident wrong Choice\", \"lint the Jev request\", \"training confronts Choice other / none-of-the-above\", \"S1 reflex keeps control / optional S2 one-use advice\", \"soft AGENTS.md rules vs linter\" (Abide / jev-pref), \"edit-phase vs turn-phase observation window\", banded confidence fail-open preference lint, \"extractive selection / pointer-not-generator\", \"encoder GLiNER compaction vs Jev Score compaction (same job; pointer not summarizer)\", \"fail-closed keep_full under mutation envelope\", \"CI flaky-vs-real merge gate\", \"fail-open VOI wake/resume (Horvitz)\", \"claim/evidence Stop integrity\", \"S1 extract + escalate-S2 indexer\", \"Harbor on/off routing\", \"policy-as-judgment PR marshal\", \"shadow-mode compaction rollout\", \"Jev Ultrafast vs GLiNER Ultrafast (observe-score-act backend-agnostic)\", \"hybrid local decide + remote fill\", \"DONE ≠ verified success\", \"observed a11y/DOM candidates vs screenshot multimodal\", \"local /v1/systemone drop-in (stub until hf scorer)\", \"dataframe-native semantic columns\", \"route≠memory\", \"advisory sidecar receipts\", \"structure induction over bags\", \"AST ∩ semantic lint\", \"extractable-from-state / retrieve first\", \"decision-model vs constrained-LLM bake-off\", \"dual-process S1 decide / S2 generate\", \"combinatorial grid ≠ extractive\", TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, or \"formally verify with Jev\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." license: MIT metadata: version: 0.3.0 @@ -54,7 +54,7 @@ classical method you already trust, substitute it, classify the win "paraphrase brittleness", "allowlist then judge", "TOCTOU-of-Noul", "Jev inside the database / sqlite-jev", "Jev picks bitrate / join order / the model", "wait for Archer", "lint the request / missing - other", "training confronts Choice other / none-of-the-above", "soft AGENTS.md rules vs the linter", "screenshot Choice / omni System One", "extractive quotes / pointer not generator", "compaction summarize vs pointer", "encoder vs Jev compaction backend", "shadow-mode compaction rollout", "CI flaky-vs-real merge gate", "fail-open VOI wake/resume", "claim vs session evidence", "S1 indexer escalate-S2", "Harbor on/off routing", "fail-open vs fail-closed wake vs CI gate", "local drop-in vs stub scorer", "route vs memory", "when does it hold / extractable from state", "decision model vs constrained LLM", "dual-process S1/S2", or "combinatorial grid vs extractive": + other", "training confronts Choice other / none-of-the-above", "soft AGENTS.md rules vs the linter", "screenshot Choice / omni System One", "extractive quotes / pointer not generator", "compaction summarize vs pointer", "encoder vs Jev compaction backend", "shadow-mode compaction rollout", "CI flaky-vs-real merge gate", "fail-open VOI wake/resume", "claim vs session evidence", "S1 indexer escalate-S2", "Harbor on/off routing", "fail-open vs fail-closed wake vs CI gate", "encoder vs Jev computer-use backend", "hybrid local decide + remote fill", "DONE vs verified success", "local drop-in vs stub scorer", "route vs memory", "when does it hold / extractable from state", "decision model vs constrained LLM", "dual-process S1/S2", or "combinatorial grid vs extractive": read `references/faq.md`, then `references/mental-models.md`, then `references/mixed-architecture.md`, then @@ -114,13 +114,13 @@ classical method you already trust, substitute it, classify the win | Familiar method | Judgment shape | Detail | |---|---|---| | Mental models across domains (not SWE-only) | EU, abstention, VOI, MCDA, SDT, search/control, Leveson, NATM/Norman/snap-fit; **extractable-from-state boundary map** (self-contained vs needs outside knowledge) | `references/mental-models.md` | -| Judgment-model class (Jev is exemplar, not monopoly) | Species: decide / locate (GLiNER) / categorize (GLiClass) / rank / perceive; open heads include encoder DeBERTa, LoRA distill, openjev-lm, kev. Compaction job is backend-agnostic (Jev Score/Noul vs GLiNER2.5 encoder). Indexer cousin: GLiNER extract + escalate-S2 (10–50× unfilled) | `references/judgment-class.md` | +| Judgment-model class (Jev is exemplar, not monopoly) | Species: decide / locate (GLiNER) / categorize (GLiClass) / rank / perceive; open heads include encoder DeBERTa, LoRA distill, openjev-lm, kev. Compaction job is backend-agnostic (Jev Score/Noul vs GLiNER2.5 encoder). Indexer cousin: GLiNER extract + escalate-S2 (10–50× unfilled). Computer-use observe→score-among-candidates→code-acts is backend-agnostic (Jev Ultrafast ↔ GLiNER2 Ultrafast) | `references/judgment-class.md` | | Open weights vs constrained decoding vs proprietary API | Three open paths: encoder open-jev / AR constrained decode (TypeAR + pcdServer; decision-token LoRA; packed one-forward logprob on open LLMs) / trained decision-only (Laya + ONNX port, Nimble, kev, **blackwood-rlcd** multimodal now, Archer Watch still Watch). Local `/v1/systemone` surfaces: jev-local (stub until `hf`), kev (trained pointer), von (tiny SAN — not a replica). Softmax over allowed tokens ≠ Noul | `references/judgment-class.md` (when-to-use table) | | Entropy as allocator (low / medium / high) | Typed low+medium decisions → System One marginals; high-entropy synthesis → frontier decoder. Product rhetoric, not a meter. **Hypothesis** | `references/judgment-class.md` | | Formal / semi-formal (proof vs judgment) | Sensor vs constraint vs searchlight; Alloy vs Apalache; DST trio; TOCTOU-of-Noul, AI×FM | `references/formal-methods.md` (one-screen: `references/formal-semi-formal.md`) | -| Mixed architecture (judgment model + LLM) | Provider judges, LLM writes, code owns control; not a stack replacement. Advisory sidecar never changes host routing. Dual-process: S1 decides, S2 generates (routing accuracy unmeasured); S1 specialists + S2 coordinator is the same split (description-only greenfield this hour). Internals ≠ FSM; placement is a **component node**. Browser-use strength = DOM-as-text + speculative fan-out, not vision. Fail polarity is per act: skip-wake fail-open vs merge-gate BLOCK fail-closed | `references/mixed-architecture.md` | +| Mixed architecture (judgment model + LLM) | Provider judges, LLM writes, code owns control; not a stack replacement. Advisory sidecar never changes host routing. Dual-process: S1 decides, S2 generates (routing accuracy unmeasured); S1 specialists + S2 coordinator is the same split (description-only greenfield this hour). Internals ≠ FSM; placement is a **component node**. Browser-use strength = DOM-as-text + speculative fan-out, not vision. Fail polarity is per act: skip-wake fail-open vs merge-gate BLOCK fail-closed. Hybrid local decide + remote fill; `DONE` ≠ verified success | `references/mixed-architecture.md` | | Context sieve | Relevance Noul per block; always-keep set in code; stub + recall key. Encoder cousin: GLiNER2.5 retention Choice + char-offset spans (gliner25-compaction); fail-closed keep_full; shadowMode default | `references/applied-mappings.md#1-context-sieve` | -| Exact-text keep / drop | Choice include/exclude/mixed over candidates code already holds. Extractive quotes / pointer-not-generator (model never writes the excerpt; char-offset compaction same species) | `references/applied-mappings.md#2-exact-text-keep--drop` | +| Exact-text keep / drop | Choice include/exclude/mixed over candidates code already holds. Extractive quotes / pointer-not-generator (model never writes the excerpt; char-offset compaction same species). Observed a11y/DOM controls: score among them; code clicks (Jev or GLiNER2) | `references/applied-mappings.md#2-exact-text-keep--drop` | | Environment / harness triage | Scan every step for env failure; LLM autopsy only on flags. Merge-gate cousin: cluster in code, judge labels cause, policy owns PASS/BLOCK (latch; judge never says ignore alone) | `references/applied-mappings.md#3-environment--harness-triage` | | Moderation and ranking | Hold-before-publish vs graded rerank; fail policy per action | `references/applied-mappings.md#4-moderation-and-ranking` | | Skill / tool routing | Choice over a closed catalog + whether-anything-fits; code dispatches. Route ≠ memory: cheap intent gate skips memory tours on easy routes | `references/applied-mappings.md#5-skill--tool-routing` | @@ -150,7 +150,7 @@ classical method you already trust, substitute it, classify the win | Effect-oriented state-machine loops | Soft predicates on transitions; code owns the transition | `references/mappings.md#19-effect-oriented-state-machine-loops-hypothesis` (**Hypothesis**; ZIO client, not Effect.ts) | | Agent self-supervision / on-track detection | Pre-gate → output judge → done-check → supervisor nouls. S1 reflex keeps control; optional S2 is one-use advice. Claim/evidence Stop (anti-hallucinated-done); bounded Pi supervisor (shadow recovery, never generates commands) | `references/agent-self-assessment.md` | | Optimizer/program frameworks (Ax, DSPy) | Typed fields → one provider request; judge metrics; threshold discipline. Ax and DSPy climb LM-program knobs only | `references/optimizer-integration.md` | -| Perception → decision pipeline / measure / hill-climb | Stages with a versioned state contract; frozen taskset; DSPy/Ax only on the LM-program slice. Specialist composition stays **Hypothesis**; open multimodal decide (blackwood-rlcd) is a named receipt. Structured observe→decide→verified-act (no screenshots) is a computer-use speed-layer receipt. Same section as the row below | `references/validation.md#eval--hill-climb` | +| Perception → decision pipeline / measure / hill-climb | Stages with a versioned state contract; frozen taskset; DSPy/Ax only on the LM-program slice. Specialist composition stays **Hypothesis**; open multimodal decide (blackwood-rlcd) is a named receipt. Structured observe→decide→verified-act (no screenshots) is a computer-use speed-layer receipt (Jev or GLiNER2 backend). Same section as the row below | `references/validation.md#eval--hill-climb` | | Eval & hill-climb | Decision-stage jevals hygiene; Harbor taskset × harness × runtime; one score-composition table. Shared bake-off exemplar: open-jev-laya-bench (ECE/NLL/Brier; LLM-as-judge is not the score). Harbor-style frozen protocol vs constrained LLMs: DMB (accuracy/calibration/latency/cost; raw logs). Feedstock: jevals-data CC-BY-4.0 boards + JSONL (recompute-from-logs). Collab-arm curriculum: llm_autonomous vs scripted_plus_jev vs llm_plus_jev (Wilson / McNemar). Negative: combinatorial grid assembly ≠ extractive (ARC-AGI Direct Jev 4/400). Harbor on/off routing: chess-engine tasks, hidden perft verifier, one-run preliminary (jev-gateway-bench). Pair CI merge-gate with Harbor + rh-guard | `references/validation.md#eval--hill-climb` | | (meta) Finding new mappings & applications | Toolbox sweep: judgment-shaped component of a known method, substituted + falsified | `references/toolbox-mapping.md` | | Named methods / operators / theorems | Substitution tiers: operand-judgments, preconditioned theorems, non-substitutable | `references/methods-catalog.md` | diff --git a/.agents/skills/augustus/references/agent-self-assessment.md b/.agents/skills/augustus/references/agent-self-assessment.md index dddeaa3..d204bc9 100644 --- a/.agents/skills/augustus/references/agent-self-assessment.md +++ b/.agents/skills/augustus/references/agent-self-assessment.md @@ -28,6 +28,10 @@ relying: foreman, pi-jev, pi-warden, winnow, fast-jev-compaction, jev-judgment. session**; keyword retriever, not semantic; below `JEV_FIRM` 0.6 never blocks; true-but-unread still flags unsupported (`notes.md` §51). Anti-hallucinated-done, not a test runner. + Computer-use cousin of the same honesty: + [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) + — loop `DONE` is termination, not verified success; apps inspect + the actual result (`notes.md` §52). 4. **Stuck-detector**: three failures with the same strategy → ask for a new hypothesis, not another retry. 5. **Supervision during long runs** (foreman): separate concurrent loop @@ -45,6 +49,10 @@ relying: foreman, pi-jev, pi-warden, winnow, fast-jev-compaction, jev-judgment. **no screenshots**; model output never becomes a selector. Harbor-style task score (Stripe API / answer key). Author table vs Codex on Solari: 60.2 s vs 194.9 s; 66 s vs 460 s; 24.2 s vs 98.4 s (`notes.md` §48). + Encoder backend of the same hole: + [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) + — local GLiNER2 scores observed controls; `DONE` ≠ verified success + (`notes.md` §52). Productized Kahneman cascade for *any* cheap-decide / expensive-write loop (business/life, not only SWE): [dual-process-ai](https://github.com/taro1985/dual-process-ai) — diff --git a/.agents/skills/augustus/references/applied-mappings.md b/.agents/skills/augustus/references/applied-mappings.md index 6f36cb7..d570b56 100644 --- a/.agents/skills/augustus/references/applied-mappings.md +++ b/.agents/skills/augustus/references/applied-mappings.md @@ -115,7 +115,16 @@ code copies those bytes; a generator summary is the rejected species (`notes.md` §50). Computer-use cousin: [solari-reflex](https://github.com/hitakshiA/solari-reflex) — structured observation → typed decision → verified act; **no screenshots**; model -output never becomes a selector (`notes.md` §48). +output never becomes a selector (`notes.md` §48). Encoder-backend +cousin of the same hole: +[gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) +— local GLiNER2 (`fastino/gliner2-multi-v1`) scores observed a11y/DOM +controls; code clicks; remote text helper only for TYPE; `DONE` ≠ +verified success (`notes.md` §52). Open-head cousin: +[laya-mind2web](https://huggingface.co/ShaunSpark/laya-mind2web-browser-agent) +— Laya operation + target index over interactive DOM elements (not +screenshot multimodal). Contrast blackwood-rlcd (letters on a +screenshot). **DOM-as-text + fan-out (Empirical as atlas browser-use *shape*):** a screenshot task translated into a structured DOM snapshot as `state`, then speculative questions over numbered candidates — not vision diff --git a/.agents/skills/augustus/references/composition-algebra.md b/.agents/skills/augustus/references/composition-algebra.md index 864cd12..bfc4f0f 100644 --- a/.agents/skills/augustus/references/composition-algebra.md +++ b/.agents/skills/augustus/references/composition-algebra.md @@ -143,7 +143,9 @@ Reusable shapes when generating applications: meeting action items ~150 ms after each utterance. 7. **Formula embedding**: JUDGE/SCORE/CHOOSE as first-class spreadsheet formulas. 8. **Pixel-free computer use**: accessibility tree → compact actionable-JSON → one - batched question set per step → execute via AX actions. + batched question set per step → execute via AX actions. Encoder-backend + cousin: gliner2-ultrafast scores observed a11y/DOM controls with + local GLiNER2; code clicks; `DONE` ≠ success (`notes.md` §52). 9. **Shadow-mode harness** (jev-harness): policy + confidence gate + shadow mode + offline eval CLI replaying fixtures, asserting on actions; 24-row filter 48.9 s (Claude CLI) vs 1.3 s Jev at concurrency 8. Compaction rollout: diff --git a/.agents/skills/augustus/references/faq.md b/.agents/skills/augustus/references/faq.md index aac8986..b10c47d 100644 --- a/.agents/skills/augustus/references/faq.md +++ b/.agents/skills/augustus/references/faq.md @@ -154,6 +154,11 @@ Hole first, logo last. These are **species**, not aliases on the ambiguous tail **if the backend loaded**. GitHub one-liner 10–50× is a **target, not a measured speedup** — table TBD (`notes.md` §51). Not a Noul. + Computer-use receipt: + [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) + uses GLiNER2 (`fastino/gliner2-multi-v1`, not 2.5) to **score among + observed** a11y/DOM controls; code acts; not a screenshot model + (`notes.md` §52). Same observe→score→act hole as Jev Ultrafast. - **GLiClass (categorize):** one forward pass over text + *all* labels; sigmoid multi-label or softmax single-label. Use for large or changing tag sets. Scores are class affinities, not automatically a gateable @@ -195,7 +200,9 @@ A GLiGuard score is not a proof. LLM I/O safety is not a coding-agent tool gate (rh-guard for reward-hacking; jevgate shape for allowlist ∩ remainder; Abide for project soft rules on diffs; gliner25-compaction for extractive context compaction — Fastino -sibling class, not GLiGuard). `judgment-class.md`. +sibling class, not GLiGuard; +gliner2-ultrafast for scoring observed browser controls — GLiNER2, +not 2.5, not a safety schema). `judgment-class.md`. ## Can I threshold CLIP / SigLIP as a safety gate? @@ -391,7 +398,10 @@ Choice). Soft judgment over pixel candidates inside deterministic code. Jev still leads general *text* (0.850 vs 0.786 on their 8,456-item table). Specialist composition (SAM / OCR → text → Jev) remains valid. Do not wait, and do not treat screenshot-vs-Jev-text as the same input. -`judgment-class.md`; `notes.md` §46. +Pixel-free computer-use (DOM/a11y candidates → score → code acts) does +not wait either: [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) +is that hole with a local encoder (`notes.md` §52). +`judgment-class.md`; `notes.md` §46, §52. ## When does a decision model hold? @@ -463,9 +473,12 @@ copies exact source spans; a prose summary of the tool result is generation, not keep/drop (`notes.md` §50). Claim/evidence Stop: [clear-head](https://github.com/VladyslavHontar/clear-head) judges against retrieved session lines, not generated prose (`notes.md` -§51). Generation is only for +§51). Computer-use encoder +cousin: [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) +scores observed controls; code clicks; no generated selectors +(`notes.md` §52). Generation is only for TYPE/prose when something must be written. `applied-mappings.md` §2; -`notes.md` §48, §50. +`notes.md` §48, §50, §52. ## Should compaction summarize? @@ -493,6 +506,23 @@ fails the Action on error / empty / low confidence. jevgate cannot block; pi-jev-approver fails closed without a key; Abide is fail-open on diffs. Same sandwich, opposite authorized act. `notes.md` §50, §51. +## Is observe→score→act Jev-only? + +No. The hole is backend-agnostic: observe controls, score among those +candidates, code acts. [jev-ultrafast](https://github.com/browser-use/jev-ultrafast) +and [solari-reflex](https://github.com/hitakshiA/solari-reflex) use +TypeSafe Jev; [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) +uses local GLiNER2 (`fastino/gliner2-multi-v1`); +[laya-mind2web](https://huggingface.co/ShaunSpark/laya-mind2web-browser-agent) +uses a Laya head over DOM element indices. Same lesson as compaction +(Jev Noul/Score vs GLiNER2.5). Screenshot multimodal (blackwood-rlcd: +letters on an image) is a **different input**, not a better version of +this hole. Hybrid local decide + remote fill is mixed-architecture +economics, not dual-process-ai. `DONE` is loop termination, not +verified success. Not GLiNER2.5. Not a bake-off against the Flights +demo clock. `judgment-class.md`; `mixed-architecture.md`; +`notes.md` §52. + ## Is routing the same as memory? No. A cheap intent gate can skip a memory/tool *tour* on diff --git a/.agents/skills/augustus/references/judgment-class.md b/.agents/skills/augustus/references/judgment-class.md index 36a4e67..ecd4f23 100644 --- a/.agents/skills/augustus/references/judgment-class.md +++ b/.agents/skills/augustus/references/judgment-class.md @@ -111,6 +111,25 @@ below, next to the when-to-use table. one-liner 10–50× is **unfilled** — not Empirical (`notes.md` §51). Same locate family as compaction; different hole (index vs keep/drop). License not on GitHub this pass. + **Computer-use selection (Empirical as README / architecture + behavior, 2026-09-18 ~16:56):** + [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) + (MIT) is a *different named job* on the GLi\* encoder family, not + a new species and **not GLiNER2.5**. Checkpoint + `fastino/gliner2-multi-v1`. Adaptation of + [jev-ultrafast](https://github.com/browser-use/jev-ultrafast): + local GLiNER2 extracts requirements and **scores observed** + a11y/DOM controls; code clicks. **No screenshots. No generated + selectors.** Same observe→score-among-candidates→code-acts hole + as jev-ultrafast / solari-reflex (Jev backends) and + [laya-mind2web](https://huggingface.co/ShaunSpark/laya-mind2web-browser-agent) + (Laya over DOM element indices). Contrast blackwood-rlcd + (screenshot + marked letters → Choice). Hybrid: local decide; + remote text helper only for TYPE. `DONE` is loop termination, not + verified success. Inspector scores are not calibrated P(task + success). Their Flights demo (12.20 s / 13.785 s / ~$0.0001 API) + is a demonstration, not a bake-off (`notes.md` §52). Do not copy + `uv` / `.env`. - **Decide.** Typed Choice/Score/Noul with a decision/proper-scoring objective. That is Jev's product claim. Open heads copy the *shape*; distillation copies the *teacher* (openjev-lm, jev-gate-student-b). @@ -155,8 +174,11 @@ below, next to the when-to-use table. [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) is extractive context compaction (locate + categorize on tool transcripts; Fastino sibling class, not a GLiGuard clone) - (`notes.md` §50). Different holes. - Do not point one model at both, and do not copy a hook install here. + (`notes.md` §50). + [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) + is browser computer-use scoring among observed controls (GLiNER2 + `gliner2-multi-v1`, not 2.5; `notes.md` §52). Different holes. + Do not point one model at every hole, and do not copy a hook install here. **Aggregation is policy-in-code, already taught.** The README's benchmark rule ORs unsafe / non-benign prompt labels and lets refusal @@ -241,7 +263,10 @@ parser. RAM / AX tree / object JSON → closed action or region set → Choice. Prices and dates stay in code. Launch-week recipes: typesafe-mario, jev-drone (classical CV → symbols, Jev advisory), lizard-agent - (visible elements only). The model never sees a screenshot. + (visible elements only). Encoder-backend cousin: + [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) + scores observed a11y/DOM controls with local GLiNER2; code clicks + (`notes.md` §52). The model never sees a screenshot. 2. **Region / label Choice over extracted boxes.** Perception (detector, grid, SAM, OCR boxes) proposes candidates; a scorer picks. `hr98w/jev-visual`: diff --git a/.agents/skills/augustus/references/mappings.md b/.agents/skills/augustus/references/mappings.md index cc642f3..958c094 100644 --- a/.agents/skills/augustus/references/mappings.md +++ b/.agents/skills/augustus/references/mappings.md @@ -431,6 +431,17 @@ immediate win missed once reversed; Fool's-mate confidence 31%/37% so a hole on the constrained-AR surface. Not a strength rating. `notes.md` §42; `validation.md`. +**Computer-use observe → score → act (Empirical as README / +architecture behavior, 2026-09-18 ~16:56):** +[gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) +— the *algorithm* is the browser loop; the substituted classifier +step is scoring among observed a11y/DOM controls. Local GLiNER2 is +one backend; Jev Ultrafast / solari-reflex are the Jev backends of +the same hole. Code owns actuators, dates, freshness. `DONE` is not +the probe — application verifiers are. Contrast blackwood-rlcd +(screenshot input). Hybrid remote TYPE is generation, not the +classifier step (`notes.md` §52). + **Structure induction over a bag (Empirical as a *shape*, 2026-09-18):** [`Joymfl/dag-jev`](https://github.com/Joymfl/dag-jev) — unordered items in, pairwise "does i depend on j?" judgments, DAG in `petgraph`. Code @@ -538,6 +549,15 @@ to `keep_full`. Soft judgment inside a hard envelope, encoder backend — not a Jev Score and not a summarizer (`notes.md` §50). Same sandwich shape as bitrate-advisor; different family. +**Named computer-use envelope (Empirical as README / architecture, +2026-09-18 ~16:56):** +[gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) +— the monitor is **code** (resolve to an observed node; freshness / +visibility / disabled / occlusion; no generated selectors or JS). +GLiNER2 may only pick among candidates the snapshot already holds. +`DONE` is not the monitor. Soft judgment inside a hard envelope, +encoder backend — not a screenshot VLM (`notes.md` §52). + ## 13. DST multiverse triage (Hypothesis) **Method**: Antithesis / Resonate DST artifacts → failure taxonomy → diff --git a/.agents/skills/augustus/references/mental-models.md b/.agents/skills/augustus/references/mental-models.md index d865c4b..6e326c5 100644 --- a/.agents/skills/augustus/references/mental-models.md +++ b/.agents/skills/augustus/references/mental-models.md @@ -150,8 +150,10 @@ reliability on *your* labels before you threshold. **Browser-use is this axis, not vision.** Strength = DOM-as-text + speculative fan-out over candidates code already numbered — a visual task translated into extractive text. Not screenshots. Same -component-node placement as lizard-agent / solari-reflex -(`applied-mappings.md` §2; `mixed-architecture.md`). +component-node placement as lizard-agent / solari-reflex / +gliner2-ultrafast (GLiNER2 encoder backend of the same hole; +`notes.md` §52). Contrast blackwood-rlcd (screenshot + marked +letters). (`applied-mappings.md` §2; `mixed-architecture.md`). **Does not:** merge Banking77 87% (atlas/jev-benchmarks) with DMB 76.3% or jevals.com 79.67% into one ranking — protocol / n / split @@ -465,13 +467,13 @@ Use these as *existence proofs of a position*. Write your own card. | Knowledge work | extract a quote / a cited fact | per-sentence or per-line-id Noul/Choice (**Empirical**: testimonial-miner, jev-reviewer) | verbatim join; place; human publish permission | | Agent context | compact completed tool results without inventing prose | retention Choice + char-offset locate (**Empirical**: gliner25-compaction; same *job* as fast-jev-compaction / pi-jev-compaction) | mutation/shell envelope → keep_full; fail-closed keep_full; shadowMode before replace; copy exact bytes | | Dataframe labeling | classify / score rows | Noul/Choice/Score + full `p__` (**Empirical** as jevframe / jevpandas *shape*) | pandas/Polars, thresholds in code | -| Computer-use speed | one verified act per step | operation + target Choice on numbered controls (**Empirical**: solari-reflex) | Guard check; deny-list absence; no screenshots | +| Computer-use speed | one verified act per step | score / Choice among numbered a11y/DOM controls (**Empirical**: solari-reflex Jev; gliner2-ultrafast GLiNER2; laya-mind2web Laya DOM indices) | Guard check; deny-list absence; no screenshots; `DONE` ≠ success | | Agent turn | skip memory tour on easy intent | intent Choice (**Empirical**: jev-hermes) | Memory still writes; complex still searches | | Document / lab routing | which pages need the expensive observation | Noul on remainder after a text layer / recipe | local extract, merge order (**Empirical** as OCR-router *shape*) | | Shell / tool allowlist | unlisted remainder after a **proof** | five Nouls on unknown verbs | Proven/Refused in code; cannot block (**Empirical**: jevgate) | | SWE | residual AGENTS.md / CLAUDE.md rules | one Score per named instruction-file rule | linter owns hard rules; bands + fail-open (**Empirical**: Abide replay, `notes.md` §47) | | Screenshot candidates → act | lettered elements code already marked | Choice over those letters | Click in code (**Empirical** as blackwood-rlcd *shape*; CC BY-NC) | -| Browser / DOM candidates → act | numbered elements from a **text** snapshot | Choice / Nouls over those ids (**Empirical** as atlas browser-use *shape*: DOM-as-text + fan-out, not vision) | Click in code; no screenshots | +| Browser / DOM candidates → act | numbered elements from a **text** snapshot | score among those ids (**Empirical**: atlas browser-use / jev-ultrafast / gliner2-ultrafast *shape*: DOM-as-text, not vision) | Click in code; no screenshots; hybrid remote TYPE optional | | Knowledge / recall | fact that is not in the document | **Do not ask.** Retrieve the passage first; then a self-contained Choice (**Empirical**: history suite A wrong@0.90 → C right@0.97) | Index, citation, the passage in `state` | | Dual-process cascade | cheap classify / route vs write | S1 typed decision + τ; S2 generates only on low conf (**Empirical as a productized metaphor**; routing accuracy **unmeasured** — dual-process-ai) | Safety still fail-closed in code | | CI merge-gate | ignore infra noise without merging a real bug | cause Choice per cluster (**Empirical**: latch demo PASS vs BLOCK) | Cluster + fingerprint + `--gate` table; reporter never fails the runner | diff --git a/.agents/skills/augustus/references/methods-catalog.md b/.agents/skills/augustus/references/methods-catalog.md index b7367c6..52102c5 100644 --- a/.agents/skills/augustus/references/methods-catalog.md +++ b/.agents/skills/augustus/references/methods-catalog.md @@ -46,6 +46,7 @@ judgment component is new). | Beam search over taxonomies | Which branches deserve expansion | Choice distributions as branch priority; keep K paths where ambiguity is early | Frontier, budget, final selection | **Empirical recipe** (beam K=3 cookbook) | | Structure induction over a bag | Pairwise "does i depend on j?" (or Choice over order) | One judgment per pair; DAG / scheduler in code | Topology, cycles, execution | **Empirical as a shape** (dag-jev experiment; empty README; no metrics, `notes.md` §48) | | Collab-arm product loop | Scripted legal set vs LLM-propose vs unconstrained | Choice over legal actions; stop on low p rather than guess | Legality, Wilson/McNemar, ceiling flags | **Empirical as a harness shape** (jev-testbench; bake into jevals/Harbor, `notes.md` §48) | +| Computer-use observe → score → act | Which observed control matches the current requirement | Score / Choice among a11y/DOM candidates (Jev *or* GLiNER2 *or* Laya); code clicks | Observation, freshness, dates; never generate selectors; independent outcome check (`DONE` ≠ success) | **Empirical recipe** (jev-ultrafast / solari-reflex Jev; gliner2-ultrafast GLiNER2; laya-mind2web DOM-index Laya; contrast blackwood-rlcd screenshot, `notes.md` §48, §52) | | Screening / Wald sequential tests | Pass / fail / keep-looking per candidate | One Noul gate per candidate in one batched request; budget in code | Sequential rule, stop boundaries | **Hypothesis** | | STPA / STAMP control structure | Sensor reading vs enforced constraint | Judgment as sensor; constraints in policy/code/interlock; STPA table if the sensor lies | The constraint, the actuator, the probe | **Contract** as ownership; **Hypothesis** as domain product (`mappings.md` §8) | | PufferLib / Ocean env contracts | Does this episode look like a known trainer-bug mode? | Cluster failing episodes; never "the policy is correct" | Seeded serial env, Ocean sanity, observed rewards | **Hypothesis** as placement; **Contract** that Ocean is not a comparative baseline (`formal-methods.md` DST trio) | @@ -69,7 +70,7 @@ judgment component is new). | Method | Judgment-shaped component | Jev substitution | Stays in code | Status | |---|---|---|---|---| | Claim–evidence entailment (NLI) | supports / contradicts / not-established per claim–source pair | One Choice per pair + review flag; judge against the cited source text only. Stop-hook cousin: claims vs **session** evidence | Quote extraction, citation graph, audit log. When the answer *is* a span you already hold, **point** at line ids and copy verbatim — the model never writes the excerpt (jev-reviewer). Keyword retrieve is not semantic (clear-head) | **Empirical recipe** (citation_check cookbook; jev-reviewer sample study, `notes.md` §48; clear-head Stop, `notes.md` §51). Atlas: paraphrase_support and reversed_meaning_high_overlap correctly judged when both texts are in `state` (`notes.md` §49) | -| Extractive selection + offline re-threshold | Keep/drop over sentences, ids, or character offsets code already holds | Per-item Noul/Choice + one broadcast; join in order; `redecide` on the log with no new calls. Compaction: retention Choice + span locate | Numbering, header skip, thresholds, publish permission; mutation envelope; copy exact bytes | **Empirical recipe** (testimonial-miner 8-request fixture; jev-reviewer; gliner25-compaction char-offset copies, `notes.md` §48, §50). Cousin of applied-mappings §2 | +| Extractive selection + offline re-threshold | Keep/drop over sentences, ids, character offsets, or observed DOM controls code already holds | Per-item Noul/Choice + one broadcast; join in order; `redecide` on the log with no new calls. Compaction: retention Choice + span locate. Computer-use: score among a11y/DOM candidates | Numbering, header skip, thresholds, publish permission; mutation envelope; copy exact bytes; click in code; never generate selectors | **Empirical recipe** (testimonial-miner 8-request fixture; jev-reviewer; gliner25-compaction char-offset copies, `notes.md` §48, §50; gliner2-ultrafast observe→score→act, `notes.md` §52). Cousin of applied-mappings §2 | | Combinatorial grid / program synthesis | Consistent whole-object from many cells | **Rejected as extractive.** Cell-wise Choice does not assemble ARC grids (4/400 Direct Jev) | Search, a program, a simulator | **Empirical as a negative** (`notes.md` §49) | | Spec vs artifact conformance (model checking *mindset*) | Property holds / violated / unverifiable for a named requirement | One Noul/Score per requirement, batched; violated → named rule back into context (pi-warden / Abide shape). This is **not** TLC/Apalache/GNATprove | Requirement enumeration, enforcement, logging; the **linter** if the rule is lintable; the real checker if you have one | **Empirical recipe** (pi-warden: 6→0 rule breaks, 150 paired runs; jev-pref: YOU define the rule; Abide: productized compile/calibrate/tune/replay, `notes.md` §47; if-ai: plain-English PR check, fail-closed on error, `notes.md` §51). Ownership split: `formal-methods.md` | | AST ∩ semantic lint | Semantic remainder after a parser already extracted units | Typed questions on Tree-sitter targets; do not execute scanned code | Parser, selection, fail-on; `tenbin` owns the lint *skill* | **Empirical as a shape** (jevscan 0.2.0rc4; not a calibration claim; `notes.md` §48) | diff --git a/.agents/skills/augustus/references/mixed-architecture.md b/.agents/skills/augustus/references/mixed-architecture.md index 1e9fa6d..aa724c2 100644 --- a/.agents/skills/augustus/references/mixed-architecture.md +++ b/.agents/skills/augustus/references/mixed-architecture.md @@ -89,7 +89,12 @@ vision.** Translate a screenshot task into a structured DOM snapshot as `state`; fan out over candidate elements in one call. Text-only models then sit in their strong zone (extractable from fed state). Same placement as lizard-agent / solari-reflex; not blackwood -screenshot-in. +screenshot-in. **Backend-agnostic this hour:** +[gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) +is that hole with local GLiNER2 instead of Jev — observe → score +among a11y/DOM candidates → code acts (`notes.md` §52). Hybrid: +local decide; remote fill only for TYPE. `DONE` is not verified +success. **Dual-process cascade (Kahneman productized; routing accuracy unmeasured).** @@ -183,6 +188,7 @@ not a global virtue: | Merge a red CI run | **Fail closed** on `--gate` (`latch`); reporter stays fail-open | False PASS merges a real bug. Missing key never fails Playwright; the gate is a separate step. Judge never says ignore alone | | Plain-English PR check | **Fail closed** on error / empty / low confidence (`if-ai`) | A skipped or timed-out check is not a pass. Threshold is policy, not measured correctness | | Route to a tool / start a side effect | **Fail closed** (don't call) | A wrong tool is an action | +| Actuate an observed browser control | **Fail closed** (code validates the node) | Freshness / visibility / disabled / occlusion in code; model never emits selectors (`gliner2-ultrafast`, jev-ultrafast, solari-reflex). `DONE` does not authorize "success" | | Rerank a retrieved list | Fail open: keep retrieval order (`WiktorB2004/llama-index-jev`, **Empirical recipe** on BEIR nfcorpus: MiniLM 0.340 nDCG@5 → MiniLM+Jev 0.396; rerank fails open, *select* fails closed). Listwise/cross-encoder scores belong here, not on the row above. | Ranking errors are quality; selection errors are control-flow | Worked placements (2026-09-18 topic:jev hour + prior archive): @@ -448,7 +454,8 @@ decision-design card. Do not clone APIs from READMEs. | Decision-as-business-tool | Named judgment; gate is part of the result | Registry, arithmetic, hard guards | jev-decision-layer (unofficial) | | NL cases → checked e2e | Jev selects observed controls | Playwright expectations; PASS/FAIL/BLOCKED | jev-e2e (alpha) | | Extractive quotes / pointer evidence | Per-sentence, per-line-id, or char-offset Noul/Choice | Verbatim join; place; `redecide` / CSV; model never writes the excerpt | testimonial-miner; jev-reviewer; gliner25-compaction | -| Structured observe → decide → act | Operation + target Choice on numbered controls | Guard check; deny-list absence; no screenshots; TYPE is the only generation | solari-reflex | +| Structured observe → decide → act | Score / Choice among numbered a11y/DOM controls | Guard check; deny-list absence; no screenshots; no generated selectors; TYPE is the only generation; `DONE` ≠ verified success | solari-reflex (Jev); jev-ultrafast (Jev); gliner2-ultrafast (GLiNER2); laya-mind2web (Laya, DOM indices) | +| Hybrid local decide + remote fill | Local encoder scores observed controls | Code owns actuators; remote OpenAI-compat helper writes field text only | gliner2-ultrafast (GLiNER2 local + Mercury 2.5 default) | | Dataframe semantic columns | Noul / Choice / Score per row; full `p__` | pandas/Polars, indexes, never silent renormalize | jevpandas; jevframe (PyPI + Polars) | | Route ≠ memory | Intent Choice before a turn | Config + flat tools on easy routes; memory stays on for hard ones | jev-hermes | | Advisory sidecar receipts | Typed answers as `no_action` evidence | Host routing / executor / policy unchanged | agent-workflow-typesafe-ai | @@ -506,7 +513,8 @@ not the class monopoly. **Local contract drop-in this hour:** [`Mattepiu/laya-onnx`](https://huggingface.co/Mattepiu/laya-onnx) — do not copy the inherited vs-Jev table. GLiNER (locate) / GLiClass (categorize) / GLiNER2.5 (local multi-head; extractive compaction is a named *job* on -that family, `notes.md` §50), listwise, and vision families: +that family, `notes.md` §50; computer-use selection is a *different* +named job on GLiNER2 `gliner2-multi-v1`, `notes.md` §52), listwise, and vision families: `judgment-class.md`. ## Design-card extras for mixed systems diff --git a/.agents/skills/augustus/references/toolbox-mapping.md b/.agents/skills/augustus/references/toolbox-mapping.md index 09134ea..3cee2b8 100644 --- a/.agents/skills/augustus/references/toolbox-mapping.md +++ b/.agents/skills/augustus/references/toolbox-mapping.md @@ -78,7 +78,7 @@ component; keep the rest of the method in code. | Discrete math: width vs depth | Fan out in width (parallel ≈ free), pay depth linearly; two-stage only when next options depend on an earlier answer | **Empirical recipe** (fan-out: 12.2× cheaper, 10× faster) | | Psychology: Kahneman | System 2 generates/proposes (LLM), System 1 discriminates (Jev); never the reverse. S1 keeps control; optional S2 is one-use advice and does not fly. Productized cascade: `conf ≥ τ` → S1 decides else S2 writes; routing fails open / safety fails closed; **routing accuracy unmeasured**; keyword fallback ≠ S1. S1 specialists + S2 coordinator is the same split (reification-labs/foreman is description-only Phoenix scaffold this pass — do not invent an Elixir API). Indexer: S1 GLiNER extract on the bulk, escalate LLM on the tail (10–50× unfilled) | **Empirical recipe** (mcts-agent role split; 2026-09-18 mixed-architecture discourse; jev-reflex-autonomy-lab as a control-loop shape, `notes.md` §46; [dual-process-ai](https://github.com/taro1985/dual-process-ai) as a business/life cascade, `notes.md` §49; s1-graphify-indexer, `notes.md` §51). **Route ≠ memory:** a cheap intent gate skips memory/tool *tours* on easy routes; memory still writes; complex still searches (jev-hermes, `notes.md` §48) | | IR / cascades | Cheap relevance / irrelevance before an expensive ranker or generator | **Empirical recipe** (RAG cookbook; jevprune; git-jev-stage; LlamaIndex Jev rerank) | -| IR / extractive keep-drop | Number candidates in code; model selects; copy verbatim; `redecide` thresholds on the log with no new calls. Compaction: same pointer job on tool results (Jev Noul/Score *or* GLiNER encoder) | **Empirical recipe** (testimonial-miner; jev-reviewer pointer-not-generator, `notes.md` §48; gliner25-compaction char-offset + fail-closed keep_full, `notes.md` §50) | +| IR / extractive keep-drop | Number candidates in code; model selects; copy verbatim; `redecide` thresholds on the log with no new calls. Compaction: same pointer job on tool results (Jev Noul/Score *or* GLiNER encoder). Computer-use: same pointer job on observed a11y/DOM controls (Jev *or* GLiNER2) | **Empirical recipe** (testimonial-miner; jev-reviewer pointer-not-generator, `notes.md` §48; gliner25-compaction char-offset + fail-closed keep_full, `notes.md` §50; gliner2-ultrafast observe→score→act, `notes.md` §52) | | Spec / lint | Project-defined semantic rules as predicates over a diff; linter owns hard rules. AST remainder: Tree-sitter units, then typed questions; do not execute scanned code. Plain-English PR check: one condition + min-confidence; fail-closed on error | **Empirical recipe** (jev-pref contract; Abide productized path — replay 93 sessions, edit precision ~26% / turn ~73% before tune, `notes.md` §47; JevLint file-level Noul; pi-warden; snifftest unsure-band; jevscan AST∩semantic, `tenbin` owns the lint skill, `notes.md` §48; if-ai, `notes.md` §51). jev-marshal is Watch / empty this pass | | Formal methods / DST / safety | Judgment triages counterexamples, failing seeds, and named-rule conformance; proof/MC/DST stay with their tools. Alloy finder ≠ Apalache BMC ≠ Quint run. DST trio: Antithesis hypervisor / Resonate HQ Lean+oracle+SDK (durable async) / PufferLib env+seed. Noul is a sensor, not a discharged PO. Semi-formal diagrams are vocabularies, not enforcers | **Hypothesis as product**, **Contract** as ownership (matching `mappings.md` §8 and `methods-catalog.md`; worked shape pi-warden — `formal-methods.md`, `formal-semi-formal.md`) | | Decision analysis: VOI | Gather as an enumerated act; pay iff expected decision-loss drop > cost. Wake/resume: skip the LLM turn only if the judge answers and p is low (Horvitz); user-message / skip-limit already answer without a model | **Hypothesis** as calculator (`mappings.md` §6; `mental-models.md`). wakegate 21/21 is smoke (`notes.md` §51) | diff --git a/.agents/skills/augustus/references/validation.md b/.agents/skills/augustus/references/validation.md index fa5d573..dc46664 100644 --- a/.agents/skills/augustus/references/validation.md +++ b/.agents/skills/augustus/references/validation.md @@ -324,7 +324,7 @@ Rules: | LM-program knobs only | DSPy/Ax (narrow) | never primary System One calibration score | | Reward-hack / eval gaming | [rh-guard](https://github.com/24601/rh-guard) | structural deny + System One sidecar | | Project soft-rule lint | [Abide](https://github.com/coldteadotai/abide) | Score per rule on the diff; bands; fail-open; replay + independent review | -| Collab / computer-use product loop | [jev-testbench](https://github.com/ufx7/jev-testbench); [solari-reflex](https://github.com/hitakshiA/solari-reflex) | Wilson/McNemar arms; independently checked task time | +| Collab / computer-use product loop | [jev-testbench](https://github.com/ufx7/jev-testbench); [solari-reflex](https://github.com/hitakshiA/solari-reflex); [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) | Wilson/McNemar arms; independently checked task time; `DONE` ≠ success | | Agent routing on vs off | [jev-gateway-bench](https://github.com/vinilana/jev-gateway-bench) | Hidden perft; cost/quality; one-run signal this pass | rh-guard is a reward-hack hook, a different surface from jevgate and @@ -343,7 +343,12 @@ in the rubric. Text/diff only. the *task* (Stripe API / answer key), not a paragraph judge. Observe → decide → verified act; no screenshots. Author table vs Codex on the same Solari machines: 60.2 s vs 194.9 s; 66 s vs 460 s; 24.2 s -vs 98.4 s (`notes.md` §48). **Collab-arm curriculum:** +vs 98.4 s (`notes.md` §48). Encoder-backend cousin: +[gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) +— same hole, local GLiNER2; their Flights demo (12.20 s visible / +13.785 s loop / ~$0.0001 API) is a **demonstration**, not a bake-off +or a vs-Jev-Ultrafast table; `DONE` is not the Harbor score +(`notes.md` §52). **Collab-arm curriculum:** [jev-testbench](https://github.com/ufx7/jev-testbench) — `llm_autonomous` vs `scripted_plus_jev` vs `llm_plus_jev`; Wilson + McNemar; Jev is not a peer arm. Bake into jevals/Harbor hygiene, do diff --git a/CHANGELOG.md b/CHANGELOG.md index 755ec74..c774197 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -226,6 +226,19 @@ Pegged against [`typesafe-ai/skills` v0.5.7](https://github.com/typesafe-ai/skil jev-marshal Watch/empty; jevons bounded Pi supervisor (shadow recovery). MED: if-ai, omp-auto-mode, downloads-sorter, label-desk, herdr-jev. Archer still Watch. No invented metrics. No wrapper. +- GLiNER2 Ultrafast observe→score→act (`research/notes.md` §52, + [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast), + MIT): architecture notes, not a browser-agent how-to. Same + observe→score-among-candidates→code-acts *job* as jev-ultrafast / + solari-reflex; local GLiNER2 (`fastino/gliner2-multi-v1`) backend, + not GLiNER2.5. No screenshots; no generated selectors; code owns + actuators. Hybrid local decide + remote fill (Mercury 2.5 default + for TYPE). `DONE` ≠ verified success. Contrast blackwood-rlcd + screenshot multimodal; laya-mind2web is DOM-index Laya (same + observed-candidate family). Fastino sibling class with + gliner25-compaction (different hole) and GLiGuard (safety schema). + Demo (theirs, not re-run): Flights 12.20 s / 13.785 s / ~$0.0001 + API — demonstration, not a bake-off. No invented metrics. - Effect-oriented loops (`notes.md` §28, `mappings.md` §19): Ward's ZIO client keeps Jev as the outer Choice and the handler as the effect. Not Effect.ts. GLiNER author: GLiNER2 "like jev" is GLiGuard diff --git a/README.md b/README.md index 44ce414..2301bbe 100644 --- a/README.md +++ b/README.md @@ -36,7 +36,9 @@ never launder a Noul as a proof. contract-compatible local `/v1/systemone` (stub until hf scorer; also kev pointer / von tiny SAN — not replicas), GLiNER/GLiClass species (locate vs categorize vs local multi-head; GLiNER2.5 extractive compaction as a named job, not a new species; - GLiNER code-graph indexer + escalate-S2, 10–50× unfilled), + GLiNER code-graph indexer + escalate-S2, 10–50× unfilled; + GLiNER2 observe→score-among-candidates computer-use as a *different* + named job, not GLiNER2.5), listwise vs decision objectives, vision scoring, when-to-use axes (including decision-model vs constrained LLM), agent-architecture portents @@ -53,12 +55,13 @@ never launder a Noul as a proof. placement: judgment-class model + LLM + code; preference lint; provider (Jev default / other family with self-eval); dual-process S1 decide / S2 generate; component node; DOM-as-text + fan-out; shadow-mode compaction rollout; - fail-open wake vs fail-closed merge-gate; Harbor on/off routing + fail-open wake vs fail-closed merge-gate; Harbor on/off routing; + hybrid local decide + remote fill; `DONE` ≠ verified success - `.agents/skills/augustus/references/applied-mappings.md` — context sieve, - exact-text keep/drop (extractive / pointer-not-generator; char-offset compaction), env triage (OpenSmoke + latch merge-gate), moderation/ranking, skill routing (route ≠ memory) + exact-text keep/drop (extractive / pointer-not-generator; char-offset compaction; observed a11y/DOM controls), env triage (OpenSmoke + latch merge-gate), moderation/ranking, skill routing (route ≠ memory) - `.agents/skills/augustus/references/faq.md` — "just classification", stack replacement, Jev vs open head vs encoder vs LoRA vs constrained AR vs kev vs blackwood, - wait-for-Archer, missing-other confident-wrong, soft project rules vs linter (Abide), extractive/pointer-not-generator, compaction summarize vs pointer, encoder vs Jev compaction, fail-closed keep_full, shadow-mode rollout, fail-open vs fail-closed wake vs CI gate, local drop-in vs stub scorer, route ≠ memory, when-it-holds / extractable-from-state, decision-model vs constrained LLM, dual-process S1/S2, combinatorial grid ≠ extractive, GLiNER vs GLiClass vs CLIP, LLM-as-judge, in-engine vs CLI store, + wait-for-Archer, missing-other confident-wrong, soft project rules vs linter (Abide), extractive/pointer-not-generator, compaction summarize vs pointer, encoder vs Jev compaction, fail-closed keep_full, shadow-mode rollout, fail-open vs fail-closed wake vs CI gate, observe→score→act backend-agnostic, hybrid local decide + remote fill, DONE ≠ verified success, local drop-in vs stub scorer, route ≠ memory, when-it-holds / extractable-from-state, decision-model vs constrained LLM, dual-process S1/S2, combinatorial grid ≠ extractive, GLiNER vs GLiClass vs CLIP, LLM-as-judge, in-engine vs CLI store, hard envelope (bitrate / planner), not-another-how-to - `.agents/skills/augustus/references/mappings.md` — classical-method mappings with boundaries, counterexamples, acceptance tests (including @@ -70,7 +73,8 @@ never launder a Noul as a proof. Harbor-style frozen protocol vs constrained LLMs; jevals-data as CC-BY-4.0 recompute-from-logs feedstock; Abide replay as Harbor-adjacent soft-rule measurement; solari-reflex Harbor-style - computer-use; jev-testbench collab arms; ARC-AGI Direct Jev as + computer-use; gliner2-ultrafast encoder-backend cousin (`DONE` ≠ + success; demo is not a bake-off); jev-testbench collab arms; ARC-AGI Direct Jev as combinatorial-≠-extractive negative; jev-gateway-bench Harbor on/off routing one-run signal) - `.agents/skills/augustus/references/boundary-audit.md` — existing-system diff --git a/docs/ecosystem.md b/docs/ecosystem.md index 9ebe5eb..8b23ee9 100644 --- a/docs/ecosystem.md +++ b/docs/ecosystem.md @@ -50,7 +50,7 @@ Per-keystroke launchers (104ms median, sequence-tagged staleness), firehose mode - **jeiel85/jevscope** — local-first visual debugger + JSONL regression for Choice/Score/Noul; compare two definitions; policy buckets are JevScope-derived. Sits next to jevals. Pointer: `research/notes.md` §25. ### Local / open heads & GLi\* species -- **GLiNER / GLiNER2.5 / GLiClass** — species map: locate spans vs categorize the sequence vs local multi-head (fastino-ai GLiNER2.5 CPU-first). Peer of Jev, not a footnote. `references/judgment-class.md`. Author primary source: GLiNER2 "like jev" is schema-conditioned categorize (GLiGuard), not a Noul (`notes.md` §28). 36× Browser Use claim is a tweet (`notes.md` §25). Named job: [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) — extractive retention Choice + char-offset copies; not a summarizer; not Jev (`notes.md` §50). +- **GLiNER / GLiNER2.5 / GLiClass** — species map: locate spans vs categorize the sequence vs local multi-head (fastino-ai GLiNER2.5 CPU-first). Peer of Jev, not a footnote. `references/judgment-class.md`. Author primary source: GLiNER2 "like jev" is schema-conditioned categorize (GLiGuard), not a Noul (`notes.md` §28). 36× Browser Use claim is a tweet (`notes.md` §25). Named jobs: [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) — extractive retention Choice + char-offset copies; not a summarizer; not Jev (`notes.md` §50). [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) — GLiNER2 `fastino/gliner2-multi-v1` scores observed a11y/DOM controls; not GLiNER2.5; not multimodal (`notes.md` §52). - **GLiGuard** (fastino-ai) — 0.3B GLiNER2 encoder, checkpoint `fastino/gliguard-LLMGuardrails-300M`. One bidirectional pass over a safety schema. Same interface shape as batched questions; different objective. Not a Jev weight clone. `judgment-class.md`; `notes.md` §30. - **DECRUX9812/openjev-lm** — Qwen2.5-0.5B+LoRA distilled from hosted Jev answers; 65/70 = 92.9% on 70 hand-labelled rows (one annotator, one domain, one seed) overnight on 6 vCPU, $0/call. Its 98.1% on fresh rows is teacher *agreement*, not gold. Receipts pattern: `notes.md` §25, §44. - **convaiinnovations/laya** — open Choice/Score/Noul head, text-only, 512 tok. Companion packaging this hour: [`laya-typed-decisions`](https://huggingface.co/convaiinnovations/laya-typed-decisions) (421.3M, acc 0.766 / Brier 0.066 unverified). Shared bake-off: [`pngwn/open-jev-laya-bench`](https://huggingface.co/datasets/pngwn/open-jev-laya-bench) (26+9 tasks, 11959 items; ECE/NLL/Brier; not TypeSafe Jev vs Laya). ONNX replica: [`Mattepiu/laya-onnx`](https://huggingface.co/Mattepiu/laya-onnx) (~15 ms CPU; do not copy vs-Jev table). `notes.md` §18, §42, §46, §48. @@ -127,7 +127,7 @@ the READMEs, not a monopoly. - **AppitStudio/testimonial-miner** — extractive selection + multi-question broadcast + offline `redecide`. Model never writes the quote. - **choxos/jev-reviewer** — pointer-not-generator: line ids; verbatim copy with place; *not found* is an answer. - **us/jev-local** — contract-compatible `POST /v1/systemone`. Default scorer is a **stub** until `JEVLOCAL_SCORER=hf`. -- **hitakshiA/solari-reflex** — observe → decide → verified act; no screenshots. Author table vs Codex on Solari ~3–7× wall. +- **hitakshiA/solari-reflex** — observe → decide → verified act; no screenshots. Author table vs Codex on Solari ~3–7× wall. Encoder-backend cousin: gliner2-ultrafast (`notes.md` §52). - **ktaletsk/jevframe** — pandas/Polars `.jev` accessor; full `p__`; sibling of jevpandas. - **de-niji/jev-hermes** — route ≠ memory: cheap intent gate skips memory tours. - **ngallodev-software/agent-workflow-typesafe-ai** — advisory sidecar receipts; never changes host routing (Apache-2.0). @@ -173,6 +173,12 @@ Architecture notes, not a plugin catalog. `notes.md` §51. TypeSafe Jev is the e - **LilDojd/jevons** — bounded Pi supervisor; shadow recovery; never generates commands. MIT. - MED: if-ai (plain-English PR checks, fail-closed on error); omp-auto-mode (safe/unsafe/ask); jev-downloads-sorter (device-loop Choice); jev-label-desk (description-only); herdr-jev (~260 ms triage + triad; no-key heuristic). +### Hourly ~16:56 Boise (GLiNER2 Ultrafast observe→score→act) + +Architecture notes, not a browser-agent catalog. `notes.md` §52. TypeSafe Jev is the exemplar, not a monopoly. **Not Jev. Not GLiNER2.5. Not multimodal.** Archer still Watch. + +- **sahibzada-allahyar/gliner2-ultrafast** — local GLiNER2 (`fastino/gliner2-multi-v1`) scores observed a11y/DOM controls. Adaptation of jev-ultrafast. No screenshots; no generated selectors; code owns actuators. Hybrid local decide + remote Mercury 2.5 fill. `DONE` ≠ verified success. Same *job* as jev-ultrafast / solari-reflex; encoder backend. Contrast blackwood-rlcd (screenshot + marked letters). Cousin: ShaunSpark/laya-mind2web-browser-agent (Laya over DOM indices). Fastino sibling class with gliner25-compaction (different hole) and GLiGuard (safety schema). Demo (theirs, not re-run): Flights 12.20 s visible / 13.785 s loop / ~$0.0001 API — demonstration, not a bake-off. MIT. + See `references/mixed-architecture.md` in the skill. Class-level family choice: `references/judgment-class.md`. Proof vs judgment (Alloy vs Apalache; DST trio Antithesis / Resonate HQ / PufferLib): diff --git a/research/archive/findings.md b/research/archive/findings.md index 6c5685c..686a4b1 100644 --- a/research/archive/findings.md +++ b/research/archive/findings.md @@ -917,8 +917,6 @@ compaction job is backend-agnostic (Jev Score/Noul vs GLiNER encoder); (bx) shadow-mode default is the rollout for memory mutation. - - ## Batch #35 (2026-09-18 ~16:48 Boise) — CI merge-gate, fail-open wake, S1 indexer, claim-evidence Note: `research/notes.md` §51. Docs-only. Folded into PR #2. Not a @@ -967,6 +965,39 @@ picking fail polarity (wake skip vs merge PASS vs compaction drop); (cc) claim/evidence Stop is anti-hallucinated-done, not a test runner; (cd) description-only greenfield is not a product receipt. +## Batch #36 (2026-09-18 ~16:56 Boise) — GLiNER2 Ultrafast observe→score→act + +Note: `research/notes.md` §52. Docs-only. Folded into PR #2. Not a +second PR. Archer still Watch. **Not Jev. Not GLiNER2.5. Not +multimodal.** No invented metrics. Do not re-fold §48 solari, §50 +compaction, §51 CI merge-gate / wake, GLiGuard species, or jev-ultrafast 7.1 s as this demo. + +- **sahibzada-allahyar/gliner2-ultrafast (Empirical as README / + architecture behavior).** MIT. Created 2026-09-18T19:35:23Z; 12★ + at capture. Adaptation of browser-use/jev-ultrafast. Local + GLiNER2 (`fastino/gliner2-multi-v1`) extracts requirements and + scores observed a11y/DOM controls. No screenshots. No generated + selectors or JS. Code owns order, dates, freshness, clicks. + Hybrid: local decide; Mercury 2.5 via OpenRouter for TYPE. + `DONE` is loop termination; apps must verify outcomes + independently. Inspector scores are not calibrated P(success). + Demo (theirs, not re-run): NYC→SFO Flights 12.20 s visible / + 13.785 s loop / ~$0.0001 API. Demonstration, not a bake-off. +- **Mental models:** (1) observe→score-among-candidates→code-acts + is backend-agnostic (Jev Ultrafast ↔ GLiNER Ultrafast; same + lesson as compaction); (2) observed DOM/a11y candidates vs + screenshot multimodal (solari / laya-mind2web DOM-index Laya vs + blackwood-rlcd); (3) hybrid local decide + remote fill; (4) + composition + independent outcome check. +- **rh-guard:** light note only (do not trust `DONE`). Not + reward-hack detection. + +Cross-repo addition: (ce) computer-use selection is +backend-agnostic on the same observed-candidate hole; (cf) +screenshot multimodal is a different input from DOM-as-text; +(cg) local decide + remote TYPE is mixed-architecture economics, +not dual-process-ai; (ch) `DONE` ≠ Harbor-verified success. + diff --git a/research/notes.md b/research/notes.md index 14dfdfe..89ef549 100644 --- a/research/notes.md +++ b/research/notes.md @@ -3709,3 +3709,168 @@ Not multimodal. Usage: merge-gate / wake / claim-evidence as text-state decisions **before** a generative turn. Archive + landscape pointer. Archer still Watch. +## 52. GLiNER2 Ultrafast — encoder backend of observe→score-among-candidates→code-acts (2026-09-18 ~16:56 Boise) + +America/Boise ~16:56 = 22:56 UTC. Docs-only fold into open PR #2 +(`cursor/augustus-store-envelope-00b4`). Not a competing PR. Archer +27B drop still **WATCH**. Identity lock vs `typesafe-ai` / `tenbin` / +`decision-first` holds. No wrapper, no `uv` / `.env` / Browser Harness +doctor how-to, no copied ports or Hub download scripts. No invented +metrics. **Not Jev. Not GLiNER2.5. Not multimodal.** Do not re-fold +§48 solari-reflex as a new product, §50 compaction as a species +rewrite, §51 CI merge-gate / wake / Harbor on/off, GLiGuard as a +safety clone, or jev-ultrafast's 7.1 s Flights +row as this demo. + +Backend-agnostic: this is an **observe → score among observed +candidates → code acts** placement (pointer/selection among a11y/DOM +controls; hybrid local decide + remote fill; `DONE` is loop +termination, not verified success). TypeSafe Jev is one backend for +that *job* (`browser-use/jev-ultrafast`, `hitakshiA/solari-reflex`); +local Fastino GLiNER2 is another. Same instinct as §50 compaction +(Jev Noul/Score vs GLiNER2.5 encoder): Augustus stays family-first. + +### HIGH + +1. **[`sahibzada-allahyar/gliner2-ultrafast`](https://github.com/sahibzada-allahyar/gliner2-ultrafast)** + (MIT; Python; created 2026-09-18T19:35:23Z; 12★ at capture). Fork / + adaptation of [`browser-use/jev-ultrafast`](https://github.com/browser-use/jev-ultrafast) + that swaps TypeSafe Jev for **local Fastino GLiNER2** + ([`fastino/gliner2-multi-v1`](https://huggingface.co/fastino/gliner2-multi-v1); + Apache-2.0 weights, 307M extractor, GLiNER2 not GLiNER2.5). Real + browser automation. **Work in progress.** **Not a prose planner. + Not a screenshot VLM. Not a hosted decision-model API.** + + Load-bearing loop (README + `docs/architecture.md`): + + ```text + goal → local GLiNER2 → requirements + page → observed controls → local matching + controller → browser action + → text helper (API) if typing needed + ``` + + GLiNER extracts requirement spans and **scores observed controls**. + Candidates come from DOM / accessible labels (`snapshot.js`); the + model chooses among them. It does **not** use screenshots. It does + **not** generate selectors or executable JavaScript. Code owns + requirement order, progress, calendar matching (English month + names, ISO, US-style numeric), form submit, freshness / visibility + / disabled / occlusion checks, and execution. Browser mutations + are not automatically retried after uncertain execution. The + inspector's scores are **not calibrated probabilities of task + success**. + + Default hybrid, not fully offline: local GLiNER2 for decide; + Mercury 2.5 via OpenRouter for typed field text (OpenAI-compatible + endpoint configurable). Websites and the text-model service need + the network. The browser uses an owned tab in an existing Chrome + profile. + + **`DONE` ≠ verified success.** Termination is heuristic. Loop + `DONE` reports that the agent stopped; applications must + independently inspect the actual result. Example verifiers run + *after* the loop and do not choose actions. Same Harbor-style + honesty as solari-reflex (`notes.md` §48). + + **Demo claims (theirs; not re-run; not a bake-off).** Live Google + Flights, one-way NYC→SFO on 9 Oct 2026; no ticket selected or + purchased. README / `docs/demo.md`: **12.20 s** to visible results + (12.201 s frame); complete action loop **13.785 s**; API usage + **~$0.0001** ($0.00010623 across three text-helper calls). Clock + starts after model loading, goal parsing, initial navigation, and + first page observation. Independent outcome verification is + *outside* that clock. Local compute and electricity excluded. + Their measurement doc: a demonstration, not a controlled + performance comparison or general reliability benchmark. Do not + invent a vs-Jev-Ultrafast table; do not merge with the atlas + 7.1 s / 9.5 s Flights figures (`notes.md` §49). + + **Four load-bearing mental models (architecture, not a plugin + catalog):** + + 1. **Same job, encoder backend.** Jev Ultrafast scores among + observed controls with TypeSafe Jev; this repo scores among + observed controls with local GLiNER2. The *hole* is + observe → score-among-candidates → code acts. Compaction + (`notes.md` §50) already taught that keep/drop is + backend-agnostic (Jev Noul/Score vs GLiNER2.5). Computer-use + selection is the same lesson on a different hole. Augustus + does not pick a vendor for the hole. + + 2. **Candidates from observation, not generation; not + screenshot multimodal.** Control set is a11y/DOM-derived. + Pointer/selection species with + [solari-reflex](https://github.com/hitakshiA/solari-reflex) + (structured obs, Jev, no screenshots; `notes.md` §48) and + [laya-mind2web-browser-agent](https://huggingface.co/ShaunSpark/laya-mind2web-browser-agent) + (Laya open head; operation + target index over a list of + interactive DOM elements; author-reported 74.3% on 68 + held-out — small n, not a ranking; Apache-2.0). Contrast + [blackwood-rlcd](https://huggingface.co/BlackwoodAI/blackwood-rlcd) + (screenshot + letters code already marked → Choice; CC BY-NC; + `notes.md` §46). Pixel-free / DOM-as-text is the preferred + computer-use placement when the environment is already + structured (`judgment-class.md` vision pattern 1; + `mental-models.md` §boundary). + + 3. **Hybrid local decide + remote fill.** Economics: tiny API + for TYPE text; local encoder for control scoring. Mixed + architecture, not dual-process-ai (that product is S1/S2 with + a confidence τ; routing accuracy unmeasured — `notes.md` + §49). Here the *split* is where compute lives: judgment on + the laptop, generation only where a string must be typed. + Code still owns actuators. A loop with nothing to type stays + local. + + 4. **Composition + independent outcome check.** Perception + (Browser Harness / DOM snapshot) → decision (GLiNER2) → + verified act (code). The post-run verifier is not the + decision policy. Same Harbor instinct as solari / jev-e2e: + score the *task*, not a self-report. + + **Limits (theirs, README + architecture; experimental).** Common + HTML and ARIA patterns; behavior varies with page structure and + goal wording. Date parsing is English-oriented. Not fully + offline. Optional traces can include page content. No published + Harbor taskset or calibration of control scores as P(success) — + do not invent them. + + **Siblings — complementary, do not merge.** + + - **`browser-use/jev-ultrafast`:** same job, Jev backend. Credit + in their README. Do not treat this demo clock as a head-to-head. + - **`hitakshiA/solari-reflex`:** same observe → decide → verified + act; Jev; Harbor-style table vs Codex on Solari (`notes.md` §48). + - **`m-newhauser/gliner25-compaction`:** Fastino encoder sibling + *class*, GLiNER2.5 (`fastino/gliner2.5-base-v1`), compaction / + keep-drop hole — not this checkpoint and not browser CU + (`notes.md` §50). + - **GLiGuard:** Fastino encoder sibling (safety-schema classify). + Not control scoring. + - **blackwood-rlcd:** screenshot multimodal decide. Different + input. Not this placement. + - **`24601/rh-guard`:** light note only — independent outcome + verification vs trusting `DONE`. Not reward-hack detection. + + **Placement.** Exact-text keep/drop among observed controls + (`applied-mappings.md` §2) + mixed architecture (code acts; LLM + writes TYPE only). Pillar: search/control (one substituted + classifier step) + runtime-assurance sandwich (freshness / + visibility / no generated selectors in code). Hole: perceive / + keep-drop / replace-one-classifier-step. Family: GLi\* encoder + (GLiNER2 local, `gliner2-multi-v1`). Fail-closed on actuation + (code validates the observed node). Eval path: their Flights + demonstration + post-run verifier; no class bake-off this pass. + **Empirical** as README / architecture behavior. **Hypothesis** + that the same envelope transfers to *your* sites. Cards: + `judgment-class.md` (primary); `mixed-architecture.md` (primary); + `applied-mappings.md` §2; `mappings.md` §9 / §12; `faq.md`; + `mental-models.md`; `validation.md`; `methods-catalog.md`; + `toolbox-mapping.md`; `agent-self-assessment.md`. No wrapper. + +### Omni / Jev-omni + +Not multimodal pixels. Strong **composition / open-weights decide** +exemplar for browser computer-use (local encoder + remote fill). +Archive + landscape. Harbor-style independent verify is already in +their framing. diff --git a/research/refresh-log.md b/research/refresh-log.md index ad59f8b..f6913d6 100644 --- a/research/refresh-log.md +++ b/research/refresh-log.md @@ -538,6 +538,30 @@ ecosystem, CHANGELOG, README. - notes.md §51; sources.json; findings.md batch #35. No wrapper. +## 2026-09-18 22:56 UTC — GLiNER2 Ultrafast observe→score→act (~16:56 Boise) + +- Folded into open PR #2 (`cursor/augustus-store-envelope-00b4`). + Docs-only. Not a competing PR. Archer 27B drop still **WATCH**. + **Not Jev. Not GLiNER2.5. Not multimodal.** No invented metrics. +- HIGH: [sahibzada-allahyar/gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) + (MIT, created 2026-09-18). Local GLiNER2 + `fastino/gliner2-multi-v1` scores observed a11y/DOM controls. + No screenshots; no generated selectors; code owns actuators. + Hybrid local decide + remote Mercury 2.5 fill. `DONE` ≠ verified + success. Same *job* as jev-ultrafast / solari-reflex; encoder + backend. Fastino sibling class with gliner25-compaction (different + hole, GLiNER2.5) and GLiGuard (safety schema). Contrast + blackwood-rlcd screenshot→Choice; laya-mind2web is DOM-index Laya + (same observed-candidate family). +- Demo (theirs, not re-run): Flights 12.20 s visible / 13.785 s loop / + ~$0.0001 API. Not a bake-off. Do not merge with atlas 7.1 s. +- Cards: SKILL.md, judgment-class (primary), mixed-architecture + (primary), applied-mappings §2, mappings §9/§12, faq, + mental-models, validation, methods-catalog, toolbox, + agent-self-assessment, composition-algebra, ecosystem, CHANGELOG, + README. +- notes.md §52; sources.json; findings.md batch #36. No wrapper. + diff --git a/research/sources.json b/research/sources.json index 7087f88..0e5a599 100644 --- a/research/sources.json +++ b/research/sources.json @@ -1,6 +1,6 @@ { "refresh_cadence": "hourly", - "retrieved": "2026-09-18T22:48Z", + "retrieved": "2026-09-18T22:56Z", "sources": [ { "kind": "docs", @@ -148,7 +148,7 @@ }, { "kind": "repo", - "note": "Operation+target heads, 7.1s Flights; ~2100 stars", + "note": "Operation+target heads, 7.1s Flights; ~2100 stars. Same observe\u2192score\u2192act *job* as sahibzada-allahyar/gliner2-ultrafast (GLiNER2 backend, notes.md \u00a752); do not merge demo clocks.", "title": "browser-use/jev-ultrafast", "url": "https://github.com/browser-use/jev-ultrafast" }, @@ -1485,6 +1485,7 @@ "note": "Apache-2.0. Local GLiNER2.5 (fastino/gliner2.5-base-v1) extractive compaction for Claude Code. Retention Choice keep_full|keep_evidence|keep_call_only|drop; exact char-offset copies; mutating tools keep_full; fail-closed to keep_full; shadowMode default true. Not Jev, not a summarizer, not multimodal. notes.md \u00a750." }, { + "kind": "github", "kind": "github", "title": "CaseReed/latch", "url": "https://github.com/CaseReed/latch", @@ -1573,6 +1574,24 @@ "title": "flaviomartil/herdr-jev", "url": "https://github.com/flaviomartil/herdr-jev", "note": "Jev triage ~260ms + triad orchestration. No-key local heuristic. Do not copy install.sh. License not stated this pass. notes.md \u00a751." + }, + { + "kind": "github", + "title": "sahibzada-allahyar/gliner2-ultrafast", + "url": "https://github.com/sahibzada-allahyar/gliner2-ultrafast", + "note": "MIT. Created 2026-09-18T19:35:23Z; 12 stars at capture. Adaptation of browser-use/jev-ultrafast: local Fastino GLiNER2 (fastino/gliner2-multi-v1) scores observed a11y/DOM controls; no screenshots; no generated selectors; code owns actuators. Hybrid local decide + remote Mercury 2.5 fill. DONE is loop termination, not verified success. Demo (theirs, not re-run): NYC\u2192SFO Flights 12.20s visible / 13.785s loop / ~$0.0001 API. Not GLiNER2.5, not multimodal. notes.md \u00a752." + }, + { + "kind": "huggingface", + "title": "fastino/gliner2-multi-v1", + "url": "https://huggingface.co/fastino/gliner2-multi-v1", + "note": "Apache-2.0 GLiNER2 multilingual extractor (~307M). Default local checkpoint for sahibzada-allahyar/gliner2-ultrafast. Distinct from fastino/gliner2.5-base-v1 (compaction, notes.md \u00a750). Do not copy Hub download into Augustus. notes.md \u00a752." + }, + { + "kind": "huggingface", + "title": "ShaunSpark/laya-mind2web-browser-agent", + "url": "https://huggingface.co/ShaunSpark/laya-mind2web-browser-agent", + "note": "Apache-2.0. Laya (convaiinnovations/laya) fine-tune: operation + target index over interactive DOM elements (Mind2Web). Observed-candidate family with gliner2-ultrafast / solari-reflex / jev-ultrafast; not screenshot multimodal (contrast blackwood-rlcd). Author-reported 74.3% on 68 held-out \u2014 small n, not a ranking. Cousin pointer only. notes.md \u00a752." } ] } From d1c30068c4cf21d0938d8ec0626e75ae2d767082 Mon Sep 17 00:00:00 2001 From: Cursor Agent Date: Fri, 18 Sep 2026 23:20:34 +0000 Subject: [PATCH 11/43] Fold jev-pruner as evidence-preserving Bash stdout prune MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Docs-only: tamaratran/jev-pruner Noul-prunes command output after a hard ≤10k/JSON-diff envelope, before the main LLM sees it. Same extractive family as fast-jev-compaction / gliner25-compaction; different job. Fail-safe keep original; archive for recovery. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com> --- .agents/skills/augustus/SKILL.md | 10 +- .../references/agent-self-assessment.md | 9 +- .../augustus/references/applied-mappings.md | 16 +- .../references/composition-algebra.md | 3 + .agents/skills/augustus/references/faq.md | 35 ++++- .../augustus/references/judgment-class.md | 10 ++ .../skills/augustus/references/mappings.md | 16 ++ .../augustus/references/mental-models.md | 5 +- .../augustus/references/methods-catalog.md | 4 +- .../augustus/references/mixed-architecture.md | 8 +- .../augustus/references/toolbox-mapping.md | 4 +- .../skills/augustus/references/validation.md | 15 ++ CHANGELOG.md | 13 ++ README.md | 10 +- docs/ecosystem.md | 8 +- research/archive/findings.md | 33 ++++ research/notes.md | 141 ++++++++++++++++++ research/refresh-log.md | 21 +++ research/sources.json | 8 +- 19 files changed, 345 insertions(+), 24 deletions(-) diff --git a/.agents/skills/augustus/SKILL.md b/.agents/skills/augustus/SKILL.md index 145e213..987c985 100644 --- a/.agents/skills/augustus/SKILL.md +++ b/.agents/skills/augustus/SKILL.md @@ -1,6 +1,6 @@ --- name: augustus -description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill, open multimodal RLCD / blackwood), announced open decision-model (Watch — still not landed), constrained-AR (TypeAR, pcdServer, decision-token LoRA), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"wait for Archer vs open multimodal RLCD\", \"screenshot/DOM candidates → typed Choice\", \"eval path\", \"jevals\", \"Harbor taskset\", \"shared bake-off ECE/NLL/Brier\", \"LLM-as-judge is not the System One score\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge (allowlist *proves*; fail-open cannot block), \"missing other → confident wrong Choice\", \"lint the Jev request\", \"training confronts Choice other / none-of-the-above\", \"S1 reflex keeps control / optional S2 one-use advice\", \"soft AGENTS.md rules vs linter\" (Abide / jev-pref), \"edit-phase vs turn-phase observation window\", banded confidence fail-open preference lint, \"extractive selection / pointer-not-generator\", \"encoder GLiNER compaction vs Jev Score compaction (same job; pointer not summarizer)\", \"fail-closed keep_full under mutation envelope\", \"CI flaky-vs-real merge gate\", \"fail-open VOI wake/resume (Horvitz)\", \"claim/evidence Stop integrity\", \"S1 extract + escalate-S2 indexer\", \"Harbor on/off routing\", \"policy-as-judgment PR marshal\", \"shadow-mode compaction rollout\", \"Jev Ultrafast vs GLiNER Ultrafast (observe-score-act backend-agnostic)\", \"hybrid local decide + remote fill\", \"DONE ≠ verified success\", \"observed a11y/DOM candidates vs screenshot multimodal\", \"local /v1/systemone drop-in (stub until hf scorer)\", \"dataframe-native semantic columns\", \"route≠memory\", \"advisory sidecar receipts\", \"structure induction over bags\", \"AST ∩ semantic lint\", \"extractable-from-state / retrieve first\", \"decision-model vs constrained-LLM bake-off\", \"dual-process S1 decide / S2 generate\", \"combinatorial grid ≠ extractive\", TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, or \"formally verify with Jev\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." +description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill, open multimodal RLCD / blackwood), announced open decision-model (Watch — still not landed), constrained-AR (TypeAR, pcdServer, decision-token LoRA), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"wait for Archer vs open multimodal RLCD\", \"screenshot/DOM candidates → typed Choice\", \"eval path\", \"jevals\", \"Harbor taskset\", \"shared bake-off ECE/NLL/Brier\", \"LLM-as-judge is not the System One score\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge (allowlist *proves*; fail-open cannot block), \"missing other → confident wrong Choice\", \"lint the Jev request\", \"training confronts Choice other / none-of-the-above\", \"S1 reflex keeps control / optional S2 one-use advice\", \"soft AGENTS.md rules vs linter\" (Abide / jev-pref), \"edit-phase vs turn-phase observation window\", banded confidence fail-open preference lint, \"extractive selection / pointer-not-generator\", \"encoder GLiNER compaction vs Jev Score compaction (same job; pointer not summarizer)\", \"fail-closed keep_full under mutation envelope\", \"CI flaky-vs-real merge gate\", \"fail-open VOI wake/resume (Horvitz)\", \"claim/evidence Stop integrity\", \"S1 extract + escalate-S2 indexer\", \"Harbor on/off routing\", \"policy-as-judgment PR marshal\", \"shadow-mode compaction rollout\", \"Jev Ultrafast vs GLiNER Ultrafast (observe-score-act backend-agnostic)\", \"hybrid local decide + remote fill\", \"DONE ≠ verified success\", \"observed a11y/DOM candidates vs screenshot multimodal\", \"evidence-preserving stdout prune (not summarize)\", \"hard token/format envelope then soft Noul\", \"fail-safe keep original on prune failure\", \"stdout prune vs session compaction\", \"local /v1/systemone drop-in (stub until hf scorer)\", \"dataframe-native semantic columns\", \"route≠memory\", \"advisory sidecar receipts\", \"structure induction over bags\", \"AST ∩ semantic lint\", \"extractable-from-state / retrieve first\", \"decision-model vs constrained-LLM bake-off\", \"dual-process S1 decide / S2 generate\", \"combinatorial grid ≠ extractive\", TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, or \"formally verify with Jev\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." license: MIT metadata: version: 0.3.0 @@ -54,7 +54,7 @@ classical method you already trust, substitute it, classify the win "paraphrase brittleness", "allowlist then judge", "TOCTOU-of-Noul", "Jev inside the database / sqlite-jev", "Jev picks bitrate / join order / the model", "wait for Archer", "lint the request / missing - other", "training confronts Choice other / none-of-the-above", "soft AGENTS.md rules vs the linter", "screenshot Choice / omni System One", "extractive quotes / pointer not generator", "compaction summarize vs pointer", "encoder vs Jev compaction backend", "shadow-mode compaction rollout", "CI flaky-vs-real merge gate", "fail-open VOI wake/resume", "claim vs session evidence", "S1 indexer escalate-S2", "Harbor on/off routing", "fail-open vs fail-closed wake vs CI gate", "encoder vs Jev computer-use backend", "hybrid local decide + remote fill", "DONE vs verified success", "local drop-in vs stub scorer", "route vs memory", "when does it hold / extractable from state", "decision model vs constrained LLM", "dual-process S1/S2", or "combinatorial grid vs extractive": + other", "training confronts Choice other / none-of-the-above", "soft AGENTS.md rules vs the linter", "screenshot Choice / omni System One", "extractive quotes / pointer not generator", "compaction summarize vs pointer", "encoder vs Jev compaction backend", "shadow-mode compaction rollout", "CI flaky-vs-real merge gate", "fail-open VOI wake/resume", "claim vs session evidence", "S1 indexer escalate-S2", "Harbor on/off routing", "fail-open vs fail-closed wake vs CI gate", "encoder vs Jev computer-use backend", "hybrid local decide + remote fill", "DONE vs verified success", "stdout prune vs session compaction", "hard envelope then Noul prune", "local drop-in vs stub scorer", "route vs memory", "when does it hold / extractable from state", "decision model vs constrained LLM", "dual-process S1/S2", or "combinatorial grid vs extractive": read `references/faq.md`, then `references/mental-models.md`, then `references/mixed-architecture.md`, then @@ -119,7 +119,7 @@ classical method you already trust, substitute it, classify the win | Entropy as allocator (low / medium / high) | Typed low+medium decisions → System One marginals; high-entropy synthesis → frontier decoder. Product rhetoric, not a meter. **Hypothesis** | `references/judgment-class.md` | | Formal / semi-formal (proof vs judgment) | Sensor vs constraint vs searchlight; Alloy vs Apalache; DST trio; TOCTOU-of-Noul, AI×FM | `references/formal-methods.md` (one-screen: `references/formal-semi-formal.md`) | | Mixed architecture (judgment model + LLM) | Provider judges, LLM writes, code owns control; not a stack replacement. Advisory sidecar never changes host routing. Dual-process: S1 decides, S2 generates (routing accuracy unmeasured); S1 specialists + S2 coordinator is the same split (description-only greenfield this hour). Internals ≠ FSM; placement is a **component node**. Browser-use strength = DOM-as-text + speculative fan-out, not vision. Fail polarity is per act: skip-wake fail-open vs merge-gate BLOCK fail-closed. Hybrid local decide + remote fill; `DONE` ≠ verified success | `references/mixed-architecture.md` | -| Context sieve | Relevance Noul per block; always-keep set in code; stub + recall key. Encoder cousin: GLiNER2.5 retention Choice + char-offset spans (gliner25-compaction); fail-closed keep_full; shadowMode default | `references/applied-mappings.md#1-context-sieve` | +| Context sieve | Relevance Noul per block; always-keep set in code; stub + recall key. Encoder cousin: GLiNER2.5 retention Choice + char-offset spans (gliner25-compaction); fail-closed keep_full; shadowMode default. Stdout cousin: jev-pruner (Jev Noul after hard ≤10k/JSON-diff envelope; fail-safe original; archive) | `references/applied-mappings.md#1-context-sieve` | | Exact-text keep / drop | Choice include/exclude/mixed over candidates code already holds. Extractive quotes / pointer-not-generator (model never writes the excerpt; char-offset compaction same species). Observed a11y/DOM controls: score among them; code clicks (Jev or GLiNER2) | `references/applied-mappings.md#2-exact-text-keep--drop` | | Environment / harness triage | Scan every step for env failure; LLM autopsy only on flags. Merge-gate cousin: cluster in code, judge labels cause, policy owns PASS/BLOCK (latch; judge never says ignore alone) | `references/applied-mappings.md#3-environment--harness-triage` | | Moderation and ranking | Hold-before-publish vs graded rerank; fail policy per action | `references/applied-mappings.md#4-moderation-and-ranking` | @@ -133,7 +133,7 @@ classical method you already trust, substitute it, classify the win | Decision tables / circuits / state machines | Judgment predicates, code owns transitions | `references/mappings.md#3-semantic-predicates--decision-circuits` | | Retrieve + expensive relevance fn | Bounded rerank of a retrieved shortlist | `references/mappings.md#4-retrieval--bounded-semantic-reranking` (independent TREC DL2019 benchmark: Jev zero-shot best MAP 0.4748, nDCG@10 0.683 vs tuned monoBERT 0.718 — competitive, not dominant) | | Store as semantic index (SQL / SQLite / zoxide / dataframe) | Cheap exact predicates first; typed questions on the remainder. In-engine extension (sqlite-jev) vs CLI rewrite (jevql) vs path index (joxide) vs dataframe columns (jevpandas / jevframe) | `references/mappings.md#4-retrieval--bounded-semantic-reranking` | -| Soft judgment inside a hard envelope | Model may only match the deterministic policy or be more conservative (bitrate ABR; query-planner override-when-confident; compaction mutations/shell operators → keep_full) | `references/mappings.md#12-runtime-assurance-sandwich-hypothesis`; `references/mappings.md#18-structural-prove--soft-remainder-hypothesis-as-domain-general-empirical-as-named-shapes` | +| Soft judgment inside a hard envelope | Model may only match the deterministic policy or be more conservative (bitrate ABR; query-planner override-when-confident; compaction mutations/shell operators → keep_full; stdout prune: ≤10k/JSON-diff-whole-doc untouched, then Noul) | `references/mappings.md#12-runtime-assurance-sandwich-hypothesis`; `references/mappings.md#18-structural-prove--soft-remainder-hypothesis-as-domain-general-empirical-as-named-shapes` | | Value of information / gather as an act | Pay for another observation only if EV(decision) improves more than cost; abstain from calling *any* model when a regex already answers (meta-VOI). Fail-open wake/resume: skip the LLM turn only if the judge answers and p(wake) is low (Horvitz) | `references/mappings.md#6-value-of-information--gather-as-an-enumerated-act` (**Hypothesis** until a labeled act/outcome log; 149-row receipt is Empirical as a shape; wakegate 21/21 is smoke) | | Signal detection / ROC | Criterion and operating point from costs and base rate, not accuracy | `references/mappings.md#7-signal-detection--criterion-not-accuracy` (**Hypothesis** for non-SWE plots) | | Org / safety control structure | Sensor ≠ constraint (Leveson); STPA if the sensor lies | `references/mappings.md#8-control-structure--sensor--constraint-leveson` | @@ -151,7 +151,7 @@ classical method you already trust, substitute it, classify the win | Agent self-supervision / on-track detection | Pre-gate → output judge → done-check → supervisor nouls. S1 reflex keeps control; optional S2 is one-use advice. Claim/evidence Stop (anti-hallucinated-done); bounded Pi supervisor (shadow recovery, never generates commands) | `references/agent-self-assessment.md` | | Optimizer/program frameworks (Ax, DSPy) | Typed fields → one provider request; judge metrics; threshold discipline. Ax and DSPy climb LM-program knobs only | `references/optimizer-integration.md` | | Perception → decision pipeline / measure / hill-climb | Stages with a versioned state contract; frozen taskset; DSPy/Ax only on the LM-program slice. Specialist composition stays **Hypothesis**; open multimodal decide (blackwood-rlcd) is a named receipt. Structured observe→decide→verified-act (no screenshots) is a computer-use speed-layer receipt (Jev or GLiNER2 backend). Same section as the row below | `references/validation.md#eval--hill-climb` | -| Eval & hill-climb | Decision-stage jevals hygiene; Harbor taskset × harness × runtime; one score-composition table. Shared bake-off exemplar: open-jev-laya-bench (ECE/NLL/Brier; LLM-as-judge is not the score). Harbor-style frozen protocol vs constrained LLMs: DMB (accuracy/calibration/latency/cost; raw logs). Feedstock: jevals-data CC-BY-4.0 boards + JSONL (recompute-from-logs). Collab-arm curriculum: llm_autonomous vs scripted_plus_jev vs llm_plus_jev (Wilson / McNemar). Negative: combinatorial grid assembly ≠ extractive (ARC-AGI Direct Jev 4/400). Harbor on/off routing: chess-engine tasks, hidden perft verifier, one-run preliminary (jev-gateway-bench). Pair CI merge-gate with Harbor + rh-guard | `references/validation.md#eval--hill-climb` | +| Eval & hill-climb | Decision-stage jevals hygiene; Harbor taskset × harness × runtime; one score-composition table. Shared bake-off exemplar: open-jev-laya-bench (ECE/NLL/Brier; LLM-as-judge is not the score). Harbor-style frozen protocol vs constrained LLMs: DMB (accuracy/calibration/latency/cost; raw logs). Feedstock: jevals-data CC-BY-4.0 boards + JSONL (recompute-from-logs). Collab-arm curriculum: llm_autonomous vs scripted_plus_jev vs llm_plus_jev (Wilson / McNemar). Negative: combinatorial grid assembly ≠ extractive (ARC-AGI Direct Jev 4/400). Harbor on/off routing: chess-engine tasks, hidden perft verifier, one-run preliminary (jev-gateway-bench). Pair CI merge-gate with Harbor + rh-guard. Harbor needle/noise stdout prune: jev-pruner (manual sweep theirs; plugin eval cannot reach Jev → fail-safe original; Terminal-Bench pilot is integration not a full bench) | `references/validation.md#eval--hill-climb` | | (meta) Finding new mappings & applications | Toolbox sweep: judgment-shaped component of a known method, substituted + falsified | `references/toolbox-mapping.md` | | Named methods / operators / theorems | Substitution tiers: operand-judgments, preconditioned theorems, non-substitutable | `references/methods-catalog.md` | | (meta) Where a judgment model sits relative to any construct | 11 positions + logical-operator rules + position×construct traversal as the application generator | `references/composition-algebra.md` | diff --git a/.agents/skills/augustus/references/agent-self-assessment.md b/.agents/skills/augustus/references/agent-self-assessment.md index d204bc9..8655a7b 100644 --- a/.agents/skills/augustus/references/agent-self-assessment.md +++ b/.agents/skills/augustus/references/agent-self-assessment.md @@ -83,6 +83,10 @@ relying: foreman, pi-jev, pi-warden, winnow, fast-jev-compaction, jev-judgment. — GLiNER2.5 retention Choice + exact character-offset copies; mutating tools stay `keep_full`; low-confidence fails closed to `keep_full`; `shadowMode` default true. Not a summarizer. Not Jev (`notes.md` §50). + Stdout-prune cousin, same family, different job: + [jev-pruner](https://github.com/tamaratran/jev-pruner) — Jev Noul on + Bash chunks after a hard envelope; fail-safe original; archive + (`notes.md` §53). Marketplace id still `fast-jev-output`. ## Non-negotiable boundaries @@ -97,7 +101,10 @@ relying: foreman, pi-jev, pi-warden, winnow, fast-jev-compaction, jev-judgment. (`mixed-architecture.md` prefilter table; `mappings.md` §18). Compaction *drop* is that second kind: [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) - fails closed to `keep_full` (`notes.md` §50). + fails closed to `keep_full` (`notes.md` §50). Stdout prune is the + same polarity: + [jev-pruner](https://github.com/tamaratran/jev-pruner) fails closed + to original output (`notes.md` §53). - Cache identical judgments (~120s) and deduplicate sibling calls into one in-flight request. - pi-warden measured cost makes continuous guarding viable: ~$0.00004 and diff --git a/.agents/skills/augustus/references/applied-mappings.md b/.agents/skills/augustus/references/applied-mappings.md index d570b56..66bfab2 100644 --- a/.agents/skills/augustus/references/applied-mappings.md +++ b/.agents/skills/augustus/references/applied-mappings.md @@ -55,6 +55,17 @@ irreversible act; from the evidence side this *looks* like keep-on-error). `shadowMode` defaults true. Characters, not tokens; no published retention-quality rates (`notes.md` §50). Family: `judgment-class.md`. Do not copy the plugin. +**Stdout prune, same family, different job (Empirical as README / +evals README, 2026-09-18 ~17:15):** +[jev-pruner](https://github.com/tamaratran/jev-pruner) — after Bash +runs, Jev Noul-prunes stdout chunks **before** the main LLM sees +them; no summary. Hard envelope first (≤10k estimated tokens; +JSON/XML/YAML/diff/binary; whole-document commands untouched), then +soft Noul. Fail-safe keep original on any failure; full archive for +recovery. Marketplace id still `fast-jev-output`. Codex is opt-in +wrapper, not automatic interception. Same author as +fast-jev-compaction; complementary, not a duplicate. Do not copy +the plugin (`notes.md` §53). Local teacher-copy for the same hole: [`SargeDev/jev-gate-student-b`](https://huggingface.co/SargeDev/jev-gate-student-b) (Qwen2.5-0.5B LoRA; P(relevant) from yes/no logits; 148,160-row @@ -112,7 +123,10 @@ not against another model's prose; `JEV_FIRM` below 0.6 never blocks ([gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction), ~16:22): the model points at **character offsets** in a tool result; code copies those bytes; a generator summary is the rejected species -(`notes.md` §50). Computer-use cousin: +(`notes.md` §50). Stdout-prune cousin: +[jev-pruner](https://github.com/tamaratran/jev-pruner) — the model +scores chunks of observed Bash stdout; code keeps verbatim lines and +archives the rest (`notes.md` §53). Computer-use cousin: [solari-reflex](https://github.com/hitakshiA/solari-reflex) — structured observation → typed decision → verified act; **no screenshots**; model output never becomes a selector (`notes.md` §48). Encoder-backend diff --git a/.agents/skills/augustus/references/composition-algebra.md b/.agents/skills/augustus/references/composition-algebra.md index bfc4f0f..ca76fa9 100644 --- a/.agents/skills/augustus/references/composition-algebra.md +++ b/.agents/skills/augustus/references/composition-algebra.md @@ -151,6 +151,9 @@ Reusable shapes when generating applications: (Claude CLI) vs 1.3 s Jev at concurrency 8. Compaction rollout: gliner25-compaction public default `shadowMode: true` (log proposed reduction; do not replace history) (`notes.md` §50). + Stdout-prune cousin: [jev-pruner](https://github.com/tamaratran/jev-pruner) + archives full stdout before scoring; fail-safe keep original + (`notes.md` §53). Marketplace id still `fast-jev-output`. Recovery cousin: [jevons](https://github.com/LilDojd/jevons) default recovery **shadow** (record, do not interrupt); steering never generates commands (`notes.md` §51). diff --git a/.agents/skills/augustus/references/faq.md b/.agents/skills/augustus/references/faq.md index b10c47d..fd54e74 100644 --- a/.agents/skills/augustus/references/faq.md +++ b/.agents/skills/augustus/references/faq.md @@ -470,7 +470,10 @@ ids; *not found* is an answer. [solari-reflex](https://github.com/hitakshiA/sola never lets model output become a selector. Compaction is the same species: [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) copies exact source spans; a prose summary of the tool result is -generation, not keep/drop (`notes.md` §50). Claim/evidence Stop: +generation, not keep/drop (`notes.md` §50). Command-output cousin: +[jev-pruner](https://github.com/tamaratran/jev-pruner) keeps verbatim +chunks of Bash stdout; dropped spans live in an archive, not a +summary (`notes.md` §53). Claim/evidence Stop: [clear-head](https://github.com/VladyslavHontar/clear-head) judges against retrieved session lines, not generated prose (`notes.md` §51). Computer-use encoder @@ -478,7 +481,7 @@ cousin: [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultraf scores observed controls; code clicks; no generated selectors (`notes.md` §52). Generation is only for TYPE/prose when something must be written. `applied-mappings.md` §2; -`notes.md` §48, §50, §52. +`notes.md` §48, §50, §52, §53. ## Should compaction summarize? @@ -491,12 +494,34 @@ stay `keep_full` in **code**. Low-confidence / invalid evidence fail **closed to `keep_full`** — the *reduction* is the irreversible act, unlike Abide / jevgate fail-open. Ship `shadowMode` first (default true: log, do not replace history). Not Jev. Not multimodal. -`judgment-class.md`; `notes.md` §50. +Stdout prune is the same *family* (evidence-preserving reduce) on a +**different job**: [jev-pruner](https://github.com/tamaratran/jev-pruner) +Noul-prunes a just-run Bash result before the main LLM sees it; +fast-jev-compaction / gliner25-compaction compact completed tool +pairs already in history. Hard ≤10k / JSON-diff-whole-doc envelope +in code; fail-safe keep original; archive for recovery. Marketplace +id still `fast-jev-output`. `judgment-class.md`; `notes.md` §50, §53. + +## Is pruning Bash stdout the same as compacting session memory? + +No. Same family (pointer, not summarizer; dropped bytes recoverable). +Different job. [jev-pruner](https://github.com/tamaratran/jev-pruner) +scores chunks of a command that just ran, *before* they enter the +generative turn. [fast-jev-compaction](https://github.com/tamaratran/fast-jev-compaction) +and [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) +reduce completed tool pairs already in session history (Jev Noul vs +GLiNER2.5 encoder). Host capability shapes the product: Claude wraps +Bash automatically; Codex is an opt-in wrapper because it cannot +replace native shell output from `PostToolUse`. `applied-mappings.md` +§1; `notes.md` §50, §53. ## Fail-open or fail-closed — which? Name the **irreversible act**, then pick polarity. Compaction *drop* -fails closed to `keep_full`. Wake *skip* fails open (wake on error / +fails closed to `keep_full`. Stdout *prune* fails closed to original +output ([jev-pruner](https://github.com/tamaratran/jev-pruner): +archive/Jev/incomplete-score failure keeps the log; Harbor plugin-eval +cannot reach Jev and therefore cannot prune). Wake *skip* fails open (wake on error / unsure): [wakegate](https://github.com/shitianfang/wakegate) skips only if Jev answers and p(wake) < 0.2. Merge *PASS* on a red run fails closed at the gate: [latch](https://github.com/CaseReed/latch) @@ -504,7 +529,7 @@ fails closed at the gate: [latch](https://github.com/CaseReed/latch) stays fail-open. [if-ai](https://github.com/Victor-Casado/if-ai) fails the Action on error / empty / low confidence. jevgate cannot block; pi-jev-approver fails closed without a key; Abide is fail-open -on diffs. Same sandwich, opposite authorized act. `notes.md` §50, §51. +on diffs. Same sandwich, opposite authorized act. `notes.md` §50, §51, §53. ## Is observe→score→act Jev-only? diff --git a/.agents/skills/augustus/references/judgment-class.md b/.agents/skills/augustus/references/judgment-class.md index ecd4f23..7e9ea7c 100644 --- a/.agents/skills/augustus/references/judgment-class.md +++ b/.agents/skills/augustus/references/judgment-class.md @@ -99,6 +99,16 @@ below, next to the when-to-use table. tokens; no published retention-quality rates (`notes.md` §50). Fastino/GLiGuard sibling *class*, not a GLiGuard safety-schema clone. Do not copy the plugin install. + **Stdout prune (Empirical as README / evals README, 2026-09-18 + ~17:15):** + [jev-pruner](https://github.com/tamaratran/jev-pruner) (MIT) is the + same evidence-preserving *family* on a **different job** and the + **Jev** backend: Noul-prune Bash stdout after a hard ≤10k / + JSON-diff-whole-doc envelope, before the main LLM sees it. Not a + summarizer. Fail-safe keep original. Archive for recovery. + Marketplace id still `fast-jev-output`. Codex is opt-in wrapper. + Same author as fast-jev-compaction; complementary (`notes.md` + §53). Do not copy the plugin. **Code-graph indexer (Empirical as README behavior; 10–50× is a target, 2026-09-18 ~16:48):** [s1-graphify-indexer](https://github.com/GreyssonEnterprises/s1-graphify-indexer) diff --git a/.agents/skills/augustus/references/mappings.md b/.agents/skills/augustus/references/mappings.md index 958c094..f8eead2 100644 --- a/.agents/skills/augustus/references/mappings.md +++ b/.agents/skills/augustus/references/mappings.md @@ -549,6 +549,16 @@ to `keep_full`. Soft judgment inside a hard envelope, encoder backend — not a Jev Score and not a summarizer (`notes.md` §50). Same sandwich shape as bitrate-advisor; different family. +**Named stdout-prune envelope (Empirical as README / evals README, +2026-09-18 ~17:15):** +[jev-pruner](https://github.com/tamaratran/jev-pruner) +— the monitor is **code** (≤10k estimated tokens; errors; +JSON/XML/YAML/diff/binary; whole-document commands). Jev Noul may +only score residual noisy chunks. Archive/Jev/incomplete-score +failure keeps the original. Soft judgment inside a hard envelope, +Jev backend — same family as gliner25-compaction, different *job* +(command output vs session memory) (`notes.md` §53). + **Named computer-use envelope (Empirical as README / architecture, 2026-09-18 ~16:56):** [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) @@ -725,6 +735,12 @@ judges only the remainder; uncertain **fails closed to `keep_full`** (`notes.md` §50). Same sandwich, opposite fail policy from jevgate (cannot block) and Abide (fail-open on diffs): the authorized act is a destructive reduction of memory. +**Stdout prune is the same polarity, different job (2026-09-18 +~17:15).** [jev-pruner](https://github.com/tamaratran/jev-pruner) — +code proves ≤10k / JSON-diff-whole-doc pass-through; Jev scores the +remainder; uncertain **fails closed to original stdout** plus an +archive (`notes.md` §53). Harbor plugin-eval cannot reach Jev and +therefore cannot prune — fail-safe, not a missing score. **Name the irreversible act (2026-09-18 ~16:48).** Wake *skip* is irreversible (the agent stays asleep) → [wakegate](https://github.com/shitianfang/wakegate) authorizes skip diff --git a/.agents/skills/augustus/references/mental-models.md b/.agents/skills/augustus/references/mental-models.md index 6e326c5..469283d 100644 --- a/.agents/skills/augustus/references/mental-models.md +++ b/.agents/skills/augustus/references/mental-models.md @@ -324,7 +324,9 @@ inside a hard envelope: bitrate-advisor (ABR) and mmalisper's JOB hybrid (Postgres plans first) — `notes.md` §44. Compaction envelope (encoder, not Jev): gliner25-compaction — mutating tools / shell operators prove `keep_full`; the model may only match that or be more -conservative (`notes.md` §50). +conservative (`notes.md` §50). Stdout-prune envelope (Jev): +jev-pruner — ≤10k / JSON-diff-whole-doc prove pass-through; Noul on +the remainder; fail-safe keep original (`notes.md` §53). ## Signal detection @@ -466,6 +468,7 @@ Use these as *existence proofs of a position*. Write your own card. | Inbox | reply / snooze / archive | urgency Noul + aboutness Choice | send, calendar | | Knowledge work | extract a quote / a cited fact | per-sentence or per-line-id Noul/Choice (**Empirical**: testimonial-miner, jev-reviewer) | verbatim join; place; human publish permission | | Agent context | compact completed tool results without inventing prose | retention Choice + char-offset locate (**Empirical**: gliner25-compaction; same *job* as fast-jev-compaction / pi-jev-compaction) | mutation/shell envelope → keep_full; fail-closed keep_full; shadowMode before replace; copy exact bytes | +| Agent context | prune Bash stdout before the LLM without inventing prose | Noul per chunk after a hard size/format envelope (**Empirical**: jev-pruner) | ≤10k / JSON-diff-whole-doc untouched; fail-safe original; archive dropped spans | | Dataframe labeling | classify / score rows | Noul/Choice/Score + full `p__` (**Empirical** as jevframe / jevpandas *shape*) | pandas/Polars, thresholds in code | | Computer-use speed | one verified act per step | score / Choice among numbered a11y/DOM controls (**Empirical**: solari-reflex Jev; gliner2-ultrafast GLiNER2; laya-mind2web Laya DOM indices) | Guard check; deny-list absence; no screenshots; `DONE` ≠ success | | Agent turn | skip memory tour on easy intent | intent Choice (**Empirical**: jev-hermes) | Memory still writes; complex still searches | diff --git a/.agents/skills/augustus/references/methods-catalog.md b/.agents/skills/augustus/references/methods-catalog.md index 52102c5..1dc59e7 100644 --- a/.agents/skills/augustus/references/methods-catalog.md +++ b/.agents/skills/augustus/references/methods-catalog.md @@ -59,7 +59,7 @@ judgment component is new). | Method | Judgment-shaped component | Jev substitution | Stays in code | Status | |---|---|---|---|---| | Entropy as uncertainty signal | Measuring "how spread is this belief" | Entropy of returned distributions across repeats or options — computed in code from returned probabilities | All arithmetic | **Empirical recipe** (cookbook pattern) | -| Detector / Neyman filter (context) | Is this artifact relevant to the current task? | One relevance Noul per block before it enters context; stub + recall key. Encoder cousin: retention Choice + span locate, copy verbatim. Indexer cousin: GLiNER extract on the bulk, escalate LLM on the tail | Cache, recall, safety keeps; mutation/shell envelope; do not dump the repo if S1 failed to load | **Empirical recipe** (winnow ≤0.22 hide; compaction 2-noul rule; pi-jev-context hide-not-delete; gliner25-compaction GLiNER2.5 keep_full/keep_evidence/keep_call_only/drop, `notes.md` §50; s1-graphify-indexer degraded-load / no-invent-edges, 10–50× unfilled, `notes.md` §51) | +| Detector / Neyman filter (context) | Is this artifact relevant to the current task? | One relevance Noul per block before it enters context; stub + recall key. Encoder cousin: retention Choice + span locate, copy verbatim. Indexer cousin: GLiNER extract on the bulk, escalate LLM on the tail. Stdout cousin: Noul per chunk after a hard size/format envelope | Cache, recall, safety keeps; mutation/shell envelope; ≤10k/JSON-diff pass-through; archive dropped spans; do not dump the repo if S1 failed to load | **Empirical recipe** (winnow ≤0.22 hide; compaction 2-noul rule; pi-jev-context hide-not-delete; gliner25-compaction GLiNER2.5 keep_full/keep_evidence/keep_call_only/drop, `notes.md` §50; s1-graphify-indexer degraded-load / no-invent-edges, 10–50× unfilled, `notes.md` §51; jev-pruner Bash stdout prune, `notes.md` §53) | | Anomaly detection | Does this deviate from expected shape? | Guard nouls + harm Score over {input, output, tool trace} | Baselines, alert thresholds | **Empirical recipe** (guardrails cookbook; pi-jev output judge) | | Allowlist ∩ remainder (code-then-model) | Unlisted / unstructured leftovers after a **proof** | Typed questions only on the unknown tier; admit iff every p < τ | Proven/refused in code; cannot block unless a sandbox sits under | **Empirical recipe** (jevgate 0/59 unsafe unasked held-out; allowlist *proves* read-only verbs; doc-router 1.74× $). Domain-general: `mappings.md` §18 | | Decision-token LoRA (constrained-AR) | Specialize a generator for parallel constrained fields | Loss only on the single decision token; KV broadcast across fields | Schema, candidate tokens, policy | **Empirical recipe** as Foodoo1 200-case / 4-field receipt (fraud_risk 64→95%, overall 85.2→98.8%, ~234 ms); **Hypothesis** as a general recipe. Synthetic; not a financial product. Softmax ≠ Noul | @@ -70,7 +70,7 @@ judgment component is new). | Method | Judgment-shaped component | Jev substitution | Stays in code | Status | |---|---|---|---|---| | Claim–evidence entailment (NLI) | supports / contradicts / not-established per claim–source pair | One Choice per pair + review flag; judge against the cited source text only. Stop-hook cousin: claims vs **session** evidence | Quote extraction, citation graph, audit log. When the answer *is* a span you already hold, **point** at line ids and copy verbatim — the model never writes the excerpt (jev-reviewer). Keyword retrieve is not semantic (clear-head) | **Empirical recipe** (citation_check cookbook; jev-reviewer sample study, `notes.md` §48; clear-head Stop, `notes.md` §51). Atlas: paraphrase_support and reversed_meaning_high_overlap correctly judged when both texts are in `state` (`notes.md` §49) | -| Extractive selection + offline re-threshold | Keep/drop over sentences, ids, character offsets, or observed DOM controls code already holds | Per-item Noul/Choice + one broadcast; join in order; `redecide` on the log with no new calls. Compaction: retention Choice + span locate. Computer-use: score among a11y/DOM candidates | Numbering, header skip, thresholds, publish permission; mutation envelope; copy exact bytes; click in code; never generate selectors | **Empirical recipe** (testimonial-miner 8-request fixture; jev-reviewer; gliner25-compaction char-offset copies, `notes.md` §48, §50; gliner2-ultrafast observe→score→act, `notes.md` §52). Cousin of applied-mappings §2 | +| Extractive selection + offline re-threshold | Keep/drop over sentences, ids, character offsets, observed DOM controls, or stdout chunks code already holds | Per-item Noul/Choice + one broadcast; join in order; `redecide` on the log with no new calls. Compaction: retention Choice + span locate. Computer-use: score among a11y/DOM candidates. Stdout prune: Noul per chunk after hard envelope | Numbering, header skip, thresholds, publish permission; mutation envelope; copy exact bytes; click in code; never generate selectors; archive dropped stdout | **Empirical recipe** (testimonial-miner 8-request fixture; jev-reviewer; gliner25-compaction char-offset copies, `notes.md` §48, §50; gliner2-ultrafast observe→score→act, `notes.md` §52; jev-pruner, `notes.md` §53). Cousin of applied-mappings §2 | | Combinatorial grid / program synthesis | Consistent whole-object from many cells | **Rejected as extractive.** Cell-wise Choice does not assemble ARC grids (4/400 Direct Jev) | Search, a program, a simulator | **Empirical as a negative** (`notes.md` §49) | | Spec vs artifact conformance (model checking *mindset*) | Property holds / violated / unverifiable for a named requirement | One Noul/Score per requirement, batched; violated → named rule back into context (pi-warden / Abide shape). This is **not** TLC/Apalache/GNATprove | Requirement enumeration, enforcement, logging; the **linter** if the rule is lintable; the real checker if you have one | **Empirical recipe** (pi-warden: 6→0 rule breaks, 150 paired runs; jev-pref: YOU define the rule; Abide: productized compile/calibrate/tune/replay, `notes.md` §47; if-ai: plain-English PR check, fail-closed on error, `notes.md` §51). Ownership split: `formal-methods.md` | | AST ∩ semantic lint | Semantic remainder after a parser already extracted units | Typed questions on Tree-sitter targets; do not execute scanned code | Parser, selection, fail-on; `tenbin` owns the lint *skill* | **Empirical as a shape** (jevscan 0.2.0rc4; not a calibration claim; `notes.md` §48) | diff --git a/.agents/skills/augustus/references/mixed-architecture.md b/.agents/skills/augustus/references/mixed-architecture.md index aa724c2..4a42af4 100644 --- a/.agents/skills/augustus/references/mixed-architecture.md +++ b/.agents/skills/augustus/references/mixed-architecture.md @@ -184,6 +184,7 @@ not a global virtue: |---|---|---| | Drop a RAG chunk or log line | **Fail open** (keep on error) | A false drop loses evidence; a false keep costs tokens | | Compact / drop a completed tool result | **Fail closed** to keep-full (`gliner25-compaction`) | Compaction is a destructive edit of memory. Uncertain *looks* like keep-on-error from the evidence side; name the *reduction* as the act. Contrast Abide / jevgate fail-open | +| Prune Bash stdout before the LLM | **Fail closed** to original (`jev-pruner`) | Dropping the log is irreversible. ≤10k / JSON-diff-whole-doc prove pass-through; archive/Jev/incomplete-score failure keeps the result. Harbor plugin-eval cannot reach Jev → cannot prune | | Skip waking a sleeping agent | **Fail open** (wake on error / unsure / no key) (`wakegate`) | Skip is the irreversible act. User-message, skip-limit, nothing-to-judge, and p in 0.2–0.5 all wake. Contrast pi-jev-approver fail-closed without a key | | Merge a red CI run | **Fail closed** on `--gate` (`latch`); reporter stays fail-open | False PASS merges a real bug. Missing key never fails Playwright; the gate is a separate step. Judge never says ignore alone | | Plain-English PR check | **Fail closed** on error / empty / low confidence (`if-ai`) | A skipped or timed-out check is not a pass. Threshold is policy, not measured correctness | @@ -203,6 +204,11 @@ Worked placements (2026-09-18 topic:jev hour + prior archive): [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) — GLiNER2.5 retention Choice + exact spans; fail closed to `keep_full`; `shadowMode` default true (`notes.md` §50). Not Jev. + Stdout-prune cousin, same family, different job: + [jev-pruner](https://github.com/tamaratran/jev-pruner) — Jev Noul on + residual noisy Bash after a hard ≤10k/format envelope; fail-safe + original; archive for recovery (`notes.md` §53). Marketplace id + still `fast-jev-output`. - **Diff hunks before `git add`** — `ibrahemid/git-jev-stage`: one Choice per hunk (`include` / `exclude` / `mixed`); mixed and low-confidence stay unstaged; lines never split; staging is an exact patch after confirm. @@ -437,7 +443,7 @@ decision-design card. Do not clone APIs from READMEs. | Hold-before-publish moderation | Hazard Nouls + harm Score | Block/review/pass policy | Near Here / firehose family | | Tool / engine / skill select | Choice + fits-Noul | Dispatch, auth, reject-all | skillranker, LlamaIndex selectors, Toolrouter | | Preference lint | Per-rule Score/Noul on a diff | Rule text, linter for hard rules, bands + fail-open | jev-pref (contract), Abide (productized), JevLint; if-ai (plain-English PR check, fail-closed on error); jev-marshal (Watch / empty repo) | -| Context / log prune | Per-line or per-block relevance; or a retention Choice + spans | Always-keep set, recall keys; mutation envelope in code; shadow before replace | jevprune, winnow; fast-jev-compaction / pi-jev-compaction (Jev); gliner25-compaction (GLiNER2.5) | +| Context / log prune | Per-line or per-block relevance; or a retention Choice + spans; or a Noul per stdout chunk | Always-keep set, recall keys; mutation envelope in code; shadow before replace; size/format envelope then Noul; archive dropped spans | jevprune, winnow; fast-jev-compaction / pi-jev-compaction (Jev session); gliner25-compaction (GLiNER2.5 session); jev-pruner (Jev Bash stdout) | | Exact hunk staging | Per-hunk include/exclude/mixed | `git diff`, atomic apply | git-jev-stage | | Semantic `WHERE` | Noul/`jev_prob` over a row | SQL, indexes, LIMIT | jevql (CLI; DB sees ordinary SQL); sqlite-jev (in-engine extension) | | Formula / query embedding | JUDGE as a function | Spreadsheet/SQL engine | judge-sheets, jevql, sqlite-jev | diff --git a/.agents/skills/augustus/references/toolbox-mapping.md b/.agents/skills/augustus/references/toolbox-mapping.md index 3cee2b8..789cc0f 100644 --- a/.agents/skills/augustus/references/toolbox-mapping.md +++ b/.agents/skills/augustus/references/toolbox-mapping.md @@ -77,8 +77,8 @@ component; keep the rest of the method in code. | Experimental design: Harbor on/off routing | Same coding-agent task with routing on vs off; hidden verifier; cheaper unsolved is not a saving | **Empirical as a *shape* and one-run signal** (jev-gateway-bench; `notes.md` §51). Pair CI merge-gate with Harbor + rh-guard | | Discrete math: width vs depth | Fan out in width (parallel ≈ free), pay depth linearly; two-stage only when next options depend on an earlier answer | **Empirical recipe** (fan-out: 12.2× cheaper, 10× faster) | | Psychology: Kahneman | System 2 generates/proposes (LLM), System 1 discriminates (Jev); never the reverse. S1 keeps control; optional S2 is one-use advice and does not fly. Productized cascade: `conf ≥ τ` → S1 decides else S2 writes; routing fails open / safety fails closed; **routing accuracy unmeasured**; keyword fallback ≠ S1. S1 specialists + S2 coordinator is the same split (reification-labs/foreman is description-only Phoenix scaffold this pass — do not invent an Elixir API). Indexer: S1 GLiNER extract on the bulk, escalate LLM on the tail (10–50× unfilled) | **Empirical recipe** (mcts-agent role split; 2026-09-18 mixed-architecture discourse; jev-reflex-autonomy-lab as a control-loop shape, `notes.md` §46; [dual-process-ai](https://github.com/taro1985/dual-process-ai) as a business/life cascade, `notes.md` §49; s1-graphify-indexer, `notes.md` §51). **Route ≠ memory:** a cheap intent gate skips memory/tool *tours* on easy routes; memory still writes; complex still searches (jev-hermes, `notes.md` §48) | -| IR / cascades | Cheap relevance / irrelevance before an expensive ranker or generator | **Empirical recipe** (RAG cookbook; jevprune; git-jev-stage; LlamaIndex Jev rerank) | -| IR / extractive keep-drop | Number candidates in code; model selects; copy verbatim; `redecide` thresholds on the log with no new calls. Compaction: same pointer job on tool results (Jev Noul/Score *or* GLiNER encoder). Computer-use: same pointer job on observed a11y/DOM controls (Jev *or* GLiNER2) | **Empirical recipe** (testimonial-miner; jev-reviewer pointer-not-generator, `notes.md` §48; gliner25-compaction char-offset + fail-closed keep_full, `notes.md` §50; gliner2-ultrafast observe→score→act, `notes.md` §52) | +| IR / cascades | Cheap relevance / irrelevance before an expensive ranker or generator | **Empirical recipe** (RAG cookbook; jevprune; git-jev-stage; LlamaIndex Jev rerank; jev-pruner stdout after size/format envelope, `notes.md` §53) | +| IR / extractive keep-drop | Number candidates in code; model selects; copy verbatim; `redecide` thresholds on the log with no new calls. Compaction: same pointer job on tool results (Jev Noul/Score *or* GLiNER encoder). Computer-use: same pointer job on observed a11y/DOM controls (Jev *or* GLiNER2). Stdout: same pointer job on Bash chunks after a hard envelope (Jev Noul) | **Empirical recipe** (testimonial-miner; jev-reviewer pointer-not-generator, `notes.md` §48; gliner25-compaction char-offset + fail-closed keep_full, `notes.md` §50; gliner2-ultrafast observe→score→act, `notes.md` §52; jev-pruner, `notes.md` §53) | | Spec / lint | Project-defined semantic rules as predicates over a diff; linter owns hard rules. AST remainder: Tree-sitter units, then typed questions; do not execute scanned code. Plain-English PR check: one condition + min-confidence; fail-closed on error | **Empirical recipe** (jev-pref contract; Abide productized path — replay 93 sessions, edit precision ~26% / turn ~73% before tune, `notes.md` §47; JevLint file-level Noul; pi-warden; snifftest unsure-band; jevscan AST∩semantic, `tenbin` owns the lint skill, `notes.md` §48; if-ai, `notes.md` §51). jev-marshal is Watch / empty this pass | | Formal methods / DST / safety | Judgment triages counterexamples, failing seeds, and named-rule conformance; proof/MC/DST stay with their tools. Alloy finder ≠ Apalache BMC ≠ Quint run. DST trio: Antithesis hypervisor / Resonate HQ Lean+oracle+SDK (durable async) / PufferLib env+seed. Noul is a sensor, not a discharged PO. Semi-formal diagrams are vocabularies, not enforcers | **Hypothesis as product**, **Contract** as ownership (matching `mappings.md` §8 and `methods-catalog.md`; worked shape pi-warden — `formal-methods.md`, `formal-semi-formal.md`) | | Decision analysis: VOI | Gather as an enumerated act; pay iff expected decision-loss drop > cost. Wake/resume: skip the LLM turn only if the judge answers and p is low (Horvitz); user-message / skip-limit already answer without a model | **Hypothesis** as calculator (`mappings.md` §6; `mental-models.md`). wakegate 21/21 is smoke (`notes.md` §51) | diff --git a/.agents/skills/augustus/references/validation.md b/.agents/skills/augustus/references/validation.md index dc46664..d550855 100644 --- a/.agents/skills/augustus/references/validation.md +++ b/.agents/skills/augustus/references/validation.md @@ -326,6 +326,7 @@ Rules: | Project soft-rule lint | [Abide](https://github.com/coldteadotai/abide) | Score per rule on the diff; bands; fail-open; replay + independent review | | Collab / computer-use product loop | [jev-testbench](https://github.com/ufx7/jev-testbench); [solari-reflex](https://github.com/hitakshiA/solari-reflex); [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) | Wilson/McNemar arms; independently checked task time; `DONE` ≠ success | | Agent routing on vs off | [jev-gateway-bench](https://github.com/vinilana/jev-gateway-bench) | Hidden perft; cost/quality; one-run signal this pass | +| Command-output prune (needle/noise) | [jev-pruner](https://github.com/tamaratran/jev-pruner) | Manual `trimOutput` sweep (theirs); plugin eval cannot reach Jev (fail-safe original); Terminal-Bench paired pilot is integration, not a full bench | rh-guard is a reward-hack hook, a different surface from jevgate and from Abide (eval-integrity vs allowlist-remainder vs project soft @@ -370,6 +371,20 @@ Jev is down. Pair CI merge-gate (frozen JUnit artifacts × PASS/BLOCK) and rh-guard (eval-integrity). Do not copy npm/ports (`notes.md` §51). +**Harbor-adjacent stdout prune (Empirical as README / evals README +behavior, not a full Terminal-Bench ranking; 2026-09-18 ~17:15).** +[jev-pruner](https://github.com/tamaratran/jev-pruner) ships +needle/noise graders and a Harbor Terminal-Bench 2.0 adapter in-repo. +Manual `trimOutput` sweep (theirs, 2026-09-18, 3 runs, `jev-latest`): +needles **24/24**; mean reduction **83% (71–92%)** on trim scenarios; +wrongly trimmed **0/12**; mean latency **240 ms**. Wider: standard +8/8 / 83%; accuracy 36/36 / 87%; real captures 10/10 / 54%; needle +matrix 9/9. `claude plugin eval` **cannot exercise pruning** (Jev +fetch refused → fail-safe original). Six-run paired Terminal-Bench +pilot is **integration, not a significance test** (full set 89 tasks +/ 178 trials; no published full-run scores this pass). Do not merge +those tables. Do not copy the Harbor launcher (`notes.md` §53). + **Harbor-style frozen protocol vs constrained LLMs (Empirical as that named receipt, not a ranking).** [`nibzard/decision-model-benchmark`](https://github.com/nibzard/decision-model-benchmark) diff --git a/CHANGELOG.md b/CHANGELOG.md index c774197..4c4c16a 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -239,6 +239,19 @@ Pegged against [`typesafe-ai/skills` v0.5.7](https://github.com/typesafe-ai/skil gliner25-compaction (different hole) and GLiGuard (safety schema). Demo (theirs, not re-run): Flights 12.20 s / 13.785 s / ~$0.0001 API — demonstration, not a bake-off. No invented metrics. +- jev-pruner evidence-preserving Bash stdout prune (`research/notes.md` + §53, [jev-pruner](https://github.com/tamaratran/jev-pruner), MIT): + architecture notes, not a plugin how-to. After Bash, Jev Noul-prunes + stdout chunks before the main LLM sees them — no summary. Hard + envelope (≤10k estimated tokens / JSON-diff-whole-doc untouched) + then soft Noul; fail-safe keep original; full archive. Marketplace + id still `fast-jev-output`. Codex is opt-in wrapper, not automatic + interception. Same evidence-preserving *family* as + fast-jev-compaction and gliner25-compaction; different *job* + (command output vs session memory) and Jev backend vs GLiNER2.5. + Manual sweep (theirs): needles 24/24; mean reduction 83% on trim + scenarios. Harbor plugin-eval cannot reach Jev. Terminal-Bench + paired pilot is integration, not a full bench. No invented metrics. - Effect-oriented loops (`notes.md` §28, `mappings.md` §19): Ward's ZIO client keeps Jev as the outer Choice and the handler as the effect. Not Effect.ts. GLiNER author: GLiNER2 "like jev" is GLiGuard diff --git a/README.md b/README.md index 2301bbe..3beaaa6 100644 --- a/README.md +++ b/README.md @@ -56,12 +56,13 @@ never launder a Noul as a proof. (Jev default / other family with self-eval); dual-process S1 decide / S2 generate; component node; DOM-as-text + fan-out; shadow-mode compaction rollout; fail-open wake vs fail-closed merge-gate; Harbor on/off routing; - hybrid local decide + remote fill; `DONE` ≠ verified success + hybrid local decide + remote fill; `DONE` ≠ verified success; + evidence-preserving stdout prune (hard envelope then Noul) - `.agents/skills/augustus/references/applied-mappings.md` — context sieve, - exact-text keep/drop (extractive / pointer-not-generator; char-offset compaction; observed a11y/DOM controls), env triage (OpenSmoke + latch merge-gate), moderation/ranking, skill routing (route ≠ memory) + exact-text keep/drop (extractive / pointer-not-generator; char-offset compaction; observed a11y/DOM controls; Bash stdout prune), env triage (OpenSmoke + latch merge-gate), moderation/ranking, skill routing (route ≠ memory) - `.agents/skills/augustus/references/faq.md` — "just classification", stack replacement, Jev vs open head vs encoder vs LoRA vs constrained AR vs kev vs blackwood, - wait-for-Archer, missing-other confident-wrong, soft project rules vs linter (Abide), extractive/pointer-not-generator, compaction summarize vs pointer, encoder vs Jev compaction, fail-closed keep_full, shadow-mode rollout, fail-open vs fail-closed wake vs CI gate, observe→score→act backend-agnostic, hybrid local decide + remote fill, DONE ≠ verified success, local drop-in vs stub scorer, route ≠ memory, when-it-holds / extractable-from-state, decision-model vs constrained LLM, dual-process S1/S2, combinatorial grid ≠ extractive, GLiNER vs GLiClass vs CLIP, LLM-as-judge, in-engine vs CLI store, + wait-for-Archer, missing-other confident-wrong, soft project rules vs linter (Abide), extractive/pointer-not-generator, compaction summarize vs pointer, encoder vs Jev compaction, fail-closed keep_full, shadow-mode rollout, fail-open vs fail-closed wake vs CI gate, observe→score→act backend-agnostic, hybrid local decide + remote fill, DONE ≠ verified success, stdout prune vs session compaction, local drop-in vs stub scorer, route ≠ memory, when-it-holds / extractable-from-state, decision-model vs constrained LLM, dual-process S1/S2, combinatorial grid ≠ extractive, GLiNER vs GLiClass vs CLIP, LLM-as-judge, in-engine vs CLI store, hard envelope (bitrate / planner), not-another-how-to - `.agents/skills/augustus/references/mappings.md` — classical-method mappings with boundaries, counterexamples, acceptance tests (including @@ -76,7 +77,8 @@ never launder a Noul as a proof. computer-use; gliner2-ultrafast encoder-backend cousin (`DONE` ≠ success; demo is not a bake-off); jev-testbench collab arms; ARC-AGI Direct Jev as combinatorial-≠-extractive negative; jev-gateway-bench Harbor on/off - routing one-run signal) + routing one-run signal; jev-pruner Harbor needle/noise + Terminal-Bench + integration pilot, not a full bench) - `.agents/skills/augustus/references/boundary-audit.md` — existing-system insertion: fit test, opportunity map, smallest boundary, red flags - `.agents/skills/augustus/scripts/evaluate_decisions.py` — offline evaluator diff --git a/docs/ecosystem.md b/docs/ecosystem.md index 8b23ee9..c5249fa 100644 --- a/docs/ecosystem.md +++ b/docs/ecosystem.md @@ -77,7 +77,7 @@ Per-keystroke launchers (104ms median, sequence-tagged staleness), firehose mode ### Agent harnesses extras (this hour) - **kevinpita/pi-jev-context** — reversible Pi context sieve: hide, do not delete; `/jev off` restores. Cousin of winnow/jevprune. -- **vava-nessa/pi-jev-compaction** (and `tamaratran/fast-jev-compaction`) — verbatim drop, never summarize. Pair with jev-gate-student-b for local memory-gating. Same *job* as [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) (GLiNER2.5 encoder backend; `notes.md` §50). +- **vava-nessa/pi-jev-compaction** (and `tamaratran/fast-jev-compaction`) — verbatim drop, never summarize. Pair with jev-gate-student-b for local memory-gating. Same *job* as [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction) (GLiNER2.5 encoder backend; `notes.md` §50). Stdout cousin [jev-pruner](https://github.com/tamaratran/jev-pruner) — same family, prune Bash before the LLM, not session memory (`notes.md` §53). - **reachjalil/jev-tree** — authored taxonomy so each Choice stays under 255; truncate silently drops the tail (`jev-tree-choice-cap`). - **HacksonClark / SREGym-Lite** — Jev ranks next tests/evidence; does not diagnose; 20/50→24/50 with 2 regressions. `notes.md` §33. - **ddfeyes/jev-mode** — bulk triage/tag/route off frontier context; synthetic 1,000: −77.8% tokens; accuracy is parity. `notes.md` §42. @@ -179,6 +179,12 @@ Architecture notes, not a browser-agent catalog. `notes.md` §52. TypeSafe Jev i - **sahibzada-allahyar/gliner2-ultrafast** — local GLiNER2 (`fastino/gliner2-multi-v1`) scores observed a11y/DOM controls. Adaptation of jev-ultrafast. No screenshots; no generated selectors; code owns actuators. Hybrid local decide + remote Mercury 2.5 fill. `DONE` ≠ verified success. Same *job* as jev-ultrafast / solari-reflex; encoder backend. Contrast blackwood-rlcd (screenshot + marked letters). Cousin: ShaunSpark/laya-mind2web-browser-agent (Laya over DOM indices). Fastino sibling class with gliner25-compaction (different hole) and GLiGuard (safety schema). Demo (theirs, not re-run): Flights 12.20 s visible / 13.785 s loop / ~$0.0001 API — demonstration, not a bake-off. MIT. +### Hourly ~17:15 Boise (jev-pruner evidence-preserving Bash stdout prune) + +Architecture notes, not a plugin catalog. `notes.md` §53. TypeSafe Jev is the exemplar, not the monopoly. **Not a summarizer. Not session compaction. Not GLiNER.** Archer still Watch. + +- **tamaratran/jev-pruner** — after Bash, Jev Noul-prunes stdout chunks before the main LLM sees them. No summary. Hard envelope (≤10k estimated tokens; JSON/XML/YAML/diff/binary; whole-document commands) then soft Noul. Fail-safe keep original; full archive. Marketplace id still `fast-jev-output`. Codex is opt-in wrapper, not automatic interception. Same evidence-preserving *family* as fast-jev-compaction / gliner25-compaction; different *job* (command output vs session memory). Manual sweep (theirs): needles 24/24; mean reduction 83% on trim scenarios. Harbor plugin-eval cannot reach Jev. Terminal-Bench paired pilot is integration, not a full bench. MIT. + See `references/mixed-architecture.md` in the skill. Class-level family choice: `references/judgment-class.md`. Proof vs judgment (Alloy vs Apalache; DST trio Antithesis / Resonate HQ / PufferLib): diff --git a/research/archive/findings.md b/research/archive/findings.md index 686a4b1..4235b3a 100644 --- a/research/archive/findings.md +++ b/research/archive/findings.md @@ -998,6 +998,39 @@ screenshot multimodal is a different input from DOM-as-text; (cg) local decide + remote TYPE is mixed-architecture economics, not dual-process-ai; (ch) `DONE` ≠ Harbor-verified success. +## Batch #37 (2026-09-18 ~17:15 Boise) — jev-pruner evidence-preserving Bash stdout prune + +Note: `research/notes.md` §53. Docs-only. Folded into PR #2. Not a +second PR. Archer still Watch. **Not a summarizer. Not session +compaction. Not GLiNER.** No invented metrics. Do not re-fold §50 +gliner25-compaction as this product, §52 observe→score-act, or +fast-jev-compaction as a duplicate. + +- **tamaratran/jev-pruner (Empirical as README / evals README + behavior).** MIT. Created 2026-09-18T03:00:58Z; 5★ attached + capture, 7★ live this pass. After Bash, Jev Noul-prunes stdout + chunks before the main LLM sees them. No summary. Hard envelope + (≤10k estimated tokens; JSON/XML/YAML/diff/binary; whole-document + commands) then soft Noul. Fail-safe keep original; full archive. + Marketplace id still `fast-jev-output`. Codex is opt-in wrapper. + Manual sweep (theirs, 2026-09-18): needles 24/24; mean reduction + 83% (71–92%) on trim scenarios; wrongly trimmed 0/12; 240 ms. + Harbor plugin-eval cannot reach Jev (fail-safe). Terminal-Bench + paired pilot is integration, not a full bench. +- **Mental models:** (1) evidence-preserving prune ≠ summarizer + (family with gliner25-compaction / jev-reviewer); (2) hard + size/format envelope then soft Noul; (3) fail-safe keep original + (reduction is the irreversible act); (4) stdout prune vs session + compaction are different jobs; host capability shapes the product. +- **rh-guard:** sibling note only (fail-safe / envelope). Not + reward-hack detection. + +Cross-repo addition: (ci) command-output sieve and session compaction +share extractive honesty, not a product; (cj) structure-first +pass-through (≤10k / JSON-diff) is the sandwich, Jev is the remainder; +(ck) Harbor plugin-eval refusing Jev is a fail-safe receipt, not a +missing metric; (cl) marketplace id may lag the repo name. + diff --git a/research/notes.md b/research/notes.md index 89ef549..d7c3095 100644 --- a/research/notes.md +++ b/research/notes.md @@ -3874,3 +3874,144 @@ Not multimodal pixels. Strong **composition / open-weights decide** exemplar for browser computer-use (local encoder + remote fill). Archive + landscape. Harbor-style independent verify is already in their framing. + +## 53. jev-pruner — evidence-preserving Bash stdout prune (2026-09-18 ~17:15 Boise) + +America/Boise ~17:15 = 23:15 UTC. Docs-only fold into open PR #2 +(`cursor/augustus-store-envelope-00b4`). Not a competing PR. Archer +27B drop still **WATCH**. Identity lock vs `typesafe-ai` / `tenbin` / +`decision-first` holds. No wrapper, no marketplace / Codex install +how-to, no copied `keepThreshold` as a class constant. No invented +metrics. **Not a summarizer. Not session compaction. Not GLiNER.** +Do not re-fold §50 gliner25-compaction as this product, §52 +observe→score-act, §48 jevprune/winnow as a new species, or +fast-jev-compaction as a duplicate. + +Family: **evidence-preserving reduce** (pointer/extractive, dropped +bytes recoverable). Same instinct as +[`tamaratran/fast-jev-compaction`](https://github.com/tamaratran/fast-jev-compaction) +and [`m-newhauser/gliner25-compaction`](https://github.com/m-newhauser/gliner25-compaction) +(`notes.md` §50). **Different job:** prune a just-run Bash result +*before* the main LLM sees it, not compact completed tool pairs +already in session history. **Different backend vs GLiNER2.5:** +TypeSafe Jev Noul, not an encoder retention Choice. Same author as +fast-jev-compaction; the READMEs say the two are independent and can +be installed together. Marketplace / Claude plugin id is still +`fast-jev-output` — name ≠ id. + +### HIGH + +1. **[`tamaratran/jev-pruner`](https://github.com/tamaratran/jev-pruner)** + (MIT; TypeScript; created 2026-09-18T03:00:58Z; 5★ at attached + capture, 7★ live this pass). Claude Code plugin: after Bash runs, + **before** the result is sent back to the main LLM, Jev + Noul-prunes stdout **without generating a summary**. Codex is an + **opt-in wrapper/skill**, not automatic `PostToolUse` + interception — the host cannot replace native shell output that + way. **Work in progress / measurement-native.** **Not a prose + compressor. Not a screenshot VLM. Not a GLiNER backend.** + + Load-bearing loop (README): + + ```text + Claude requests Bash → command runs → Jev prunes stdout → Claude receives result + (+ archive path for dropped spans) + ``` + + One Noul per chunk: “does any line in this chunk need to remain + available?” A single needed line protects the chunk. Chunks of + `chunkLines` lines (default 20), capped at 200 chunks. Categories + (build/test, search/excerpt) add guidance only — they never mark + a whole command disposable or change the keep threshold. + + **Four load-bearing mental models (architecture, not a plugin + catalog):** + + 1. **Evidence-preserving prune, not summarize.** Same extractive + honesty as gliner25-compaction / jev-reviewer / testimonial-miner: + the model **selects**; code **keeps verbatim chunks**. Dropped + text is archived under `.claude/fast-jev-output/` (Claude) or + `.jev-pruner/` (Codex) *before* the first scoring request; + markers point at the recovery path. A generator summary of + stdout is a different species. Credential-like commands/output + are **not** archived; markers tell the agent to re-run. That + check only skips local archive — it does **not** redact + secrets from Jev (`notes.md` this section). + + 2. **Hard envelope, then soft Noul.** Code proves pass-through + *before* Jev runs: ≤10,000 estimated tokens (`estimateTokens` + on raw stdout; `minTokens` can raise this, not lower it); + errors; JSON/XML/YAML/diff/binary; whole-document commands + (`cat`, `jq`, `git diff`, `git show`, `base64`, `openssl`). + Format detection beats a build/search category. Jev only + scores the residual noisy log. Same sandwich as + bitrate-advisor / jevgate / gliner25 mutation monitor + (`mappings.md` §12 / §18): structure first, remainder judged. + + 3. **Fail-safe keep original.** Archive write failure, Jev + failure, unfit state, incomplete scoring against every history + segment, first/last chunks, error/warning patterns → original + stdout untouched. The *reduction* is the irreversible act, so + uncertainty fails closed to keep-full — same polarity as + gliner25 `keep_full`, opposite slogan from Abide / jevgate + fail-open. Plugin-eval receipt of that envelope: Harbor's + `CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC` refuses the Jev + fetch, the hook falls back to original output, and the suite + **cannot exercise pruning** (evals README). That is the + fail-safe working, not a missing metric. + + 4. **Host capability shapes the product.** Claude: automatic + `tool.call` wrap of Bash `next()`. Codex CLI 0.152.1 cannot + replace native shell output from `PostToolUse`, so the product + is a wrapper + skill. Same judgment engine; different + insertion. Do not copy the wrapper. + + **Eval (theirs, not re-run; 2026-09-18).** Manual `trimOutput` + sweep, 3 runs/scenario, `jev-latest`: needles kept **24/24**; + mean reduction **83% (71–92%)** on scenarios meant to trim; + wrongly trimmed **0/12** pass-through; mean latency **240 ms**. + Wider sweeps: standard 8/8 / 83%; accuracy 36/36 / 87%; real + captures 10/10 / 54% (three correctly left whole); needle matrix + 9/9. README demonstration: a 76,379-char log whose 2,227-char + preview missed the error line became 4,013 chars of pruned + output that kept it — illustration, not a bake-off. Harbor + Terminal-Bench 2.0 adapter + six-run paired pilot is + **integration, not a full benchmark or significance test** (full + set is 89 tasks / 178 trials; no published full-run scores this + pass). Plugin `claude plugin eval` cases predate the 10k gate + and cannot reach Jev. Do not merge those tables. `keepThreshold` + default 0.5 is **their** knob. + + **Siblings — complementary, do not merge.** + + - **`tamaratran/fast-jev-compaction`:** same author; session + compaction of completed tool pairs (two Nouls). Different job. + - **`m-newhauser/gliner25-compaction`:** same evidence-preserving + *family*; GLiNER2.5 encoder backend; session compaction; + `shadowMode` default true (`notes.md` §50). + - **`ibrahemid/jevprune` / winnow:** per-line / per-block + relevance before context. Same *sieve* hole; this product adds + the 10k/format envelope + archive + Harbor harness. + - **`24601/rh-guard`:** light note only (fail-safe / envelope). + Not reward-hack detection. + + **Placement.** Context sieve + exact-text keep/drop + (`applied-mappings.md` §1–§2) + mixed architecture (code owns + envelope and archive; Jev scores residual chunks; LLM never + writes the kept bytes). Pillar: selective classification / SDT + (false drop >> false keep) + runtime-assurance sandwich. Hole: + sieve / keep-drop. Family: TypeSafe Jev (Noul). Fail-closed on + the reduction. Eval path: their manual sweep + in-repo Harbor + adapter (pilot ≠ full bench). **Empirical** as README / evals + README behavior. **Hypothesis** that the envelope transfers to + *your* command mix. Cards: `applied-mappings.md` §1 (primary); + `mixed-architecture.md`; `mappings.md` §12 / §18; `faq.md`; + `validation.md`; `judgment-class.md`; `mental-models.md`; + `methods-catalog.md`; `toolbox-mapping.md`; + `agent-self-assessment.md`. No wrapper. + +### Omni / Jev-omni + +Not multimodal. Usage: command-output sieve as a perception/memory +hygiene stage **before** the generative turn. Harbor-adjacent +eval harness in-repo. Archive + landscape pointer. diff --git a/research/refresh-log.md b/research/refresh-log.md index f6913d6..41cbcea 100644 --- a/research/refresh-log.md +++ b/research/refresh-log.md @@ -562,6 +562,27 @@ README. - notes.md §52; sources.json; findings.md batch #36. No wrapper. +## 2026-09-18 23:15 UTC — jev-pruner evidence-preserving Bash stdout prune (~17:15 Boise) + +- Folded into open PR #2 (`cursor/augustus-store-envelope-00b4`). + Docs-only. Not a competing PR. Archer 27B drop still **WATCH**. + **Not a summarizer. Not session compaction. Not GLiNER.** No + invented metrics. No wrapper. +- HIGH: [tamaratran/jev-pruner](https://github.com/tamaratran/jev-pruner) + (MIT, created 2026-09-18). After Bash, Jev Noul-prunes stdout + chunks before the main LLM sees them. Hard ≤10k / JSON-diff + envelope then soft Noul. Fail-safe keep original. Full archive. + Marketplace id still `fast-jev-output`. Codex opt-in wrapper. + Same *family* as fast-jev-compaction / gliner25-compaction; + different *job*. Manual sweep (theirs): needles 24/24; mean + reduction 83% on trim scenarios. Harbor plugin-eval cannot reach + Jev. Terminal-Bench paired pilot is integration, not a full bench. +- Cards: SKILL.md, applied-mappings §1 (primary), mixed-architecture, + mappings §12/§18, faq, validation, judgment-class, mental-models, + methods-catalog, toolbox, agent-self-assessment, composition-algebra, + ecosystem, CHANGELOG, README. +- notes.md §53; sources.json; findings.md batch #37. No wrapper. + diff --git a/research/sources.json b/research/sources.json index 0e5a599..3b26650 100644 --- a/research/sources.json +++ b/research/sources.json @@ -1,6 +1,6 @@ { "refresh_cadence": "hourly", - "retrieved": "2026-09-18T22:56Z", + "retrieved": "2026-09-18T23:15Z", "sources": [ { "kind": "docs", @@ -1592,6 +1592,12 @@ "title": "ShaunSpark/laya-mind2web-browser-agent", "url": "https://huggingface.co/ShaunSpark/laya-mind2web-browser-agent", "note": "Apache-2.0. Laya (convaiinnovations/laya) fine-tune: operation + target index over interactive DOM elements (Mind2Web). Observed-candidate family with gliner2-ultrafast / solari-reflex / jev-ultrafast; not screenshot multimodal (contrast blackwood-rlcd). Author-reported 74.3% on 68 held-out \u2014 small n, not a ranking. Cousin pointer only. notes.md \u00a752." + }, + { + "kind": "github", + "title": "tamaratran/jev-pruner", + "url": "https://github.com/tamaratran/jev-pruner", + "note": "MIT. Created 2026-09-18T03:00:58Z; 5 stars at attached capture, 7 live this pass. Claude Code plugin (marketplace id still fast-jev-output): after Bash, Jev Noul-prunes stdout chunks before the main LLM sees them \u2014 no summary. Hard envelope (\u226410k estimated tokens / JSON-diff-whole-doc untouched) then soft Noul; fail-safe keep original; full archive. Codex is opt-in wrapper. Same evidence-preserving family as fast-jev-compaction / gliner25-compaction; different job (command output vs session memory). Manual sweep (theirs): needles 24/24; mean reduction 83% on trim scenarios. Harbor plugin-eval cannot reach Jev. Terminal-Bench paired pilot is integration, not a full bench. notes.md \u00a753." } ] } From d2d9ea3d02dd3038f989551cc9b964efae074a1d Mon Sep 17 00:00:00 2001 From: Cursor Agent Date: Fri, 18 Sep 2026 23:27:49 +0000 Subject: [PATCH 12/43] Fold Cua-S1 as specialist form System One computer-use MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Docs-only: trycua/cua libs/cua-s1 option-attention among observed elements (fill/check/click/skip). Plan ≠ execute; dry-run default; source-only. Not TypeSafe Jev — parallel S1 naming in CUA. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com> --- .agents/skills/augustus/SKILL.md | 16 +- .../references/agent-self-assessment.md | 4 + .../augustus/references/applied-mappings.md | 6 +- .../references/composition-algebra.md | 2 + .agents/skills/augustus/references/faq.md | 29 +++- .../augustus/references/judgment-class.md | 11 ++ .../skills/augustus/references/mappings.md | 14 ++ .../augustus/references/mental-models.md | 4 +- .../augustus/references/methods-catalog.md | 2 +- .../augustus/references/mixed-architecture.md | 5 +- .../augustus/references/toolbox-mapping.md | 2 +- .../skills/augustus/references/validation.md | 11 +- CHANGELOG.md | 16 ++ README.md | 8 +- docs/ecosystem.md | 8 +- research/archive/findings.md | 37 +++++ research/notes.md | 143 ++++++++++++++++++ research/refresh-log.md | 23 +++ research/sources.json | 8 +- 19 files changed, 321 insertions(+), 28 deletions(-) diff --git a/.agents/skills/augustus/SKILL.md b/.agents/skills/augustus/SKILL.md index 987c985..3450eae 100644 --- a/.agents/skills/augustus/SKILL.md +++ b/.agents/skills/augustus/SKILL.md @@ -1,6 +1,6 @@ --- name: augustus -description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill, open multimodal RLCD / blackwood), announced open decision-model (Watch — still not landed), constrained-AR (TypeAR, pcdServer, decision-token LoRA), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"wait for Archer vs open multimodal RLCD\", \"screenshot/DOM candidates → typed Choice\", \"eval path\", \"jevals\", \"Harbor taskset\", \"shared bake-off ECE/NLL/Brier\", \"LLM-as-judge is not the System One score\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge (allowlist *proves*; fail-open cannot block), \"missing other → confident wrong Choice\", \"lint the Jev request\", \"training confronts Choice other / none-of-the-above\", \"S1 reflex keeps control / optional S2 one-use advice\", \"soft AGENTS.md rules vs linter\" (Abide / jev-pref), \"edit-phase vs turn-phase observation window\", banded confidence fail-open preference lint, \"extractive selection / pointer-not-generator\", \"encoder GLiNER compaction vs Jev Score compaction (same job; pointer not summarizer)\", \"fail-closed keep_full under mutation envelope\", \"CI flaky-vs-real merge gate\", \"fail-open VOI wake/resume (Horvitz)\", \"claim/evidence Stop integrity\", \"S1 extract + escalate-S2 indexer\", \"Harbor on/off routing\", \"policy-as-judgment PR marshal\", \"shadow-mode compaction rollout\", \"Jev Ultrafast vs GLiNER Ultrafast (observe-score-act backend-agnostic)\", \"hybrid local decide + remote fill\", \"DONE ≠ verified success\", \"observed a11y/DOM candidates vs screenshot multimodal\", \"evidence-preserving stdout prune (not summarize)\", \"hard token/format envelope then soft Noul\", \"fail-safe keep original on prune failure\", \"stdout prune vs session compaction\", \"local /v1/systemone drop-in (stub until hf scorer)\", \"dataframe-native semantic columns\", \"route≠memory\", \"advisory sidecar receipts\", \"structure induction over bags\", \"AST ∩ semantic lint\", \"extractable-from-state / retrieve first\", \"decision-model vs constrained-LLM bake-off\", \"dual-process S1 decide / S2 generate\", \"combinatorial grid ≠ extractive\", TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, or \"formally verify with Jev\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." +description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill, open multimodal RLCD / blackwood), announced open decision-model (Watch — still not landed), constrained-AR (TypeAR, pcdServer, decision-token LoRA), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"wait for Archer vs open multimodal RLCD\", \"screenshot/DOM candidates → typed Choice\", \"eval path\", \"jevals\", \"Harbor taskset\", \"shared bake-off ECE/NLL/Brier\", \"LLM-as-judge is not the System One score\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge (allowlist *proves*; fail-open cannot block), \"missing other → confident wrong Choice\", \"lint the Jev request\", \"training confronts Choice other / none-of-the-above\", \"S1 reflex keeps control / optional S2 one-use advice\", \"soft AGENTS.md rules vs linter\" (Abide / jev-pref), \"edit-phase vs turn-phase observation window\", banded confidence fail-open preference lint, \"extractive selection / pointer-not-generator\", \"encoder GLiNER compaction vs Jev Score compaction (same job; pointer not summarizer)\", \"fail-closed keep_full under mutation envelope\", \"CI flaky-vs-real merge gate\", \"fail-open VOI wake/resume (Horvitz)\", \"claim/evidence Stop integrity\", \"S1 extract + escalate-S2 indexer\", \"Harbor on/off routing\", \"policy-as-judgment PR marshal\", \"shadow-mode compaction rollout\", \"Jev Ultrafast vs GLiNER Ultrafast (observe-score-act backend-agnostic)\", \"hybrid local decide + remote fill\", \"DONE ≠ verified success\", \"observed a11y/DOM candidates vs screenshot multimodal\", \"evidence-preserving stdout prune (not summarize)\", \"hard token/format envelope then soft Noul\", \"fail-safe keep original on prune failure\", \"stdout prune vs session compaction\", \"specialist S1 computer-use (Cua-S1 form-v0; not TypeSafe Jev)\", \"plan ≠ execute / dry-run default\", \"observed-element option head (fill/check/click/skip)\", \"local /v1/systemone drop-in (stub until hf scorer)\", \"dataframe-native semantic columns\", \"route≠memory\", \"advisory sidecar receipts\", \"structure induction over bags\", \"AST ∩ semantic lint\", \"extractable-from-state / retrieve first\", \"decision-model vs constrained-LLM bake-off\", \"dual-process S1 decide / S2 generate\", \"combinatorial grid ≠ extractive\", TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, or \"formally verify with Jev\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." license: MIT metadata: version: 0.3.0 @@ -54,7 +54,7 @@ classical method you already trust, substitute it, classify the win "paraphrase brittleness", "allowlist then judge", "TOCTOU-of-Noul", "Jev inside the database / sqlite-jev", "Jev picks bitrate / join order / the model", "wait for Archer", "lint the request / missing - other", "training confronts Choice other / none-of-the-above", "soft AGENTS.md rules vs the linter", "screenshot Choice / omni System One", "extractive quotes / pointer not generator", "compaction summarize vs pointer", "encoder vs Jev compaction backend", "shadow-mode compaction rollout", "CI flaky-vs-real merge gate", "fail-open VOI wake/resume", "claim vs session evidence", "S1 indexer escalate-S2", "Harbor on/off routing", "fail-open vs fail-closed wake vs CI gate", "encoder vs Jev computer-use backend", "hybrid local decide + remote fill", "DONE vs verified success", "stdout prune vs session compaction", "hard envelope then Noul prune", "local drop-in vs stub scorer", "route vs memory", "when does it hold / extractable from state", "decision model vs constrained LLM", "dual-process S1/S2", or "combinatorial grid vs extractive": + other", "training confronts Choice other / none-of-the-above", "soft AGENTS.md rules vs the linter", "screenshot Choice / omni System One", "extractive quotes / pointer not generator", "compaction summarize vs pointer", "encoder vs Jev compaction backend", "shadow-mode compaction rollout", "CI flaky-vs-real merge gate", "fail-open VOI wake/resume", "claim vs session evidence", "S1 indexer escalate-S2", "Harbor on/off routing", "fail-open vs fail-closed wake vs CI gate", "encoder vs Jev computer-use backend", "hybrid local decide + remote fill", "DONE vs verified success", "stdout prune vs session compaction", "hard envelope then Noul prune", "Cua-S1 vs TypeSafe Jev", "plan vs execute dry-run", "specialist computer-use vs general agent", "local drop-in vs stub scorer", "route vs memory", "when does it hold / extractable from state", "decision model vs constrained LLM", "dual-process S1/S2", or "combinatorial grid vs extractive": read `references/faq.md`, then `references/mental-models.md`, then `references/mixed-architecture.md`, then @@ -114,13 +114,13 @@ classical method you already trust, substitute it, classify the win | Familiar method | Judgment shape | Detail | |---|---|---| | Mental models across domains (not SWE-only) | EU, abstention, VOI, MCDA, SDT, search/control, Leveson, NATM/Norman/snap-fit; **extractable-from-state boundary map** (self-contained vs needs outside knowledge) | `references/mental-models.md` | -| Judgment-model class (Jev is exemplar, not monopoly) | Species: decide / locate (GLiNER) / categorize (GLiClass) / rank / perceive; open heads include encoder DeBERTa, LoRA distill, openjev-lm, kev. Compaction job is backend-agnostic (Jev Score/Noul vs GLiNER2.5 encoder). Indexer cousin: GLiNER extract + escalate-S2 (10–50× unfilled). Computer-use observe→score-among-candidates→code-acts is backend-agnostic (Jev Ultrafast ↔ GLiNER2 Ultrafast) | `references/judgment-class.md` | +| Judgment-model class (Jev is exemplar, not monopoly) | Species: decide / locate (GLiNER) / categorize (GLiClass) / rank / perceive; open heads include encoder DeBERTa, LoRA distill, openjev-lm, kev. Compaction job is backend-agnostic (Jev Score/Noul vs GLiNER2.5 encoder). Indexer cousin: GLiNER extract + escalate-S2 (10–50× unfilled). Computer-use observe→score-among-candidates→code-acts is backend-agnostic (Jev Ultrafast ↔ GLiNER2 Ultrafast ↔ Cua-S1 specialist; Cua-S1 is not TypeSafe Jev) | `references/judgment-class.md` | | Open weights vs constrained decoding vs proprietary API | Three open paths: encoder open-jev / AR constrained decode (TypeAR + pcdServer; decision-token LoRA; packed one-forward logprob on open LLMs) / trained decision-only (Laya + ONNX port, Nimble, kev, **blackwood-rlcd** multimodal now, Archer Watch still Watch). Local `/v1/systemone` surfaces: jev-local (stub until `hf`), kev (trained pointer), von (tiny SAN — not a replica). Softmax over allowed tokens ≠ Noul | `references/judgment-class.md` (when-to-use table) | | Entropy as allocator (low / medium / high) | Typed low+medium decisions → System One marginals; high-entropy synthesis → frontier decoder. Product rhetoric, not a meter. **Hypothesis** | `references/judgment-class.md` | | Formal / semi-formal (proof vs judgment) | Sensor vs constraint vs searchlight; Alloy vs Apalache; DST trio; TOCTOU-of-Noul, AI×FM | `references/formal-methods.md` (one-screen: `references/formal-semi-formal.md`) | -| Mixed architecture (judgment model + LLM) | Provider judges, LLM writes, code owns control; not a stack replacement. Advisory sidecar never changes host routing. Dual-process: S1 decides, S2 generates (routing accuracy unmeasured); S1 specialists + S2 coordinator is the same split (description-only greenfield this hour). Internals ≠ FSM; placement is a **component node**. Browser-use strength = DOM-as-text + speculative fan-out, not vision. Fail polarity is per act: skip-wake fail-open vs merge-gate BLOCK fail-closed. Hybrid local decide + remote fill; `DONE` ≠ verified success | `references/mixed-architecture.md` | +| Mixed architecture (judgment model + LLM) | Provider judges, LLM writes, code owns control; not a stack replacement. Advisory sidecar never changes host routing. Dual-process: S1 decides, S2 generates (routing accuracy unmeasured); S1 specialists + S2 coordinator is the same split (description-only greenfield this hour). Internals ≠ FSM; placement is a **component node**. Browser-use strength = DOM-as-text + speculative fan-out, not vision. Fail polarity is per act: skip-wake fail-open vs merge-gate BLOCK fail-closed. Hybrid local decide + remote fill; `DONE` ≠ verified success. Specialist computer-use: plan ≠ execute, dry-run default (Cua-S1; not TypeSafe Jev) | `references/mixed-architecture.md` | | Context sieve | Relevance Noul per block; always-keep set in code; stub + recall key. Encoder cousin: GLiNER2.5 retention Choice + char-offset spans (gliner25-compaction); fail-closed keep_full; shadowMode default. Stdout cousin: jev-pruner (Jev Noul after hard ≤10k/JSON-diff envelope; fail-safe original; archive) | `references/applied-mappings.md#1-context-sieve` | -| Exact-text keep / drop | Choice include/exclude/mixed over candidates code already holds. Extractive quotes / pointer-not-generator (model never writes the excerpt; char-offset compaction same species). Observed a11y/DOM controls: score among them; code clicks (Jev or GLiNER2) | `references/applied-mappings.md#2-exact-text-keep--drop` | +| Exact-text keep / drop | Choice include/exclude/mixed over candidates code already holds. Extractive quotes / pointer-not-generator (model never writes the excerpt; char-offset compaction same species). Observed a11y/DOM controls: score among them; code clicks (Jev or GLiNER2 or Cua-S1 option-attention) | `references/applied-mappings.md#2-exact-text-keep--drop` | | Environment / harness triage | Scan every step for env failure; LLM autopsy only on flags. Merge-gate cousin: cluster in code, judge labels cause, policy owns PASS/BLOCK (latch; judge never says ignore alone) | `references/applied-mappings.md#3-environment--harness-triage` | | Moderation and ranking | Hold-before-publish vs graded rerank; fail policy per action | `references/applied-mappings.md#4-moderation-and-ranking` | | Skill / tool routing | Choice over a closed catalog + whether-anything-fits; code dispatches. Route ≠ memory: cheap intent gate skips memory tours on easy routes | `references/applied-mappings.md#5-skill--tool-routing` | @@ -133,7 +133,7 @@ classical method you already trust, substitute it, classify the win | Decision tables / circuits / state machines | Judgment predicates, code owns transitions | `references/mappings.md#3-semantic-predicates--decision-circuits` | | Retrieve + expensive relevance fn | Bounded rerank of a retrieved shortlist | `references/mappings.md#4-retrieval--bounded-semantic-reranking` (independent TREC DL2019 benchmark: Jev zero-shot best MAP 0.4748, nDCG@10 0.683 vs tuned monoBERT 0.718 — competitive, not dominant) | | Store as semantic index (SQL / SQLite / zoxide / dataframe) | Cheap exact predicates first; typed questions on the remainder. In-engine extension (sqlite-jev) vs CLI rewrite (jevql) vs path index (joxide) vs dataframe columns (jevpandas / jevframe) | `references/mappings.md#4-retrieval--bounded-semantic-reranking` | -| Soft judgment inside a hard envelope | Model may only match the deterministic policy or be more conservative (bitrate ABR; query-planner override-when-confident; compaction mutations/shell operators → keep_full; stdout prune: ≤10k/JSON-diff-whole-doc untouched, then Noul) | `references/mappings.md#12-runtime-assurance-sandwich-hypothesis`; `references/mappings.md#18-structural-prove--soft-remainder-hypothesis-as-domain-general-empirical-as-named-shapes` | +| Soft judgment inside a hard envelope | Model may only match the deterministic policy or be more conservative (bitrate ABR; query-planner override-when-confident; compaction mutations/shell operators → keep_full; stdout prune: ≤10k/JSON-diff-whole-doc untouched, then Noul; Cua-S1: plan≠execute, dry-run, fail-closed checkbox/fill) | `references/mappings.md#12-runtime-assurance-sandwich-hypothesis`; `references/mappings.md#18-structural-prove--soft-remainder-hypothesis-as-domain-general-empirical-as-named-shapes` | | Value of information / gather as an act | Pay for another observation only if EV(decision) improves more than cost; abstain from calling *any* model when a regex already answers (meta-VOI). Fail-open wake/resume: skip the LLM turn only if the judge answers and p(wake) is low (Horvitz) | `references/mappings.md#6-value-of-information--gather-as-an-enumerated-act` (**Hypothesis** until a labeled act/outcome log; 149-row receipt is Empirical as a shape; wakegate 21/21 is smoke) | | Signal detection / ROC | Criterion and operating point from costs and base rate, not accuracy | `references/mappings.md#7-signal-detection--criterion-not-accuracy` (**Hypothesis** for non-SWE plots) | | Org / safety control structure | Sensor ≠ constraint (Leveson); STPA if the sensor lies | `references/mappings.md#8-control-structure--sensor--constraint-leveson` | @@ -150,8 +150,8 @@ classical method you already trust, substitute it, classify the win | Effect-oriented state-machine loops | Soft predicates on transitions; code owns the transition | `references/mappings.md#19-effect-oriented-state-machine-loops-hypothesis` (**Hypothesis**; ZIO client, not Effect.ts) | | Agent self-supervision / on-track detection | Pre-gate → output judge → done-check → supervisor nouls. S1 reflex keeps control; optional S2 is one-use advice. Claim/evidence Stop (anti-hallucinated-done); bounded Pi supervisor (shadow recovery, never generates commands) | `references/agent-self-assessment.md` | | Optimizer/program frameworks (Ax, DSPy) | Typed fields → one provider request; judge metrics; threshold discipline. Ax and DSPy climb LM-program knobs only | `references/optimizer-integration.md` | -| Perception → decision pipeline / measure / hill-climb | Stages with a versioned state contract; frozen taskset; DSPy/Ax only on the LM-program slice. Specialist composition stays **Hypothesis**; open multimodal decide (blackwood-rlcd) is a named receipt. Structured observe→decide→verified-act (no screenshots) is a computer-use speed-layer receipt (Jev or GLiNER2 backend). Same section as the row below | `references/validation.md#eval--hill-climb` | -| Eval & hill-climb | Decision-stage jevals hygiene; Harbor taskset × harness × runtime; one score-composition table. Shared bake-off exemplar: open-jev-laya-bench (ECE/NLL/Brier; LLM-as-judge is not the score). Harbor-style frozen protocol vs constrained LLMs: DMB (accuracy/calibration/latency/cost; raw logs). Feedstock: jevals-data CC-BY-4.0 boards + JSONL (recompute-from-logs). Collab-arm curriculum: llm_autonomous vs scripted_plus_jev vs llm_plus_jev (Wilson / McNemar). Negative: combinatorial grid assembly ≠ extractive (ARC-AGI Direct Jev 4/400). Harbor on/off routing: chess-engine tasks, hidden perft verifier, one-run preliminary (jev-gateway-bench). Pair CI merge-gate with Harbor + rh-guard. Harbor needle/noise stdout prune: jev-pruner (manual sweep theirs; plugin eval cannot reach Jev → fail-safe original; Terminal-Bench pilot is integration not a full bench) | `references/validation.md#eval--hill-climb` | +| Perception → decision pipeline / measure / hill-climb | Stages with a versioned state contract; frozen taskset; DSPy/Ax only on the LM-program slice. Specialist composition stays **Hypothesis**; open multimodal decide (blackwood-rlcd) is a named receipt. Structured observe→decide→verified-act (no screenshots) is a computer-use speed-layer receipt (Jev or GLiNER2 or Cua-S1 specialist; Cua-S1 source-only, not TypeSafe Jev). Same section as the row below | `references/validation.md#eval--hill-climb` | +| Eval & hill-climb | Decision-stage jevals hygiene; Harbor taskset × harness × runtime; one score-composition table. Shared bake-off exemplar: open-jev-laya-bench (ECE/NLL/Brier; LLM-as-judge is not the score). Harbor-style frozen protocol vs constrained LLMs: DMB (accuracy/calibration/latency/cost; raw logs). Feedstock: jevals-data CC-BY-4.0 boards + JSONL (recompute-from-logs). Collab-arm curriculum: llm_autonomous vs scripted_plus_jev vs llm_plus_jev (Wilson / McNemar). Negative: combinatorial grid assembly ≠ extractive (ARC-AGI Direct Jev 4/400). Harbor on/off routing: chess-engine tasks, hidden perft verifier, one-run preliminary (jev-gateway-bench). Pair CI merge-gate with Harbor + rh-guard. Harbor needle/noise stdout prune: jev-pruner (manual sweep theirs; plugin eval cannot reach Jev → fail-safe original; Terminal-Bench pilot is integration not a full bench). Cua-S1 specialist form: source-only (metric names, no checkpoint scores; not TypeSafe Jev) | `references/validation.md#eval--hill-climb` | | (meta) Finding new mappings & applications | Toolbox sweep: judgment-shaped component of a known method, substituted + falsified | `references/toolbox-mapping.md` | | Named methods / operators / theorems | Substitution tiers: operand-judgments, preconditioned theorems, non-substitutable | `references/methods-catalog.md` | | (meta) Where a judgment model sits relative to any construct | 11 positions + logical-operator rules + position×construct traversal as the application generator | `references/composition-algebra.md` | diff --git a/.agents/skills/augustus/references/agent-self-assessment.md b/.agents/skills/augustus/references/agent-self-assessment.md index 8655a7b..8f12f3c 100644 --- a/.agents/skills/augustus/references/agent-self-assessment.md +++ b/.agents/skills/augustus/references/agent-self-assessment.md @@ -53,6 +53,10 @@ relying: foreman, pi-jev, pi-warden, winnow, fast-jev-compaction, jev-judgment. [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) — local GLiNER2 scores observed controls; `DONE` ≠ verified success (`notes.md` §52). + Specialist-form cousin, **not TypeSafe Jev:** + [Cua-S1](https://github.com/trycua/cua/tree/main/libs/cua-s1) — + option-attention among observed elements; plan ≠ execute; dry-run + default; source-only (`notes.md` §54). Productized Kahneman cascade for *any* cheap-decide / expensive-write loop (business/life, not only SWE): [dual-process-ai](https://github.com/taro1985/dual-process-ai) — diff --git a/.agents/skills/augustus/references/applied-mappings.md b/.agents/skills/augustus/references/applied-mappings.md index 66bfab2..53f53b6 100644 --- a/.agents/skills/augustus/references/applied-mappings.md +++ b/.agents/skills/augustus/references/applied-mappings.md @@ -138,7 +138,11 @@ verified success (`notes.md` §52). Open-head cousin: [laya-mind2web](https://huggingface.co/ShaunSpark/laya-mind2web-browser-agent) — Laya operation + target index over interactive DOM elements (not screenshot multimodal). Contrast blackwood-rlcd (letters on a -screenshot). +screenshot). Specialist-form cousin, **not TypeSafe Jev:** +[Cua-S1](https://github.com/trycua/cua/tree/main/libs/cua-s1) — +option-attention among observed elements (fill/check/click/skip); +code owns execution order; dry-run default; source-only +(`notes.md` §54). **DOM-as-text + fan-out (Empirical as atlas browser-use *shape*):** a screenshot task translated into a structured DOM snapshot as `state`, then speculative questions over numbered candidates — not vision diff --git a/.agents/skills/augustus/references/composition-algebra.md b/.agents/skills/augustus/references/composition-algebra.md index ca76fa9..28dbec3 100644 --- a/.agents/skills/augustus/references/composition-algebra.md +++ b/.agents/skills/augustus/references/composition-algebra.md @@ -146,6 +146,8 @@ Reusable shapes when generating applications: batched question set per step → execute via AX actions. Encoder-backend cousin: gliner2-ultrafast scores observed a11y/DOM controls with local GLiNER2; code clicks; `DONE` ≠ success (`notes.md` §52). + Specialist-form cousin: Cua-S1 option-attention (fill/check/click/skip); + not TypeSafe Jev; plan ≠ execute; source-only (`notes.md` §54). 9. **Shadow-mode harness** (jev-harness): policy + confidence gate + shadow mode + offline eval CLI replaying fixtures, asserting on actions; 24-row filter 48.9 s (Claude CLI) vs 1.3 s Jev at concurrency 8. Compaction rollout: diff --git a/.agents/skills/augustus/references/faq.md b/.agents/skills/augustus/references/faq.md index fd54e74..e6b2961 100644 --- a/.agents/skills/augustus/references/faq.md +++ b/.agents/skills/augustus/references/faq.md @@ -479,9 +479,12 @@ against retrieved session lines, not generated prose (`notes.md` §51). Computer-use encoder cousin: [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) scores observed controls; code clicks; no generated selectors -(`notes.md` §52). Generation is only for +(`notes.md` §52). Specialist-form cousin: +[Cua-S1](https://github.com/trycua/cua/tree/main/libs/cua-s1) selects +fill/check/click/skip among observed elements; does not generate +values or selectors (`notes.md` §54). Generation is only for TYPE/prose when something must be written. `applied-mappings.md` §2; -`notes.md` §48, §50, §52, §53. +`notes.md` §48, §50, §52, §53, §54. ## Should compaction summarize? @@ -539,14 +542,28 @@ and [solari-reflex](https://github.com/hitakshiA/solari-reflex) use TypeSafe Jev; [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) uses local GLiNER2 (`fastino/gliner2-multi-v1`); [laya-mind2web](https://huggingface.co/ShaunSpark/laya-mind2web-browser-agent) -uses a Laya head over DOM element indices. Same lesson as compaction +uses a Laya head over DOM element indices; +[Cua-S1](https://github.com/trycua/cua/tree/main/libs/cua-s1) uses a +byte encoder + option-attention head (fill/check/click/skip) — **not +TypeSafe Jev**, source-only this pass. Same lesson as compaction (Jev Noul/Score vs GLiNER2.5). Screenshot multimodal (blackwood-rlcd: letters on an image) is a **different input**, not a better version of this hole. Hybrid local decide + remote fill is mixed-architecture economics, not dual-process-ai. `DONE` is loop termination, not -verified success. Not GLiNER2.5. Not a bake-off against the Flights -demo clock. `judgment-class.md`; `mixed-architecture.md`; -`notes.md` §52. +verified success. Plan ≠ execute; dry-run default on Cua-S1. Not +GLiNER2.5. Not a bake-off against the Flights demo clock. +`judgment-class.md`; `mixed-architecture.md`; `notes.md` §52, §54. + +## Is Cua-S1 TypeSafe Jev? + +No. [Cua-S1](https://github.com/trycua/cua/tree/main/libs/cua-s1) +uses "System One" as a computer-use research label for a small +specialist decide head. It does not ship Choice/Score/Noul, a +`/v1/systemone` drop-in, or TypeSafe contracts. Augustus places the +*hole* (observe candidates → select among a closed option set → code +acts under a hard envelope), not the logo. Same lesson as GLiNER2 +Ultrafast vs Jev Ultrafast. Source-only; no checkpoint scores. +`judgment-class.md`; `notes.md` §54. ## Is routing the same as memory? diff --git a/.agents/skills/augustus/references/judgment-class.md b/.agents/skills/augustus/references/judgment-class.md index 7e9ea7c..74b1e8c 100644 --- a/.agents/skills/augustus/references/judgment-class.md +++ b/.agents/skills/augustus/references/judgment-class.md @@ -140,6 +140,17 @@ below, next to the when-to-use table. success). Their Flights demo (12.20 s / 13.785 s / ~$0.0001 API) is a demonstration, not a bake-off (`notes.md` §52). Do not copy `uv` / `.env`. + **Specialist computer-use S1 (Empirical as README / MODEL_CARD + behavior, 2026-09-18 ~17:21; weights Watch):** + [Cua-S1](https://github.com/trycua/cua/tree/main/libs/cua-s1) + (`trycua/cua`, MIT source) is a **parallel "System One" name**, not + TypeSafe Jev and not GLiNER. Reference `tinyx`: byte encoder + + option-attention head. Per observed element: fill (from extracted + `Label: value`) / check / click / skip. Does not generate values + or selectors. Plan ≠ execute; dry-run default; `execute` and + `submit` independent opt-ins; fail-closed on unknown checkbox + state. Profile `cua-s1-form-v0` is source-only — no weights, no + checkpoint scores (`notes.md` §54). Do not copy `uv` / MCP. - **Decide.** Typed Choice/Score/Noul with a decision/proper-scoring objective. That is Jev's product claim. Open heads copy the *shape*; distillation copies the *teacher* (openjev-lm, jev-gate-student-b). diff --git a/.agents/skills/augustus/references/mappings.md b/.agents/skills/augustus/references/mappings.md index f8eead2..d21c625 100644 --- a/.agents/skills/augustus/references/mappings.md +++ b/.agents/skills/augustus/references/mappings.md @@ -441,6 +441,10 @@ the same hole. Code owns actuators, dates, freshness. `DONE` is not the probe — application verifiers are. Contrast blackwood-rlcd (screenshot input). Hybrid remote TYPE is generation, not the classifier step (`notes.md` §52). +[Cua-S1](https://github.com/trycua/cua/tree/main/libs/cua-s1) is the +same substituted-classifier *job* on a specialist form contract +(option-attention; plan ≠ execute; not TypeSafe Jev; source-only, +`notes.md` §54). **Structure induction over a bag (Empirical as a *shape*, 2026-09-18):** [`Joymfl/dag-jev`](https://github.com/Joymfl/dag-jev) — unordered items @@ -568,6 +572,16 @@ GLiNER2 may only pick among candidates the snapshot already holds. `DONE` is not the monitor. Soft judgment inside a hard envelope, encoder backend — not a screenshot VLM (`notes.md` §52). +**Named specialist-form envelope (Empirical as README / MODEL_CARD, +2026-09-18 ~17:21; weights Watch):** +[Cua-S1](https://github.com/trycua/cua/tree/main/libs/cua-s1) +— the monitor is **code** (plan ≠ execute; dry-run default; one +window; snapshot-bound tokens; reobserve; `execute`/`submit` +opt-ins; fail-closed unknown checkbox; fill execution fails closed +without advertised token `set_value`). The option-attention head may +only pick among observed elements and extracted `Label: value` +entities. Not TypeSafe Jev. No checkpoint scores (`notes.md` §54). + ## 13. DST multiverse triage (Hypothesis) **Method**: Antithesis / Resonate DST artifacts → failure taxonomy → diff --git a/.agents/skills/augustus/references/mental-models.md b/.agents/skills/augustus/references/mental-models.md index 469283d..d0d847f 100644 --- a/.agents/skills/augustus/references/mental-models.md +++ b/.agents/skills/augustus/references/mental-models.md @@ -470,13 +470,13 @@ Use these as *existence proofs of a position*. Write your own card. | Agent context | compact completed tool results without inventing prose | retention Choice + char-offset locate (**Empirical**: gliner25-compaction; same *job* as fast-jev-compaction / pi-jev-compaction) | mutation/shell envelope → keep_full; fail-closed keep_full; shadowMode before replace; copy exact bytes | | Agent context | prune Bash stdout before the LLM without inventing prose | Noul per chunk after a hard size/format envelope (**Empirical**: jev-pruner) | ≤10k / JSON-diff-whole-doc untouched; fail-safe original; archive dropped spans | | Dataframe labeling | classify / score rows | Noul/Choice/Score + full `p__` (**Empirical** as jevframe / jevpandas *shape*) | pandas/Polars, thresholds in code | -| Computer-use speed | one verified act per step | score / Choice among numbered a11y/DOM controls (**Empirical**: solari-reflex Jev; gliner2-ultrafast GLiNER2; laya-mind2web Laya DOM indices) | Guard check; deny-list absence; no screenshots; `DONE` ≠ success | +| Computer-use speed | one verified act per step | score / Choice among numbered a11y/DOM controls (**Empirical**: solari-reflex Jev; gliner2-ultrafast GLiNER2; laya-mind2web Laya DOM indices; cua-s1 option-attention, source-only, not TypeSafe Jev) | Guard check; deny-list absence; no screenshots; `DONE` ≠ success; plan ≠ execute | | Agent turn | skip memory tour on easy intent | intent Choice (**Empirical**: jev-hermes) | Memory still writes; complex still searches | | Document / lab routing | which pages need the expensive observation | Noul on remainder after a text layer / recipe | local extract, merge order (**Empirical** as OCR-router *shape*) | | Shell / tool allowlist | unlisted remainder after a **proof** | five Nouls on unknown verbs | Proven/Refused in code; cannot block (**Empirical**: jevgate) | | SWE | residual AGENTS.md / CLAUDE.md rules | one Score per named instruction-file rule | linter owns hard rules; bands + fail-open (**Empirical**: Abide replay, `notes.md` §47) | | Screenshot candidates → act | lettered elements code already marked | Choice over those letters | Click in code (**Empirical** as blackwood-rlcd *shape*; CC BY-NC) | -| Browser / DOM candidates → act | numbered elements from a **text** snapshot | score among those ids (**Empirical**: atlas browser-use / jev-ultrafast / gliner2-ultrafast *shape*: DOM-as-text, not vision) | Click in code; no screenshots; hybrid remote TYPE optional | +| Browser / DOM candidates → act | numbered elements from a **text** snapshot | score among those ids (**Empirical**: atlas browser-use / jev-ultrafast / gliner2-ultrafast *shape*: DOM-as-text, not vision; cua-s1 specialist form, source-only) | Click in code; no screenshots; hybrid remote TYPE optional; plan ≠ execute | | Knowledge / recall | fact that is not in the document | **Do not ask.** Retrieve the passage first; then a self-contained Choice (**Empirical**: history suite A wrong@0.90 → C right@0.97) | Index, citation, the passage in `state` | | Dual-process cascade | cheap classify / route vs write | S1 typed decision + τ; S2 generates only on low conf (**Empirical as a productized metaphor**; routing accuracy **unmeasured** — dual-process-ai) | Safety still fail-closed in code | | CI merge-gate | ignore infra noise without merging a real bug | cause Choice per cluster (**Empirical**: latch demo PASS vs BLOCK) | Cluster + fingerprint + `--gate` table; reporter never fails the runner | diff --git a/.agents/skills/augustus/references/methods-catalog.md b/.agents/skills/augustus/references/methods-catalog.md index 1dc59e7..7330fe4 100644 --- a/.agents/skills/augustus/references/methods-catalog.md +++ b/.agents/skills/augustus/references/methods-catalog.md @@ -46,7 +46,7 @@ judgment component is new). | Beam search over taxonomies | Which branches deserve expansion | Choice distributions as branch priority; keep K paths where ambiguity is early | Frontier, budget, final selection | **Empirical recipe** (beam K=3 cookbook) | | Structure induction over a bag | Pairwise "does i depend on j?" (or Choice over order) | One judgment per pair; DAG / scheduler in code | Topology, cycles, execution | **Empirical as a shape** (dag-jev experiment; empty README; no metrics, `notes.md` §48) | | Collab-arm product loop | Scripted legal set vs LLM-propose vs unconstrained | Choice over legal actions; stop on low p rather than guess | Legality, Wilson/McNemar, ceiling flags | **Empirical as a harness shape** (jev-testbench; bake into jevals/Harbor, `notes.md` §48) | -| Computer-use observe → score → act | Which observed control matches the current requirement | Score / Choice among a11y/DOM candidates (Jev *or* GLiNER2 *or* Laya); code clicks | Observation, freshness, dates; never generate selectors; independent outcome check (`DONE` ≠ success) | **Empirical recipe** (jev-ultrafast / solari-reflex Jev; gliner2-ultrafast GLiNER2; laya-mind2web DOM-index Laya; contrast blackwood-rlcd screenshot, `notes.md` §48, §52) | +| Computer-use observe → score → act | Which observed control matches the current requirement | Score / Choice / option-attention among a11y/DOM candidates (Jev *or* GLiNER2 *or* Laya *or* Cua-S1); code clicks | Observation, freshness, dates; never generate selectors or values; independent outcome check (`DONE` ≠ success); plan ≠ execute | **Empirical recipe** as shipped loops (jev-ultrafast / solari-reflex Jev; gliner2-ultrafast GLiNER2; laya-mind2web DOM-index Laya); **Empirical as README/MODEL_CARD** for Cua-S1 (source-only, not TypeSafe Jev, `notes.md` §54); contrast blackwood-rlcd screenshot, `notes.md` §48, §52 | | Screening / Wald sequential tests | Pass / fail / keep-looking per candidate | One Noul gate per candidate in one batched request; budget in code | Sequential rule, stop boundaries | **Hypothesis** | | STPA / STAMP control structure | Sensor reading vs enforced constraint | Judgment as sensor; constraints in policy/code/interlock; STPA table if the sensor lies | The constraint, the actuator, the probe | **Contract** as ownership; **Hypothesis** as domain product (`mappings.md` §8) | | PufferLib / Ocean env contracts | Does this episode look like a known trainer-bug mode? | Cluster failing episodes; never "the policy is correct" | Seeded serial env, Ocean sanity, observed rewards | **Hypothesis** as placement; **Contract** that Ocean is not a comparative baseline (`formal-methods.md` DST trio) | diff --git a/.agents/skills/augustus/references/mixed-architecture.md b/.agents/skills/augustus/references/mixed-architecture.md index 4a42af4..7ee9d4c 100644 --- a/.agents/skills/augustus/references/mixed-architecture.md +++ b/.agents/skills/augustus/references/mixed-architecture.md @@ -189,7 +189,7 @@ not a global virtue: | Merge a red CI run | **Fail closed** on `--gate` (`latch`); reporter stays fail-open | False PASS merges a real bug. Missing key never fails Playwright; the gate is a separate step. Judge never says ignore alone | | Plain-English PR check | **Fail closed** on error / empty / low confidence (`if-ai`) | A skipped or timed-out check is not a pass. Threshold is policy, not measured correctness | | Route to a tool / start a side effect | **Fail closed** (don't call) | A wrong tool is an action | -| Actuate an observed browser control | **Fail closed** (code validates the node) | Freshness / visibility / disabled / occlusion in code; model never emits selectors (`gliner2-ultrafast`, jev-ultrafast, solari-reflex). `DONE` does not authorize "success" | +| Actuate an observed browser control | **Fail closed** (code validates the node) | Freshness / visibility / disabled / occlusion in code; model never emits selectors (`gliner2-ultrafast`, jev-ultrafast, solari-reflex). `DONE` does not authorize "success". Cua-S1: dry-run default; `execute`/`submit` opt-in; fail-closed unknown checkbox; fill execution fails closed without token `set_value` | | Rerank a retrieved list | Fail open: keep retrieval order (`WiktorB2004/llama-index-jev`, **Empirical recipe** on BEIR nfcorpus: MiniLM 0.340 nDCG@5 → MiniLM+Jev 0.396; rerank fails open, *select* fails closed). Listwise/cross-encoder scores belong here, not on the row above. | Ranking errors are quality; selection errors are control-flow | Worked placements (2026-09-18 topic:jev hour + prior archive): @@ -460,7 +460,8 @@ decision-design card. Do not clone APIs from READMEs. | Decision-as-business-tool | Named judgment; gate is part of the result | Registry, arithmetic, hard guards | jev-decision-layer (unofficial) | | NL cases → checked e2e | Jev selects observed controls | Playwright expectations; PASS/FAIL/BLOCKED | jev-e2e (alpha) | | Extractive quotes / pointer evidence | Per-sentence, per-line-id, or char-offset Noul/Choice | Verbatim join; place; `redecide` / CSV; model never writes the excerpt | testimonial-miner; jev-reviewer; gliner25-compaction | -| Structured observe → decide → act | Score / Choice among numbered a11y/DOM controls | Guard check; deny-list absence; no screenshots; no generated selectors; TYPE is the only generation; `DONE` ≠ verified success | solari-reflex (Jev); jev-ultrafast (Jev); gliner2-ultrafast (GLiNER2); laya-mind2web (Laya, DOM indices) | +| Structured observe → decide → act | Score / Choice among numbered a11y/DOM controls | Guard check; deny-list absence; no screenshots; no generated selectors; TYPE is the only generation; `DONE` ≠ verified success | solari-reflex (Jev); jev-ultrafast (Jev); gliner2-ultrafast (GLiNER2); laya-mind2web (Laya, DOM indices); cua-s1 (option-attention fill/check/click/skip; not TypeSafe Jev; source-only) | +| Specialist form S1 (plan ≠ execute) | Option-attention among observed elements | Dry-run default; snapshot-bound tokens; reobserve; submit opt-in; fail-closed checkbox/fill | cua-s1 (`cua-s1-form-v0` profile; no weights this pass) | | Hybrid local decide + remote fill | Local encoder scores observed controls | Code owns actuators; remote OpenAI-compat helper writes field text only | gliner2-ultrafast (GLiNER2 local + Mercury 2.5 default) | | Dataframe semantic columns | Noul / Choice / Score per row; full `p__` | pandas/Polars, indexes, never silent renormalize | jevpandas; jevframe (PyPI + Polars) | | Route ≠ memory | Intent Choice before a turn | Config + flat tools on easy routes; memory stays on for hard ones | jev-hermes | diff --git a/.agents/skills/augustus/references/toolbox-mapping.md b/.agents/skills/augustus/references/toolbox-mapping.md index 789cc0f..952f000 100644 --- a/.agents/skills/augustus/references/toolbox-mapping.md +++ b/.agents/skills/augustus/references/toolbox-mapping.md @@ -78,7 +78,7 @@ component; keep the rest of the method in code. | Discrete math: width vs depth | Fan out in width (parallel ≈ free), pay depth linearly; two-stage only when next options depend on an earlier answer | **Empirical recipe** (fan-out: 12.2× cheaper, 10× faster) | | Psychology: Kahneman | System 2 generates/proposes (LLM), System 1 discriminates (Jev); never the reverse. S1 keeps control; optional S2 is one-use advice and does not fly. Productized cascade: `conf ≥ τ` → S1 decides else S2 writes; routing fails open / safety fails closed; **routing accuracy unmeasured**; keyword fallback ≠ S1. S1 specialists + S2 coordinator is the same split (reification-labs/foreman is description-only Phoenix scaffold this pass — do not invent an Elixir API). Indexer: S1 GLiNER extract on the bulk, escalate LLM on the tail (10–50× unfilled) | **Empirical recipe** (mcts-agent role split; 2026-09-18 mixed-architecture discourse; jev-reflex-autonomy-lab as a control-loop shape, `notes.md` §46; [dual-process-ai](https://github.com/taro1985/dual-process-ai) as a business/life cascade, `notes.md` §49; s1-graphify-indexer, `notes.md` §51). **Route ≠ memory:** a cheap intent gate skips memory/tool *tours* on easy routes; memory still writes; complex still searches (jev-hermes, `notes.md` §48) | | IR / cascades | Cheap relevance / irrelevance before an expensive ranker or generator | **Empirical recipe** (RAG cookbook; jevprune; git-jev-stage; LlamaIndex Jev rerank; jev-pruner stdout after size/format envelope, `notes.md` §53) | -| IR / extractive keep-drop | Number candidates in code; model selects; copy verbatim; `redecide` thresholds on the log with no new calls. Compaction: same pointer job on tool results (Jev Noul/Score *or* GLiNER encoder). Computer-use: same pointer job on observed a11y/DOM controls (Jev *or* GLiNER2). Stdout: same pointer job on Bash chunks after a hard envelope (Jev Noul) | **Empirical recipe** (testimonial-miner; jev-reviewer pointer-not-generator, `notes.md` §48; gliner25-compaction char-offset + fail-closed keep_full, `notes.md` §50; gliner2-ultrafast observe→score→act, `notes.md` §52; jev-pruner, `notes.md` §53) | +| IR / extractive keep-drop | Number candidates in code; model selects; copy verbatim; `redecide` thresholds on the log with no new calls. Compaction: same pointer job on tool results (Jev Noul/Score *or* GLiNER encoder). Computer-use: same pointer job on observed a11y/DOM controls (Jev *or* GLiNER2 *or* Cua-S1 option-attention). Stdout: same pointer job on Bash chunks after a hard envelope (Jev Noul) | **Empirical recipe** (testimonial-miner; jev-reviewer pointer-not-generator, `notes.md` §48; gliner25-compaction char-offset + fail-closed keep_full, `notes.md` §50; gliner2-ultrafast observe→score→act, `notes.md` §52; jev-pruner, `notes.md` §53; cua-s1 specialist form, source-only, `notes.md` §54) | | Spec / lint | Project-defined semantic rules as predicates over a diff; linter owns hard rules. AST remainder: Tree-sitter units, then typed questions; do not execute scanned code. Plain-English PR check: one condition + min-confidence; fail-closed on error | **Empirical recipe** (jev-pref contract; Abide productized path — replay 93 sessions, edit precision ~26% / turn ~73% before tune, `notes.md` §47; JevLint file-level Noul; pi-warden; snifftest unsure-band; jevscan AST∩semantic, `tenbin` owns the lint skill, `notes.md` §48; if-ai, `notes.md` §51). jev-marshal is Watch / empty this pass | | Formal methods / DST / safety | Judgment triages counterexamples, failing seeds, and named-rule conformance; proof/MC/DST stay with their tools. Alloy finder ≠ Apalache BMC ≠ Quint run. DST trio: Antithesis hypervisor / Resonate HQ Lean+oracle+SDK (durable async) / PufferLib env+seed. Noul is a sensor, not a discharged PO. Semi-formal diagrams are vocabularies, not enforcers | **Hypothesis as product**, **Contract** as ownership (matching `mappings.md` §8 and `methods-catalog.md`; worked shape pi-warden — `formal-methods.md`, `formal-semi-formal.md`) | | Decision analysis: VOI | Gather as an enumerated act; pay iff expected decision-loss drop > cost. Wake/resume: skip the LLM turn only if the judge answers and p is low (Horvitz); user-message / skip-limit already answer without a model | **Hypothesis** as calculator (`mappings.md` §6; `mental-models.md`). wakegate 21/21 is smoke (`notes.md` §51) | diff --git a/.agents/skills/augustus/references/validation.md b/.agents/skills/augustus/references/validation.md index d550855..7329623 100644 --- a/.agents/skills/augustus/references/validation.md +++ b/.agents/skills/augustus/references/validation.md @@ -324,7 +324,7 @@ Rules: | LM-program knobs only | DSPy/Ax (narrow) | never primary System One calibration score | | Reward-hack / eval gaming | [rh-guard](https://github.com/24601/rh-guard) | structural deny + System One sidecar | | Project soft-rule lint | [Abide](https://github.com/coldteadotai/abide) | Score per rule on the diff; bands; fail-open; replay + independent review | -| Collab / computer-use product loop | [jev-testbench](https://github.com/ufx7/jev-testbench); [solari-reflex](https://github.com/hitakshiA/solari-reflex); [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast) | Wilson/McNemar arms; independently checked task time; `DONE` ≠ success | +| Collab / computer-use product loop | [jev-testbench](https://github.com/ufx7/jev-testbench); [solari-reflex](https://github.com/hitakshiA/solari-reflex); [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast); [cua-s1](https://github.com/trycua/cua/tree/main/libs/cua-s1) | Wilson/McNemar arms; independently checked task time; `DONE` ≠ success; Cua-S1 source-only (metric names, no checkpoint scores) | | Agent routing on vs off | [jev-gateway-bench](https://github.com/vinilana/jev-gateway-bench) | Hidden perft; cost/quality; one-run signal this pass | | Command-output prune (needle/noise) | [jev-pruner](https://github.com/tamaratran/jev-pruner) | Manual `trimOutput` sweep (theirs); plugin eval cannot reach Jev (fail-safe original); Terminal-Bench paired pilot is integration, not a full bench | @@ -349,7 +349,14 @@ vs 98.4 s (`notes.md` §48). Encoder-backend cousin: — same hole, local GLiNER2; their Flights demo (12.20 s visible / 13.785 s loop / ~$0.0001 API) is a **demonstration**, not a bake-off or a vs-Jev-Ultrafast table; `DONE` is not the Harbor score -(`notes.md` §52). **Collab-arm curriculum:** +(`notes.md` §52). Specialist-form cousin, **not TypeSafe Jev:** +[Cua-S1](https://github.com/trycua/cua/tree/main/libs/cua-s1) — +option-attention among observed elements; plan ≠ execute; dry-run +default; source-only this pass. Offline utilities *name* accuracy, +abstention, coverage, wrong actions/targets, and unsafe-when-should- +abstain; **no checkpoint scores**. Tests exercise implementation, not +quality. Do not invent a vs-Jev table. Watch for `cua-s1-form-v0` +(`notes.md` §54). **Collab-arm curriculum:** [jev-testbench](https://github.com/ufx7/jev-testbench) — `llm_autonomous` vs `scripted_plus_jev` vs `llm_plus_jev`; Wilson + McNemar; Jev is not a peer arm. Bake into jevals/Harbor hygiene, do diff --git a/CHANGELOG.md b/CHANGELOG.md index 4c4c16a..ed3294e 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -252,6 +252,22 @@ Pegged against [`typesafe-ai/skills` v0.5.7](https://github.com/typesafe-ai/skil Manual sweep (theirs): needles 24/24; mean reduction 83% on trim scenarios. Harbor plugin-eval cannot reach Jev. Terminal-Bench paired pilot is integration, not a full bench. No invented metrics. +- Cua-S1 specialist System One computer-use (`research/notes.md` §54, + [cua-s1](https://github.com/trycua/cua/tree/main/libs/cua-s1), + parent MIT, ~23.3k★ this pass): architecture notes, not a Driver / + MCP / `uv` how-to. Form-oriented profile `cua-s1-form-v0`. Byte + encoder + option-attention head chooses fill/check/click/skip per + observed element; does not generate values or selectors. Plan ≠ + execute; dry-run default; `execute`/`submit` independent opt-ins; + fail-closed on unknown checkbox / fill without advertised token + `set_value`. **Not TypeSafe Jev** — parallel "System One" naming in + CUA research. Same observe→score-among-candidates→code-acts *job* + as jev-ultrafast / gliner2-ultrafast / solari-reflex / laya-mind2web; + specialist form contract, source-only this pass (no weights, no + checkpoint scores). Offline metric *names* only (accuracy, + abstention, coverage, wrong actions/targets, unsafe when should + abstain). Tests exercise implementation, not checkpoint quality. + Watch for a `cua-s1-form-v0` artifact drop. No invented metrics. - Effect-oriented loops (`notes.md` §28, `mappings.md` §19): Ward's ZIO client keeps Jev as the outer Choice and the handler as the effect. Not Effect.ts. GLiNER author: GLiNER2 "like jev" is GLiGuard diff --git a/README.md b/README.md index 3beaaa6..1fcbcfc 100644 --- a/README.md +++ b/README.md @@ -57,12 +57,13 @@ never launder a Noul as a proof. generate; component node; DOM-as-text + fan-out; shadow-mode compaction rollout; fail-open wake vs fail-closed merge-gate; Harbor on/off routing; hybrid local decide + remote fill; `DONE` ≠ verified success; - evidence-preserving stdout prune (hard envelope then Noul) + evidence-preserving stdout prune (hard envelope then Noul); + specialist S1 computer-use (Cua-S1 form-v0; plan ≠ execute; not TypeSafe Jev) - `.agents/skills/augustus/references/applied-mappings.md` — context sieve, exact-text keep/drop (extractive / pointer-not-generator; char-offset compaction; observed a11y/DOM controls; Bash stdout prune), env triage (OpenSmoke + latch merge-gate), moderation/ranking, skill routing (route ≠ memory) - `.agents/skills/augustus/references/faq.md` — "just classification", stack replacement, Jev vs open head vs encoder vs LoRA vs constrained AR vs kev vs blackwood, - wait-for-Archer, missing-other confident-wrong, soft project rules vs linter (Abide), extractive/pointer-not-generator, compaction summarize vs pointer, encoder vs Jev compaction, fail-closed keep_full, shadow-mode rollout, fail-open vs fail-closed wake vs CI gate, observe→score→act backend-agnostic, hybrid local decide + remote fill, DONE ≠ verified success, stdout prune vs session compaction, local drop-in vs stub scorer, route ≠ memory, when-it-holds / extractable-from-state, decision-model vs constrained LLM, dual-process S1/S2, combinatorial grid ≠ extractive, GLiNER vs GLiClass vs CLIP, LLM-as-judge, in-engine vs CLI store, + wait-for-Archer, missing-other confident-wrong, soft project rules vs linter (Abide), extractive/pointer-not-generator, compaction summarize vs pointer, encoder vs Jev compaction, fail-closed keep_full, shadow-mode rollout, fail-open vs fail-closed wake vs CI gate, observe→score→act backend-agnostic, hybrid local decide + remote fill, DONE ≠ verified success, stdout prune vs session compaction, Cua-S1 vs TypeSafe Jev, plan ≠ execute / dry-run, local drop-in vs stub scorer, route ≠ memory, when-it-holds / extractable-from-state, decision-model vs constrained LLM, dual-process S1/S2, combinatorial grid ≠ extractive, GLiNER vs GLiClass vs CLIP, LLM-as-judge, in-engine vs CLI store, hard envelope (bitrate / planner), not-another-how-to - `.agents/skills/augustus/references/mappings.md` — classical-method mappings with boundaries, counterexamples, acceptance tests (including @@ -75,7 +76,8 @@ never launder a Noul as a proof. CC-BY-4.0 recompute-from-logs feedstock; Abide replay as Harbor-adjacent soft-rule measurement; solari-reflex Harbor-style computer-use; gliner2-ultrafast encoder-backend cousin (`DONE` ≠ - success; demo is not a bake-off); jev-testbench collab arms; ARC-AGI Direct Jev as + success; demo is not a bake-off); Cua-S1 specialist form source-only + (metric names, no checkpoint scores; not TypeSafe Jev); jev-testbench collab arms; ARC-AGI Direct Jev as combinatorial-≠-extractive negative; jev-gateway-bench Harbor on/off routing one-run signal; jev-pruner Harbor needle/noise + Terminal-Bench integration pilot, not a full bench) diff --git a/docs/ecosystem.md b/docs/ecosystem.md index c5249fa..446d8cb 100644 --- a/docs/ecosystem.md +++ b/docs/ecosystem.md @@ -127,7 +127,7 @@ the READMEs, not a monopoly. - **AppitStudio/testimonial-miner** — extractive selection + multi-question broadcast + offline `redecide`. Model never writes the quote. - **choxos/jev-reviewer** — pointer-not-generator: line ids; verbatim copy with place; *not found* is an answer. - **us/jev-local** — contract-compatible `POST /v1/systemone`. Default scorer is a **stub** until `JEVLOCAL_SCORER=hf`. -- **hitakshiA/solari-reflex** — observe → decide → verified act; no screenshots. Author table vs Codex on Solari ~3–7× wall. Encoder-backend cousin: gliner2-ultrafast (`notes.md` §52). +- **hitakshiA/solari-reflex** — observe → decide → verified act; no screenshots. Author table vs Codex on Solari ~3–7× wall. Encoder-backend cousin: gliner2-ultrafast (`notes.md` §52). Specialist-form cousin: cua-s1 (`notes.md` §54). - **ktaletsk/jevframe** — pandas/Polars `.jev` accessor; full `p__`; sibling of jevpandas. - **de-niji/jev-hermes** — route ≠ memory: cheap intent gate skips memory tours. - **ngallodev-software/agent-workflow-typesafe-ai** — advisory sidecar receipts; never changes host routing (Apache-2.0). @@ -185,6 +185,12 @@ Architecture notes, not a plugin catalog. `notes.md` §53. TypeSafe Jev is the e - **tamaratran/jev-pruner** — after Bash, Jev Noul-prunes stdout chunks before the main LLM sees them. No summary. Hard envelope (≤10k estimated tokens; JSON/XML/YAML/diff/binary; whole-document commands) then soft Noul. Fail-safe keep original; full archive. Marketplace id still `fast-jev-output`. Codex is opt-in wrapper, not automatic interception. Same evidence-preserving *family* as fast-jev-compaction / gliner25-compaction; different *job* (command output vs session memory). Manual sweep (theirs): needles 24/24; mean reduction 83% on trim scenarios. Harbor plugin-eval cannot reach Jev. Terminal-Bench paired pilot is integration, not a full bench. MIT. +### Hourly ~17:21 Boise (Cua-S1 specialist form System One, source-only) + +Architecture notes, not a Driver / MCP catalog. `notes.md` §54. TypeSafe Jev is the exemplar, not the monopoly. **Not TypeSafe Jev. Not GLiNER. Not a general CUA. Not multimodal pixels-in.** Archer still Watch. Weights Watch. + +- **trycua/cua `libs/cua-s1`** — specialist System One computer-use research. Profile `cua-s1-form-v0` (form-oriented UI). Parent MIT; ~23.3k★ this pass. Byte-level `tinyx` encoder + option-attention head: per observed element fill (from extracted `Label: value`) / check / click / skip. Does not generate values or selectors. Code owns execution order. Plan ≠ execute; dry-run default; `execute` and `submit` independent opt-ins; submit at most one high-confidence Button labeled exactly `Submit` / `Submit Form`; fail-closed on missing checkbox state; fill execution fails closed unless the runtime advertises token-based `set_value`. Source-only: no weights, no checkpoint scores. Offline metric *names* only (accuracy, abstention, coverage, wrong actions/targets, unsafe when should abstain). Tests exercise implementation, not checkpoint quality. Same observe→score-among-candidates→code-acts *job* as jev-ultrafast / gliner2-ultrafast / solari-reflex / laya-mind2web; parallel "System One" name in CUA, not a TypeSafe contract. Watch for a `cua-s1-form-v0` artifact drop. + See `references/mixed-architecture.md` in the skill. Class-level family choice: `references/judgment-class.md`. Proof vs judgment (Alloy vs Apalache; DST trio Antithesis / Resonate HQ / PufferLib): diff --git a/research/archive/findings.md b/research/archive/findings.md index 4235b3a..e15d84f 100644 --- a/research/archive/findings.md +++ b/research/archive/findings.md @@ -1031,6 +1031,43 @@ pass-through (≤10k / JSON-diff) is the sandwich, Jev is the remainder; (ck) Harbor plugin-eval refusing Jev is a fail-safe receipt, not a missing metric; (cl) marketplace id may lag the repo name. +## Batch #38 (2026-09-18 ~17:21 Boise) — Cua-S1 specialist form System One (source-only) + +Note: `research/notes.md` §54. Docs-only. Folded into PR #2. Not a +second PR. Archer still Watch. **Not TypeSafe Jev. Not GLiNER. Not a +general CUA. Not multimodal pixels-in.** No invented metrics. Do not +re-fold §52 gliner2-ultrafast as this product, §48 solari-reflex, +§53 jev-pruner, blackwood-rlcd as a screenshot cousin, or +laya-mind2web as a Laya DOM-index cousin. + +- **trycua/cua `libs/cua-s1` (Empirical as README / MODEL_CARD + behavior; weights Watch).** Parent MIT; ~23.3k★ this pass + (updated 2026-09-18T23:21:08Z). Research family of small specialist + computer-use models. Profile `cua-s1-form-v0`. Source-only: no + weights, datasets, or checkpoint scores. `tinyx` byte encoder + + option-attention: per observed element fill / check / click / + skip. Fill values selected from extracted `Label: value` pairs, + not generated. Code owns execution order. Plan ≠ execute; dry-run + default; `execute`/`submit` opt-in; fail-closed unknown checkbox / + fill without advertised token `set_value`. Tests exercise + implementation, not checkpoint quality. Offline metric *names* + only. +- **Mental models:** (1) specialist S1 vs general agent (narrow task + contract; membership ≠ general CUA); (2) Choice among observed + elements / fixed actions (same hole as jev-ultrafast / + gliner2-ultrafast / solari / laya-mind2web; contrast blackwood + screenshot); (3) plan ≠ execute, dry-run default, fail-closed + envelope; (4) parallel "System One" naming, not TypeSafe Jev. +- **rh-guard:** light note only (dry-run / submit opt-in / + fail-closed state). Not reward-hack detection. + +Cross-repo addition: (cm) observe→score-among-candidates→code-acts +is backend-agnostic including a specialist CUA head; (cn) "System +One" in CUA research is a parallel name, not a TypeSafe contract; +(co) plan ≠ execute / dry-run is the sandwich around a soft +specialist; (cp) source-only drops publish metric *names*, not +checkpoint scores — Watch for `cua-s1-form-v0`. + diff --git a/research/notes.md b/research/notes.md index d7c3095..7e944e5 100644 --- a/research/notes.md +++ b/research/notes.md @@ -4015,3 +4015,146 @@ be installed together. Marketplace / Claude plugin id is still Not multimodal. Usage: command-output sieve as a perception/memory hygiene stage **before** the generative turn. Harbor-adjacent eval harness in-repo. Archive + landscape pointer. + +## 54. Cua-S1 — specialist System One computer-use (form-v0 profile, source-only) (2026-09-18 ~17:21 Boise) + +America/Boise ~17:21 = 23:21 UTC. Docs-only fold into open PR #2 +(`cursor/augustus-store-envelope-00b4`). Not a competing PR. Archer +27B drop still **WATCH**. Identity lock vs `typesafe-ai` / `tenbin` / +`decision-first` holds. No wrapper, no `uv` / MCP / Driver how-to, no +copied factory env vars. No invented metrics. **Not TypeSafe Jev. +Not GLiNER. Not a general CUA. Not multimodal pixels-in. Source-only +— no weights, no checkpoint scores.** Do not re-fold §52 +gliner2-ultrafast as this product, §48 solari-reflex, §53 +jev-pruner, blackwood-rlcd as a screenshot cousin, or laya-mind2web +as a Laya DOM-index cousin. + +Family: **observe → score-among-candidates → code acts** (selection +head over observed elements; code owns execution order). Same hole +as jev-ultrafast / solari-reflex (Jev), gliner2-ultrafast (GLiNER2), +laya-mind2web (Laya). **Different naming:** CUA's "System One" is a +parallel research label, not a TypeSafe Jev contract. **Different +scope:** specialist form-oriented checkpoint profile +(`cua-s1-form-v0`), not a general computer-use agent. Weights +**Watch**. + +### HIGH + +1. **[`trycua/cua` `libs/cua-s1`](https://github.com/trycua/cua/tree/main/libs/cua-s1)** + (parent MIT; Python package `cua-s1` / `cua_s1`; ~23.3k★ parent + this pass). Research project for **small, specialist computer-use + models** with a defined task class. First checkpoint profile: + `cua-s1-form-v0` (form-oriented UI). This component ships model, + synth data, training, eval utilities, and optional Cua Driver + + MCP server **source**. It does **not** include or download + weights, datasets, demo binaries, or recordings. **No checkpoint + performance claim.** Future official weights may use separate + terms. **Not a TypeSafe Jev drop-in. Not a screenshot VLM. Not a + generator of selectors or field values.** + + Load-bearing loop (README + MODEL_CARD): + + ```text + snapshot / a11y elements + Label:value entities + → tinyx byte encoder + option-attention + → per element: fill | check | click | skip + → code orders execution (plan ≠ execute; dry-run default) + ``` + + Reference `tinyx`: byte-level transformer encoder + + **option-attention classification head**. Per observed interface + element, one option from a **fixed set**: fill with an entity + extracted from the source document, check, click, or skip. The + prototype scores elements independently. Document parser only + extracts `Label: value` pairs. **Code** turns selected options + into an execution order. Fill values are **selected**, not + generated. + + **Four load-bearing mental models (architecture, not a Driver + how-to):** + + 1. **Specialist S1 vs general agent.** Membership in the Cua-S1 + family does not imply general computer-use capability. A + checkpoint has a narrow task contract and checkpoint-specific + eval. Same philosophy as "Jev-class for a job," not omnimodal + AGI. Do not treat `form-v0` as evidence outside its evaluated + boundaries — and there is **no evaluated checkpoint** in this + source-only drop. + + 2. **Choice among observed elements / fixed actions.** Selection + head, not a generator of selectors or values. Family with + gliner2-ultrafast (score observed controls), solari-reflex + (structured observe → typed act), jev-ultrafast, laya-mind2web + (DOM indices). Contrast blackwood-rlcd (screenshot + marked + letters). If the fill entity is not already a `Label: value` + pair the parser holds, this card does not apply. + + 3. **Plan ≠ execute; dry-run default; fail-closed.** Planning and + execution are separate. Optional runtime defaults to dry run. + One unambiguous target window; snapshot-bound element tokens; + reobserve after each mutation. `execute` and `submit` are + independent opt-ins. Submit is narrow: at most one + high-confidence Button / AXButton whose normalized label is + exactly `Submit` or `Submit Form`. Fail-closed on missing + checkbox role/checked state; already-checked boxes skipped; + checked postcondition verified. PDF confined to allowed roots + (cwd default; production should use a dedicated directory). + Portable Cua Driver contract does not currently expose + `set_value` — fill **execution** fails closed unless the + connected runtime advertises token-based value mutation; + planning remains available. Inspect the dry-run plan before + enabling both execution flags. + + 4. **Not TypeSafe Jev.** Parallel "System One" naming in + computer-use research. No Choice/Score/Noul contract, no + `/v1/systemone` drop-in. Augustus stays family-first: the + *hole* is specialist decide among observed candidates under a + hard envelope. Backend-agnostic judgment class still applies. + + **Eval honesty (theirs; no scores this pass).** Included tests + exercise **implementation behavior, not checkpoint quality**. + Offline utilities *report* abstention, coverage, selective + accuracy, wrong actions, wrong targets, and unsafe actions when + the expected behavior was to abstain. Synthetic splits are + disjoint by form signature; model selection uses validation + rather than test. A future checkpoint **must** report exact + revisions, task set, environment, action space, independent + outcome verification, and failure categories. Responsible-use + text: do not treat model output or apparent task completion as + proof the action was correct. **Watch** for a `cua-s1-form-v0` + artifact drop. Do not invent metrics. + + **Siblings — complementary, do not merge.** + + - **gliner2-ultrafast / solari-reflex / jev-ultrafast / + laya-mind2web:** same observe→act *job*; Jev, GLiNER2, or Laya + backends with shipped loops. This is a source-only specialist + head. + - **blackwood-rlcd:** screenshot multimodal decide. Different + input. + - **`24601/rh-guard`:** light note only (dry-run / submit opt-in / + fail-closed state checks). Not reward-hack detection. + + **Placement.** Exact-text keep/drop among observed elements + (`applied-mappings.md` §2) + mixed architecture (code owns + envelope, dry-run, submit gate; model selects among candidates). + Pillar: search/control (one substituted classifier step) + + runtime-assurance sandwich (plan≠execute, fail-closed checkbox / + fill). Hole: perceive / keep-drop / replace-one-classifier-step. + Family: specialist encoder decide head (`tinyx` option-attention) + — **not** TypeSafe Jev, **not** GLiNER. Fail-closed on actuation. + Eval path: none published (source-only); metric *names* are + specified. **Empirical** as README / MODEL_CARD behavior. + **Hypothesis** that a future `form-v0` checkpoint fills the + profile. Cards: `judgment-class.md` (primary); + `mixed-architecture.md` (primary); `applied-mappings.md` §2; + `mappings.md` §9 / §12; `faq.md`; `mental-models.md`; + `validation.md`; `methods-catalog.md`; `toolbox-mapping.md`; + `agent-self-assessment.md`. No wrapper. + +### Omni / Jev-omni + +Not multimodal pixels. Strong **computer-use composition** signal: +perception (a11y/snapshots) → specialist decide → verified act. +Form specialist, not pixels-in. Weights TBD — Watch for +`cua-s1-form-v0`. Archive + landscape pointer. diff --git a/research/refresh-log.md b/research/refresh-log.md index 41cbcea..2d96899 100644 --- a/research/refresh-log.md +++ b/research/refresh-log.md @@ -583,6 +583,29 @@ ecosystem, CHANGELOG, README. - notes.md §53; sources.json; findings.md batch #37. No wrapper. +## 2026-09-18 23:21 UTC — Cua-S1 specialist form System One (~17:21 Boise) + +- Folded into open PR #2 (`cursor/augustus-store-envelope-00b4`). + Docs-only. Not a competing PR. Archer 27B drop still **WATCH**. + **Not TypeSafe Jev. Not GLiNER. Not a general CUA. Not multimodal + pixels-in.** Source-only — no weights, no checkpoint scores. No + invented metrics. No wrapper. No `uv` / MCP / Driver how-to. +- HIGH: [trycua/cua `libs/cua-s1`](https://github.com/trycua/cua/tree/main/libs/cua-s1) + (parent MIT; ~23.3k★ this pass). Specialist computer-use research. + Profile `cua-s1-form-v0`. Byte encoder + option-attention: + fill/check/click/skip per observed element. Fill values selected + from extracted `Label: value` pairs. Plan ≠ execute; dry-run + default; fail-closed checkbox/fill. Parallel "System One" naming + in CUA, not a TypeSafe contract. Same observe→score→act *job* as + jev-ultrafast / gliner2-ultrafast / solari-reflex / laya-mind2web. + Offline metric *names* only. Watch for `cua-s1-form-v0` artifact. +- Cards: SKILL.md, judgment-class (primary), mixed-architecture + (primary), applied-mappings §2, mappings §9/§12, faq, + mental-models, validation, methods-catalog, toolbox, + agent-self-assessment, composition-algebra, ecosystem, CHANGELOG, + README. +- notes.md §54; sources.json; findings.md batch #38. No wrapper. + diff --git a/research/sources.json b/research/sources.json index 3b26650..ba20aa4 100644 --- a/research/sources.json +++ b/research/sources.json @@ -1,6 +1,6 @@ { "refresh_cadence": "hourly", - "retrieved": "2026-09-18T23:15Z", + "retrieved": "2026-09-18T23:21Z", "sources": [ { "kind": "docs", @@ -1598,6 +1598,12 @@ "title": "tamaratran/jev-pruner", "url": "https://github.com/tamaratran/jev-pruner", "note": "MIT. Created 2026-09-18T03:00:58Z; 5 stars at attached capture, 7 live this pass. Claude Code plugin (marketplace id still fast-jev-output): after Bash, Jev Noul-prunes stdout chunks before the main LLM sees them \u2014 no summary. Hard envelope (\u226410k estimated tokens / JSON-diff-whole-doc untouched) then soft Noul; fail-safe keep original; full archive. Codex is opt-in wrapper. Same evidence-preserving family as fast-jev-compaction / gliner25-compaction; different job (command output vs session memory). Manual sweep (theirs): needles 24/24; mean reduction 83% on trim scenarios. Harbor plugin-eval cannot reach Jev. Terminal-Bench paired pilot is integration, not a full bench. notes.md \u00a753." + }, + { + "kind": "github", + "title": "trycua/cua libs/cua-s1", + "url": "https://github.com/trycua/cua/tree/main/libs/cua-s1", + "note": "Parent MIT; ~23.3k stars this pass (23339 live; updated 2026-09-18T23:21:08Z). Specialist System One computer-use research (form-v0 profile). Source-only: no weights, no checkpoint scores. tinyx byte encoder + option-attention head chooses fill/check/click/skip per observed element; does not generate values or selectors. Plan \u2260 execute; dry-run default; execute/submit independent opt-ins; fail-closed unknown checkbox / fill without advertised token set_value. Not TypeSafe Jev \u2014 parallel System One naming in CUA. Same observe\u2192score-among-candidates\u2192code-acts job as jev-ultrafast / gliner2-ultrafast / solari-reflex / laya-mind2web. Watch for cua-s1-form-v0 artifact. notes.md \u00a754." } ] } From 4aeca62439c939bd4e9ceeac8e746dbd1869727b Mon Sep 17 00:00:00 2001 From: Cursor Agent Date: Sat, 19 Sep 2026 00:03:41 +0000 Subject: [PATCH 13/43] Fold 17:48 Boise watch: CUDA replica, decision-native RAG, AMBIGUOUS eval Fold jevify (uncalibrated local likelihoods), decision-native RAG, carryforward, hunch, explore-typesafe-ai, and jev-baselines-eval (honest negative + cascade sign-flip) into PR #2 skill cards. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com> --- .agents/skills/augustus/SKILL.md | 20 +- .../references/agent-self-assessment.md | 23 +- .../augustus/references/applied-mappings.md | 34 +- .agents/skills/augustus/references/faq.md | 60 ++- .../augustus/references/judgment-class.md | 13 +- .../skills/augustus/references/mappings.md | 48 +- .../augustus/references/mental-models.md | 5 + .../augustus/references/methods-catalog.md | 12 +- .../augustus/references/mixed-architecture.md | 20 +- .../augustus/references/toolbox-mapping.md | 7 +- .../skills/augustus/references/validation.md | 32 ++ CHANGELOG.md | 28 ++ README.md | 17 +- docs/ecosystem.md | 18 +- research/archive/findings.md | 64 +++ research/notes.md | 429 ++++++++++++++++++ research/refresh-log.md | 29 ++ research/sources.json | 74 ++- 18 files changed, 895 insertions(+), 38 deletions(-) diff --git a/.agents/skills/augustus/SKILL.md b/.agents/skills/augustus/SKILL.md index 3450eae..b8c3ac2 100644 --- a/.agents/skills/augustus/SKILL.md +++ b/.agents/skills/augustus/SKILL.md @@ -1,6 +1,6 @@ --- name: augustus -description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill, open multimodal RLCD / blackwood), announced open decision-model (Watch — still not landed), constrained-AR (TypeAR, pcdServer, decision-token LoRA), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"wait for Archer vs open multimodal RLCD\", \"screenshot/DOM candidates → typed Choice\", \"eval path\", \"jevals\", \"Harbor taskset\", \"shared bake-off ECE/NLL/Brier\", \"LLM-as-judge is not the System One score\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge (allowlist *proves*; fail-open cannot block), \"missing other → confident wrong Choice\", \"lint the Jev request\", \"training confronts Choice other / none-of-the-above\", \"S1 reflex keeps control / optional S2 one-use advice\", \"soft AGENTS.md rules vs linter\" (Abide / jev-pref), \"edit-phase vs turn-phase observation window\", banded confidence fail-open preference lint, \"extractive selection / pointer-not-generator\", \"encoder GLiNER compaction vs Jev Score compaction (same job; pointer not summarizer)\", \"fail-closed keep_full under mutation envelope\", \"CI flaky-vs-real merge gate\", \"fail-open VOI wake/resume (Horvitz)\", \"claim/evidence Stop integrity\", \"S1 extract + escalate-S2 indexer\", \"Harbor on/off routing\", \"policy-as-judgment PR marshal\", \"shadow-mode compaction rollout\", \"Jev Ultrafast vs GLiNER Ultrafast (observe-score-act backend-agnostic)\", \"hybrid local decide + remote fill\", \"DONE ≠ verified success\", \"observed a11y/DOM candidates vs screenshot multimodal\", \"evidence-preserving stdout prune (not summarize)\", \"hard token/format envelope then soft Noul\", \"fail-safe keep original on prune failure\", \"stdout prune vs session compaction\", \"specialist S1 computer-use (Cua-S1 form-v0; not TypeSafe Jev)\", \"plan ≠ execute / dry-run default\", \"observed-element option head (fill/check/click/skip)\", \"local /v1/systemone drop-in (stub until hf scorer)\", \"dataframe-native semantic columns\", \"route≠memory\", \"advisory sidecar receipts\", \"structure induction over bags\", \"AST ∩ semantic lint\", \"extractable-from-state / retrieve first\", \"decision-model vs constrained-LLM bake-off\", \"dual-process S1 decide / S2 generate\", \"combinatorial grid ≠ extractive\", TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, or \"formally verify with Jev\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." +description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill, open multimodal RLCD / blackwood), announced open decision-model (Watch — still not landed), constrained-AR (TypeAR, pcdServer, decision-token LoRA), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"wait for Archer vs open multimodal RLCD\", \"screenshot/DOM candidates → typed Choice\", \"eval path\", \"jevals\", \"Harbor taskset\", \"shared bake-off ECE/NLL/Brier\", \"LLM-as-judge is not the System One score\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge (allowlist *proves*; fail-open cannot block), \"missing other → confident wrong Choice\", \"lint the Jev request\", \"training confronts Choice other / none-of-the-above\", \"S1 reflex keeps control / optional S2 one-use advice\", \"soft AGENTS.md rules vs linter\" (Abide / jev-pref), \"edit-phase vs turn-phase observation window\", banded confidence fail-open preference lint, \"extractive selection / pointer-not-generator\", \"encoder GLiNER compaction vs Jev Score compaction (same job; pointer not summarizer)\", \"fail-closed keep_full under mutation envelope\", \"CI flaky-vs-real merge gate\", \"fail-open VOI wake/resume (Horvitz)\", \"claim/evidence Stop integrity\", \"S1 extract + escalate-S2 indexer\", \"Harbor on/off routing\", \"policy-as-judgment PR marshal\", \"shadow-mode compaction rollout\", \"Jev Ultrafast vs GLiNER Ultrafast (observe-score-act backend-agnostic)\", \"hybrid local decide + remote fill\", \"DONE ≠ verified success\", \"observed a11y/DOM candidates vs screenshot multimodal\", \"evidence-preserving stdout prune (not summarize)\", \"hard token/format envelope then soft Noul\", \"fail-safe keep original on prune failure\", \"stdout prune vs session compaction\", \"specialist S1 computer-use (Cua-S1 form-v0; not TypeSafe Jev)\", \"plan ≠ execute / dry-run default\", \"observed-element option head (fill/check/click/skip)\", \"local /v1/systemone drop-in (stub until hf scorer)\", \"dataframe-native semantic columns\", \"route≠memory\", \"advisory sidecar receipts\", \"structure induction over bags\", \"AST ∩ semantic lint\", \"extractable-from-state / retrieve first\", \"decision-model vs constrained-LLM bake-off\", \"dual-process S1 decide / S2 generate\", \"combinatorial grid ≠ extractive\", TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, \"uncalibrated local likelihoods vs Noul / CUDA replica\", \"decision-native RAG retrieve wide then decide then evidence set\", \"verbatim session ledger / scored recall\", \"judgment as language primitive / English-as-config\", \"pre-registered AMBIGUOUS eval / cascade sign-flip\", \"healthcare Harbor-shaped S1+S2\", \"pre-exec tool gate allow/block/review\", or \"formally verify with Jev\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." license: MIT metadata: version: 0.3.0 @@ -54,7 +54,7 @@ classical method you already trust, substitute it, classify the win "paraphrase brittleness", "allowlist then judge", "TOCTOU-of-Noul", "Jev inside the database / sqlite-jev", "Jev picks bitrate / join order / the model", "wait for Archer", "lint the request / missing - other", "training confronts Choice other / none-of-the-above", "soft AGENTS.md rules vs the linter", "screenshot Choice / omni System One", "extractive quotes / pointer not generator", "compaction summarize vs pointer", "encoder vs Jev compaction backend", "shadow-mode compaction rollout", "CI flaky-vs-real merge gate", "fail-open VOI wake/resume", "claim vs session evidence", "S1 indexer escalate-S2", "Harbor on/off routing", "fail-open vs fail-closed wake vs CI gate", "encoder vs Jev computer-use backend", "hybrid local decide + remote fill", "DONE vs verified success", "stdout prune vs session compaction", "hard envelope then Noul prune", "Cua-S1 vs TypeSafe Jev", "plan vs execute dry-run", "specialist computer-use vs general agent", "local drop-in vs stub scorer", "route vs memory", "when does it hold / extractable from state", "decision model vs constrained LLM", "dual-process S1/S2", or "combinatorial grid vs extractive": + other", "training confronts Choice other / none-of-the-above", "soft AGENTS.md rules vs the linter", "screenshot Choice / omni System One", "extractive quotes / pointer not generator", "compaction summarize vs pointer", "encoder vs Jev compaction backend", "shadow-mode compaction rollout", "CI flaky-vs-real merge gate", "fail-open VOI wake/resume", "claim vs session evidence", "S1 indexer escalate-S2", "Harbor on/off routing", "fail-open vs fail-closed wake vs CI gate", "encoder vs Jev computer-use backend", "hybrid local decide + remote fill", "DONE vs verified success", "stdout prune vs session compaction", "hard envelope then Noul prune", "Cua-S1 vs TypeSafe Jev", "plan vs execute dry-run", "specialist computer-use vs general agent", "local drop-in vs stub scorer", "route vs memory", "when does it hold / extractable from state", "decision model vs constrained LLM", "dual-process S1/S2", "combinatorial grid vs extractive", "uncalibrated local likelihoods", "decision-native RAG", "verbatim ledger vs summary", "judgment as language primitive", or "cascade sign-flip / calibration theater": read `references/faq.md`, then `references/mental-models.md`, then `references/mixed-architecture.md`, then @@ -115,11 +115,11 @@ classical method you already trust, substitute it, classify the win |---|---|---| | Mental models across domains (not SWE-only) | EU, abstention, VOI, MCDA, SDT, search/control, Leveson, NATM/Norman/snap-fit; **extractable-from-state boundary map** (self-contained vs needs outside knowledge) | `references/mental-models.md` | | Judgment-model class (Jev is exemplar, not monopoly) | Species: decide / locate (GLiNER) / categorize (GLiClass) / rank / perceive; open heads include encoder DeBERTa, LoRA distill, openjev-lm, kev. Compaction job is backend-agnostic (Jev Score/Noul vs GLiNER2.5 encoder). Indexer cousin: GLiNER extract + escalate-S2 (10–50× unfilled). Computer-use observe→score-among-candidates→code-acts is backend-agnostic (Jev Ultrafast ↔ GLiNER2 Ultrafast ↔ Cua-S1 specialist; Cua-S1 is not TypeSafe Jev) | `references/judgment-class.md` | -| Open weights vs constrained decoding vs proprietary API | Three open paths: encoder open-jev / AR constrained decode (TypeAR + pcdServer; decision-token LoRA; packed one-forward logprob on open LLMs) / trained decision-only (Laya + ONNX port, Nimble, kev, **blackwood-rlcd** multimodal now, Archer Watch still Watch). Local `/v1/systemone` surfaces: jev-local (stub until `hf`), kev (trained pointer), von (tiny SAN — not a replica). Softmax over allowed tokens ≠ Noul | `references/judgment-class.md` (when-to-use table) | +| Open weights vs constrained decoding vs proprietary API | Three open paths: encoder open-jev / AR constrained decode (TypeAR + pcdServer; decision-token LoRA; packed one-forward logprob on open LLMs; **CUDA/PyTorch local replica** jevify — uncalibrated likelihoods ≠ Noul) / trained decision-only (Laya + ONNX port, Nimble, kev, **blackwood-rlcd** multimodal now, Archer Watch still Watch). Local `/v1/systemone` surfaces: jev-local (stub until `hf`), kev (trained pointer), von (tiny SAN — not a replica). Softmax over allowed tokens ≠ Noul | `references/judgment-class.md` (when-to-use table) | | Entropy as allocator (low / medium / high) | Typed low+medium decisions → System One marginals; high-entropy synthesis → frontier decoder. Product rhetoric, not a meter. **Hypothesis** | `references/judgment-class.md` | | Formal / semi-formal (proof vs judgment) | Sensor vs constraint vs searchlight; Alloy vs Apalache; DST trio; TOCTOU-of-Noul, AI×FM | `references/formal-methods.md` (one-screen: `references/formal-semi-formal.md`) | -| Mixed architecture (judgment model + LLM) | Provider judges, LLM writes, code owns control; not a stack replacement. Advisory sidecar never changes host routing. Dual-process: S1 decides, S2 generates (routing accuracy unmeasured); S1 specialists + S2 coordinator is the same split (description-only greenfield this hour). Internals ≠ FSM; placement is a **component node**. Browser-use strength = DOM-as-text + speculative fan-out, not vision. Fail polarity is per act: skip-wake fail-open vs merge-gate BLOCK fail-closed. Hybrid local decide + remote fill; `DONE` ≠ verified success. Specialist computer-use: plan ≠ execute, dry-run default (Cua-S1; not TypeSafe Jev) | `references/mixed-architecture.md` | -| Context sieve | Relevance Noul per block; always-keep set in code; stub + recall key. Encoder cousin: GLiNER2.5 retention Choice + char-offset spans (gliner25-compaction); fail-closed keep_full; shadowMode default. Stdout cousin: jev-pruner (Jev Noul after hard ≤10k/JSON-diff envelope; fail-safe original; archive) | `references/applied-mappings.md#1-context-sieve` | +| Mixed architecture (judgment model + LLM) | Provider judges, LLM writes, code owns control; not a stack replacement. Advisory sidecar never changes host routing. Dual-process: S1 decides, S2 generates (routing accuracy unmeasured); S1 specialists + S2 coordinator is the same split (description-only greenfield this hour). Internals ≠ FSM; placement is a **component node**. Browser-use strength = DOM-as-text + speculative fan-out, not vision. Fail polarity is per act: skip-wake fail-open vs merge-gate BLOCK fail-closed. Hybrid local decide + remote fill; `DONE` ≠ verified success. Specialist computer-use: plan ≠ execute, dry-run default (Cua-S1; not TypeSafe Jev). Judgment as a language primitive (Ruby `almost_certain?`/`pick`/`rate`). Decision-native RAG: retrieve wide → decide → evidence set → LLM | `references/mixed-architecture.md` | +| Context sieve | Relevance Noul per block; always-keep set in code; stub + recall key. Encoder cousin: GLiNER2.5 retention Choice + char-offset spans (gliner25-compaction); fail-closed keep_full; shadowMode default. Stdout cousin: jev-pruner (Jev Noul after hard ≤10k/JSON-diff envelope; fail-safe original; archive). Session-ledger cousin: carryforward (verbatim facts; Jev scores recall; rules never judged; fail-open dump) | `references/applied-mappings.md#1-context-sieve` | | Exact-text keep / drop | Choice include/exclude/mixed over candidates code already holds. Extractive quotes / pointer-not-generator (model never writes the excerpt; char-offset compaction same species). Observed a11y/DOM controls: score among them; code clicks (Jev or GLiNER2 or Cua-S1 option-attention) | `references/applied-mappings.md#2-exact-text-keep--drop` | | Environment / harness triage | Scan every step for env failure; LLM autopsy only on flags. Merge-gate cousin: cluster in code, judge labels cause, policy owns PASS/BLOCK (latch; judge never says ignore alone) | `references/applied-mappings.md#3-environment--harness-triage` | | Moderation and ranking | Hold-before-publish vs graded rerank; fail policy per action | `references/applied-mappings.md#4-moderation-and-ranking` | @@ -130,11 +130,11 @@ classical method you already trust, substitute it, classify the win | "It's just classification" / "not probabilistic programming" / stack-replacement FAQ | Typed judgment is a software primitive, not a new task; marginals are not a joint; Jev is not the only model | `references/faq.md` | | Feature engineering / multi-criteria analysis | Nouls + Score distributions as named features, weights in code | `references/mappings.md#1-semantic-judgments--features-and-explicit-utility` | | Selective classification / decision theory | Thresholds from action costs, abstention paths | `references/mappings.md#2-probabilistic-judgments--cost-sensitive-decisions` | -| Decision tables / circuits / state machines | Judgment predicates, code owns transitions | `references/mappings.md#3-semantic-predicates--decision-circuits` | -| Retrieve + expensive relevance fn | Bounded rerank of a retrieved shortlist | `references/mappings.md#4-retrieval--bounded-semantic-reranking` (independent TREC DL2019 benchmark: Jev zero-shot best MAP 0.4748, nDCG@10 0.683 vs tuned monoBERT 0.718 — competitive, not dominant) | +| Decision tables / circuits / state machines | Judgment predicates, code owns transitions. Language primitive: Ruby `chance`/`pick`/`rate` as control flow (hunch; English-as-config; fail polarity per action) | `references/mappings.md#3-semantic-predicates--decision-circuits` | +| Retrieve + expensive relevance fn | Bounded rerank of a retrieved shortlist. Decision-native RAG: retrieve wide → decide explicitly → evidence set → conflict resolve → reason only over kept evidence (embeddings stay candidate generators; no universal benchmark) | `references/mappings.md#4-retrieval--bounded-semantic-reranking` (independent TREC DL2019 benchmark: Jev zero-shot best MAP 0.4748, nDCG@10 0.683 vs tuned monoBERT 0.718 — competitive, not dominant) | | Store as semantic index (SQL / SQLite / zoxide / dataframe) | Cheap exact predicates first; typed questions on the remainder. In-engine extension (sqlite-jev) vs CLI rewrite (jevql) vs path index (joxide) vs dataframe columns (jevpandas / jevframe) | `references/mappings.md#4-retrieval--bounded-semantic-reranking` | -| Soft judgment inside a hard envelope | Model may only match the deterministic policy or be more conservative (bitrate ABR; query-planner override-when-confident; compaction mutations/shell operators → keep_full; stdout prune: ≤10k/JSON-diff-whole-doc untouched, then Noul; Cua-S1: plan≠execute, dry-run, fail-closed checkbox/fill) | `references/mappings.md#12-runtime-assurance-sandwich-hypothesis`; `references/mappings.md#18-structural-prove--soft-remainder-hypothesis-as-domain-general-empirical-as-named-shapes` | -| Value of information / gather as an act | Pay for another observation only if EV(decision) improves more than cost; abstain from calling *any* model when a regex already answers (meta-VOI). Fail-open wake/resume: skip the LLM turn only if the judge answers and p(wake) is low (Horvitz) | `references/mappings.md#6-value-of-information--gather-as-an-enumerated-act` (**Hypothesis** until a labeled act/outcome log; 149-row receipt is Empirical as a shape; wakegate 21/21 is smoke) | +| Soft judgment inside a hard envelope | Model may only match the deterministic policy or be more conservative (bitrate ABR; query-planner override-when-confident; compaction mutations/shell operators → keep_full; stdout prune: ≤10k/JSON-diff-whole-doc untouched, then Noul; Cua-S1: plan≠execute, dry-run, fail-closed checkbox/fill; pre-exec toolgate: allow/block/review — Jev is not authorization; guard error/timeout stops) | `references/mappings.md#12-runtime-assurance-sandwich-hypothesis`; `references/mappings.md#18-structural-prove--soft-remainder-hypothesis-as-domain-general-empirical-as-named-shapes` | +| Value of information / gather as an act | Pay for another observation only if EV(decision) improves more than cost; abstain from calling *any* model when a regex already answers (meta-VOI). Fail-open wake/resume: skip the LLM turn only if the judge answers and p(wake) is low (Horvitz). Selective memory: verbatim ledger + scored recall (carryforward; rules never judged; 9×3 is a hint) | `references/mappings.md#6-value-of-information--gather-as-an-enumerated-act` (**Hypothesis** until a labeled act/outcome log; 149-row receipt is Empirical as a shape; wakegate 21/21 is smoke) | | Signal detection / ROC | Criterion and operating point from costs and base rate, not accuracy | `references/mappings.md#7-signal-detection--criterion-not-accuracy` (**Hypothesis** for non-SWE plots) | | Org / safety control structure | Sensor ≠ constraint (Leveson); STPA if the sensor lies | `references/mappings.md#8-control-structure--sensor--constraint-leveson` | | Search / control loops (any domain) | Algorithm stays yours; judgment substitutes one classifier step | `references/mappings.md#9-search--control-loops--one-substituted-classifier-step` | @@ -151,7 +151,7 @@ classical method you already trust, substitute it, classify the win | Agent self-supervision / on-track detection | Pre-gate → output judge → done-check → supervisor nouls. S1 reflex keeps control; optional S2 is one-use advice. Claim/evidence Stop (anti-hallucinated-done); bounded Pi supervisor (shadow recovery, never generates commands) | `references/agent-self-assessment.md` | | Optimizer/program frameworks (Ax, DSPy) | Typed fields → one provider request; judge metrics; threshold discipline. Ax and DSPy climb LM-program knobs only | `references/optimizer-integration.md` | | Perception → decision pipeline / measure / hill-climb | Stages with a versioned state contract; frozen taskset; DSPy/Ax only on the LM-program slice. Specialist composition stays **Hypothesis**; open multimodal decide (blackwood-rlcd) is a named receipt. Structured observe→decide→verified-act (no screenshots) is a computer-use speed-layer receipt (Jev or GLiNER2 or Cua-S1 specialist; Cua-S1 source-only, not TypeSafe Jev). Same section as the row below | `references/validation.md#eval--hill-climb` | -| Eval & hill-climb | Decision-stage jevals hygiene; Harbor taskset × harness × runtime; one score-composition table. Shared bake-off exemplar: open-jev-laya-bench (ECE/NLL/Brier; LLM-as-judge is not the score). Harbor-style frozen protocol vs constrained LLMs: DMB (accuracy/calibration/latency/cost; raw logs). Feedstock: jevals-data CC-BY-4.0 boards + JSONL (recompute-from-logs). Collab-arm curriculum: llm_autonomous vs scripted_plus_jev vs llm_plus_jev (Wilson / McNemar). Negative: combinatorial grid assembly ≠ extractive (ARC-AGI Direct Jev 4/400). Harbor on/off routing: chess-engine tasks, hidden perft verifier, one-run preliminary (jev-gateway-bench). Pair CI merge-gate with Harbor + rh-guard. Harbor needle/noise stdout prune: jev-pruner (manual sweep theirs; plugin eval cannot reach Jev → fail-safe original; Terminal-Bench pilot is integration not a full bench). Cua-S1 specialist form: source-only (metric names, no checkpoint scores; not TypeSafe Jev) | `references/validation.md#eval--hill-climb` | +| Eval & hill-climb | Decision-stage jevals hygiene; Harbor taskset × harness × runtime; one score-composition table. Shared bake-off exemplar: open-jev-laya-bench (ECE/NLL/Brier; LLM-as-judge is not the score). Harbor-style frozen protocol vs constrained LLMs: DMB (accuracy/calibration/latency/cost; raw logs). Feedstock: jevals-data CC-BY-4.0 boards + JSONL (recompute-from-logs). Collab-arm curriculum: llm_autonomous vs scripted_plus_jev vs llm_plus_jev (Wilson / McNemar). Negative: combinatorial grid assembly ≠ extractive (ARC-AGI Direct Jev 4/400). Harbor on/off routing: chess-engine tasks, hidden perft verifier, one-run preliminary (jev-gateway-bench). Pair CI merge-gate with Harbor + rh-guard. Harbor needle/noise stdout prune: jev-pruner (manual sweep theirs; plugin eval cannot reach Jev → fail-safe original; Terminal-Bench pilot is integration not a full bench). Cua-S1 specialist form: source-only (metric names, no checkpoint scores; not TypeSafe Jev). Pre-registered independent eval: jev-baselines-eval (**both AMBIGUOUS**; cascade sign-flip at exact parity; confidence=1.0 theater; encoder-with-labels wins; serving-path ≠ model-speed). Healthcare Harbor-shaped: explore-typesafe-ai (synthetic FHIR; not clinically validated). Honest-negative PDF: databricks-jev-pdf-lab (no quality-equivalent Jev payoff) | `references/validation.md#eval--hill-climb` | | (meta) Finding new mappings & applications | Toolbox sweep: judgment-shaped component of a known method, substituted + falsified | `references/toolbox-mapping.md` | | Named methods / operators / theorems | Substitution tiers: operand-judgments, preconditioned theorems, non-substitutable | `references/methods-catalog.md` | | (meta) Where a judgment model sits relative to any construct | 11 positions + logical-operator rules + position×construct traversal as the application generator | `references/composition-algebra.md` | diff --git a/.agents/skills/augustus/references/agent-self-assessment.md b/.agents/skills/augustus/references/agent-self-assessment.md index 8f12f3c..839a180 100644 --- a/.agents/skills/augustus/references/agent-self-assessment.md +++ b/.agents/skills/augustus/references/agent-self-assessment.md @@ -12,7 +12,12 @@ relying: foreman, pi-jev, pi-warden, winnow, fast-jev-compaction, jev-judgment. parallel, thresholded in code: `destructive` (noul, hold ≥0.90), `exfiltrates_secrets` (≥0.70), `beyond_request_scope` (≥0.85), `impact_if_unwanted` (Score 4 levels, escalate ≥2.5). Ship shadow mode - first; enforce only after observing real traffic. + first; enforce only after observing real traffic. Productized pre-exec + cousin: [toolgate](https://github.com/fdemir/toolgate) — `allow` / + `block` / `review` before execution; guard error/timeout **stops**. + Jev is not authorization. 72-case synthetic, not independently + annotated. Distinct from the ndolinschi *vocabulary* (allow / + ask_human / deny) below (`notes.md` §55). 2. **Post-action output judge** (after the tool result exists, not before): `leaks_secret` (noul ≥0.90) and `failure_class` (Choice ~6 options). The gate sees intent; only the output judge sees what the command printed. @@ -91,6 +96,12 @@ relying: foreman, pi-jev, pi-warden, winnow, fast-jev-compaction, jev-judgment. [jev-pruner](https://github.com/tamaratran/jev-pruner) — Jev Noul on Bash chunks after a hard envelope; fail-safe original; archive (`notes.md` §53). Marketplace id still `fast-jev-output`. + Session-ledger cousin: + [carryforward](https://github.com/Dharundp6/jev-carryforward) — + verbatim JSONL; Jev scores which facts are still live; constraints + and corrections always return; fail-open dump if the scorer is + down. Nine entries × three tasks is a hint, not proof + (`notes.md` §55). Do not copy mcp add. ## Non-negotiable boundaries @@ -108,7 +119,11 @@ relying: foreman, pi-jev, pi-warden, winnow, fast-jev-compaction, jev-judgment. fails closed to `keep_full` (`notes.md` §50). Stdout prune is the same polarity: [jev-pruner](https://github.com/tamaratran/jev-pruner) fails closed - to original output (`notes.md` §53). + to original output (`notes.md` §53). Tool *execution* is the + other polarity: [toolgate](https://github.com/fdemir/toolgate) + stops on block / review-without-approval / guard error + (`notes.md` §55). Session-memory *omit* fails open (dump the + ledger): [carryforward](https://github.com/Dharundp6/jev-carryforward). - Cache identical judgments (~120s) and deduplicate sibling calls into one in-flight request. - pi-warden measured cost makes continuous guarding viable: ~$0.00004 and @@ -177,5 +192,7 @@ is Watch / empty repo this pass (`notes.md` §51). gets this variance check first, over frozen outputs, before its numbers mean anything. - Gate vocabulary is converging across implementations; reuse it rather - than inventing: allow / ask_human / deny (toolgate), ok / retry / + than inventing: allow / ask_human / deny (ndolinschi *vocab*), + allow / block / review ([toolgate](https://github.com/fdemir/toolgate) + *product* — Jev is not authorization; `notes.md` §55), ok / retry / escalate / stop (harnessjudge). Same shape as the lifecycle gates above. diff --git a/.agents/skills/augustus/references/applied-mappings.md b/.agents/skills/augustus/references/applied-mappings.md index 53f53b6..7f9f409 100644 --- a/.agents/skills/augustus/references/applied-mappings.md +++ b/.agents/skills/augustus/references/applied-mappings.md @@ -66,13 +66,23 @@ recovery. Marketplace id still `fast-jev-output`. Codex is opt-in wrapper, not automatic interception. Same author as fast-jev-compaction; complementary, not a duplicate. Do not copy the plugin (`notes.md` §53). +**Session-ledger cousin, same family, different job (Empirical as +README behavior, 2026-09-18 ~17:48):** +[carryforward](https://github.com/Dharundp6/jev-carryforward) — +verbatim JSONL facts (`record`); Jev Noul-scores `recall` against +the current task. Nothing summarised or deleted. Constraints and +corrections **always return in full** (Jev never votes on a rule). +Fail-open: no key → whole list. Thresholds 0.60 full / 0.30–0.60 +one line are *theirs*. Nine entries × three tasks is a **hint, not +proof** (`notes.md` §55). Do not copy `mcp add`. Local teacher-copy for the same hole: [`SargeDev/jev-gate-student-b`](https://huggingface.co/SargeDev/jev-gate-student-b) (Qwen2.5-0.5B LoRA; P(relevant) from yes/no logits; 148,160-row [`jev-distill-corpus`](https://huggingface.co/datasets/SargeDev/jev-distill-corpus); -card: fail-open on errors). That is System One as a **memory/context +card: fail-open on errors; HF card **unchanged** this pass vs +`notes.md` §33). That is System One as a **memory/context gate**, not an action permit: vector recall → local yes/no → inject or -stub (`notes.md` §33, §44). Agreement with Jev labels is not independent +stub (`notes.md` §33, §44, §55). Agreement with Jev labels is not independent gold (`notes.md` §33). Official cousin: classifying RAG passages cookbook (**Contract**). **Counterexample**: one Noul "is this log useful?" over 3k lines — that is nine judgments pretending to be one. **Test**: recall @@ -238,9 +248,25 @@ re-filters); LlamaIndex Jev rerank **Empirical** BEIR nfcorpus MiniLM zero-shot MAP 0.4748 / nDCG@10 0.683 vs monoBERT 0.718 (`mappings.md` §4). Realtime ~200ms chat claims remain **Hypothesis** as a number. **Counterexample**: using top-1 Choice as a relevance score across queries; dropping RAG -chunks fail-closed so a timeout empties the context. **Test**: moderation +chunks fail-closed so a timeout empties the context. +**Decision-native RAG (Empirical as architecture; Hypothesis as a +universal win, 2026-09-18 ~17:48):** +[decision-native-rag-skills](https://github.com/emergency-lee/decision-native-rag-skills) +— retrieve wide → decide explicitly → evidence set → conflict +resolve → reason only over kept evidence. Embeddings stay candidate +generators. Provider-agnostic; no bundled Python harness; **no +universal benchmark**. Default migration gates are starting +targets. Offline replay → shadow → canary → A/B (`notes.md` §55). +Do not ship because an LLM judge prefers it. +**Recursive file search (MED; distinguish from federated web):** +[kazuhideoki/jev-search](https://github.com/kazuhideoki/jev-search) +scores local files then fzf — **not** +[superagents-lab/jev-search](https://github.com/superagents-lab/jev-search) +(web lanes). Max-over-chunks is not a calibrated whole-file +probability. No LICENSE this pass. +**Test**: moderation cost/coverage + false-hold vs false-publish; ranking recall *separate* -from nDCG; select misroute rate. +from nDCG; select misroute rate; required-evidence recall vs Top-K. ## 5. Skill / tool routing diff --git a/.agents/skills/augustus/references/faq.md b/.agents/skills/augustus/references/faq.md index e6b2961..59164af 100644 --- a/.agents/skills/augustus/references/faq.md +++ b/.agents/skills/augustus/references/faq.md @@ -455,8 +455,55 @@ the speed/econ class — not the stub, not kev, **not a calibrated Jev replica**. Do not copy its vs-Jev table. [open-alternative-jev](https://github.com/ikermoel/open-alternative-jev) packs one-forward logprobs on an open LLM you already have (RACE-H -92.9% @ 4.55 q/s); **not a Jev reproduction**. A green smoke test on -the stub is not a bake-off. `judgment-class.md`; `notes.md` §48, §49. +92.9% @ 4.55 q/s); **not a Jev reproduction**. +[jevify](https://github.com/Mintzs/jevify) is a CUDA/PyTorch cousin +on Qwen2.5-1.5B (`ora_decision_engine`): CUDA graphs, branch kernels, +literal-label scoring. **Uncalibrated model likelihoods, not +measured correctness** — softmax over A/B/C is not a Noul. Independent +of Distillation. No LICENSE this pass. Default refund workflow is +not a validated policy. A green smoke test on +the stub is not a bake-off. `judgment-class.md`; `notes.md` §48, §49, §55. + +## Are local CUDA likelihoods a Noul? + +No. [jevify](https://github.com/Mintzs/jevify) (and packed-logprob +cousins) return **uncalibrated model likelihoods**. Do not threshold +them as P(permit) or as calibrated abstention. Temperature / ECE on +*your* labels if you use the surface. Softmax over allowed tokens ≠ +Noul. `judgment-class.md`; `notes.md` §55. + +## Should RAG stop at Top-K / a reranker? + +Not if the hole is **evidence**. +[decision-native-rag-skills](https://github.com/emergency-lee/decision-native-rag-skills): +retrieve wide → decide explicitly → evidence set → resolve conflicts +→ reason only over kept evidence. Embeddings stay candidate +generators. Provider-agnostic; no bundled harness; **no universal +benchmark**. Default migration gates are starting targets, not +promises. Do not ship because an LLM judge prefers it. +`mappings.md` §4; `notes.md` §55. + +## Did Jev beat nano as an escalation gate? + +Not in the pre-registered independent eval +[jev-baselines-eval](https://github.com/ickma2311/jev-baselines-eval) +(2026-09-18). **Both experiments AMBIGUOUS.** Cascade Δ +0.265 at a +1pp-below-frontier target; **at exact parity the sign flips** +(R_jev=1.000 vs nano 0.730) because Jev confidence is exactly 1.0 on +102/200 items including 6 wrong. AUROC error-ranking neither +direction established; **no ECE**. Encoder with labels wins Banking77 +(0.933 / 9 ms). Recorded call duration ~2.2×, **serving-path not +model-speed**, not 40–200×. Same-day errata three rounds. This is +the jevals/Harbor practice exemplar this hour (honest negative + +calibration theater). `validation.md`; `notes.md` §55. + +## Did Jev pay off for Precision PDF extraction? + +Not in +[databricks-jev-pdf-lab](https://github.com/laurentfabre/databricks-jev-pdf-lab). +**No quality-equivalent, end-to-end Jev payoff demonstrated.** +Compact requests cut tokens but changed 26/236 recommendations. +**No OSS license selected.** Typed output is not truth. `notes.md` §55. ## Should the model write the quote / the citation / the click? @@ -532,7 +579,14 @@ fails closed at the gate: [latch](https://github.com/CaseReed/latch) stays fail-open. [if-ai](https://github.com/Victor-Casado/if-ai) fails the Action on error / empty / low confidence. jevgate cannot block; pi-jev-approver fails closed without a key; Abide is fail-open -on diffs. Same sandwich, opposite authorized act. `notes.md` §50, §51, §53. +on diffs. Session-memory *omit* fails open (dump the ledger): +[carryforward](https://github.com/Dharundp6/jev-carryforward). +Tool *execution* fails closed on block/timeout: +[toolgate](https://github.com/fdemir/toolgate) (Jev is not +authorization). Ruby validations in +[hunch](https://github.com/carldaws/hunch) `rescue nil` at save — +spam gates should not. Same sandwich, opposite authorized act. +`notes.md` §50, §51, §53, §55. ## Is observe→score→act Jev-only? diff --git a/.agents/skills/augustus/references/judgment-class.md b/.agents/skills/augustus/references/judgment-class.md index 74b1e8c..4fac6c5 100644 --- a/.agents/skills/augustus/references/judgment-class.md +++ b/.agents/skills/augustus/references/judgment-class.md @@ -497,7 +497,7 @@ is the generator, not a sixth surface. | **Archer open decision-model** | **Watch.** No Hub weights this pass. "Smarter than Jev" is a claim against *his* calibration/order warnings | Same *hole* as Jev when it ships | 27B dense for one-forward-pass local speed once AR is removed; MoE next, then shrink. Quant-friendly is a claim | Healthcare AU data-residency / deployment control, **not** anti-TypeSafe | Multimodal, no audio. Text post-training reportedly generalizes to images with little intentional multimodal training | Unknown until the drop | | **TypeAR / pcdServer** (constrained AR) | Next-token constraint ≠ Noul. No abstention primitive. Public logit dump: [`Mikhail/mini-jev-runs`](https://huggingface.co/datasets/Mikhail/mini-jev-runs) (27.9k; scores "deliberately *not* calibrated"). Decision-token QLoRA trains *that* token under parallel constrained decode (`Foodoo1/Qwen3-14B-RLCD-Decision-LoRA`; synthetic fraud receipt, not a financial product; `notes.md` §46) | TypeAR sequential conditions later fields; pcdServer batches independent fields after one prefix. Neither is gather-as-act | TypeAR 5.8× is *their* K=16 boolean example. pcdServer: native llama.cpp, Apple+Linux. Foodoo1: ~234 ms / 4-field broadcast on RTX 3090 4-bit (their figure) | Self-host the generator / GGUF / adapter | Whatever the base model has | TypeAR enums ≤16; pcdServer 2–256 strings, 1–63 fields | | **Encoder open-jev** (DeBERTa-v3-large 434M) | Public gold, CE+Brier, val temperature. In-domain ECE 0.022 / acc 0.854; OOD acc 0.690 / ECE 0.035. **Not** a Jev teacher-copy | One pass over state + all questions; 512 tok | Author: 28 ms / 10 questions H100; 1.8 s / 4q M1 Max CPU | apache-2.0, self-host | Text | Jev-shaped 255 / Score 2–10 / Noul; 512 ctx | -| **Tiny LoRA distill** (jev-gate-student-b) | Teacher-copy. P(relevant) from yes/no logits. Held-out n=60 vs vanilla 0.5B; 148,160-row corpus | Memory-gating / context sieve; **fail-open** on errors | Qwen2.5-0.5B LoRA; ~59 ms RTX 3060 | Local, apache-2.0 | Text | Binary relevance | +| **Tiny LoRA distill** (jev-gate-student-b) | Teacher-copy. P(relevant) from yes/no logits. Held-out n=60 vs vanilla 0.5B; 148,160-row corpus. HF card **unchanged** ~17:48 vs §33 (MAE 0.187 / Pearson 0.791 / 90%; ~59 ms RTX 3060; fail-open) | Memory-gating / context sieve; **fail-open** on errors | Qwen2.5-0.5B LoRA; ~59 ms RTX 3060 | Local, apache-2.0 | Text | Binary relevance | | **Nimble** (open LoRA recipe, not a distill) | Hard synthetic labels. They say temperature was not tuned to correctness rates. 324-row agreement is their receipt, not an ECE (`notes.md` §35) | Not a gather primitive | Their latency table, not re-run | Self-host the adapter. Model card Apache-2.0; repo license absent | Text only | Enum ≤26; 2,048 tokens | | **kev** (Qwen2.5-0.5B LoRA + pointer; Apache-2.0) | Public gold, CE. Held-out ECE 0.065 (0.031 after T=1.47); acc 0.799 on 1,350 ID questions. Isolation exact. **Not** a Jev teacher-copy (`notes.md` §45) | Laptop-local System One drop-in for development/eval; independent questions, one prefill | ~160 ms / 6 questions; ~1h45m train on M5; 38 MB adapter | Self-host; official `typesafe-sdk` with `base_url` | Text. Not multimodal. 0.5B knowledge | noul / choice 2–255 / score | | **Diffusion structured reads** (djev-spark) | Interface claim only. **Hypothesis** it beats a decision head on your labels (`notes.md` §36) | Optional sequential chunks, text-only | Their GX10 tables, not a class benchmark | DGX Spark container. Do not copy the route | Images are an extension; think and sequential reject images | README criteria, not copied here | @@ -538,6 +538,17 @@ not a Jev reproduction):** 8-bit; interference 6–9%; temperature scaling on *your* labels. Space demo. Economics of packing a shared state, not trained decision-only (`notes.md` §49). +**CUDA/PyTorch local replica (constrained-AR / logprob cousin, not +a Jev reproduction, not Distillation; 2026-09-18 ~17:48):** +[`Mintzs/jevify`](https://github.com/Mintzs/jevify) — Qwen2.5-1.5B, +package `ora_decision_engine` / CLI `ora-decision`. CUDA graphs, +branch kernels, literal-label scoring. Default `--answer-encoding +letters`. **Uncalibrated model likelihoods, not measured +correctness.** Default refund `workflow.json` is not a validated +policy. **No LICENSE file this pass.** Do not copy Windows CUDA/venv +(`notes.md` §55). Independent of Distillation; independent of +open-alternative-jev's RACE-H receipt — same *class*, different +repo. **Tiny SAN local surface (extreme speed/econ class, not a replica):** [`wfzyx/von`](https://github.com/wfzyx/von) — 14 MB Needle; `POST /v1/systemone`; sub-15 ms CPU *claim* / ~38 ms embed in their table; diff --git a/.agents/skills/augustus/references/mappings.md b/.agents/skills/augustus/references/mappings.md index d21c625..a8f2401 100644 --- a/.agents/skills/augustus/references/mappings.md +++ b/.agents/skills/augustus/references/mappings.md @@ -133,7 +133,17 @@ behavior, 2026-09-18 ~16:48):** [latch](https://github.com/CaseReed/latch) — Jev labels a clustered cause; a **table** maps cause × confidence × fingerprint → PASS / BLOCK / needs_human. The judge is a sensor, not the merge act -(`notes.md` §51). **Counterexample**: decomposing tool-trace +(`notes.md` §51). **Language primitive (Empirical as README / +example suite, 2026-09-18 ~17:48):** +[hunch](https://github.com/carldaws/hunch) — Ruby `chance` / +`pick` / `rate` map to Noul / Choice / Score; English is the +configuration; `Hunch.decide` batches over one `given:`. +Validations `rescue nil` = fail-open at save; spam gates should +fail closed. Stub backend for tests. Same interface ≠ same +guarantees for a future LLM backend. Cousin of probably-lang +(a language whose loop conditions are feelings) — this is a +library, not a new language. Do not copy gem/Rails +(`notes.md` §55). **Counterexample**: decomposing tool-trace verification into per-call schema nouls works; asking "is the trace correct" as one Noul hides nine judgments. **Test**: full truth table / transition cases incl. contradictory outputs, stale observations, invalid combos. @@ -189,6 +199,24 @@ contents leave the store (same residency warning as AU health). Do not copy SQL, env, or CLI flags. `notes.md` §42, §44, §46, §48. +**Decision-native evidence set (Empirical as architecture; +Hypothesis as a measured win, 2026-09-18 ~17:48):** +[decision-native-rag-skills](https://github.com/emergency-lee/decision-native-rag-skills) +promotes this card from "rerank a shortlist" to **retrieve wide → +decide → build an evidence set → resolve conflicts → generate only +over kept evidence**. Embeddings remain candidate generators; they +do not settle relevance, sufficiency, redundancy, conflict, time, +or authority. No bundled harness; no universal benchmark; default +migration gates are starting targets (`notes.md` §55). PubMed +title/abstract screening is the same *shape* on literature +([typesafe-screening-mcp](https://github.com/masa-med-ai/typesafe-screening-mcp): +include/maybe/exclude in code; 326 hits ~17 s ~$0.014 one run; +thresholds not calibrated; screening aid, not an SR replacement). +Local-file cousin: +[kazuhideoki/jev-search](https://github.com/kazuhideoki/jev-search) +(recursive files + fzf) — **not** superagents-lab/jev-search +(federated web). Max-over-chunks ≠ calibrated whole-file p. + ## 5. Hierarchy → bounded heuristic search **Method**: beam search over a meaningful taxonomy or candidate graph. @@ -305,6 +333,14 @@ Horvitz mixed-initiative: pay for the turn iff EV(decision) beats the token cost. Savings unmeasured. Same-author scenarios+question; not a benchmark (`notes.md` §51). Contrast pi-jev-approver fail-closed without a key and jevgate cannot-block. +**Selective memory / scored recall (Empirical as README behavior; +9×3 is a hint, 2026-09-18 ~17:48):** +[carryforward](https://github.com/Dharundp6/jev-carryforward) — +verbatim ledger; Jev scores which facts are still live for the +task; constraints/corrections always return (never judged). Fail- +open dump if the scorer is down. Pay for a scored brief iff it +beats dumping the whole file. No accuracy claim until a proper +test (`notes.md` §55). Do not copy mcp add. **Beyond SWE (Hypothesis until you log act/outcome pairs):** full PDF vs abstract; customer call vs CRM fields that already fail a hard rule (credit limit is exact); blood test vs @@ -737,6 +773,16 @@ Light sibling: typed Score/Nouls on the remainder; **fail-closed** without a key (different polarity from jevgate). rh-guard-adjacent; light note only (`notes.md` §48). +**Pre-exec tool product (Empirical as README wiring, not as +accuracy; 2026-09-18 ~17:48):** +[toolgate](https://github.com/fdemir/toolgate) — `allow` / `block` / +`review` before execution; guard error or timeout **stops** (fail- +closed on the execution act). Jev is a probabilistic check, **not +authorization**. 72-case synthetic set is not independently +annotated. Distinct from the ndolinschi *vocabulary* (allow / +ask_human / deny) already in `agent-self-assessment.md`. +`onReview` must obtain authenticated human approval +(`notes.md` §55). Do not copy pnpm. [`coldteadotai/abide`](https://github.com/coldteadotai/abide) is the same *family* on project instructions: the **linter proves** lintable rules; Jev Scores only residual soft AGENTS.md rules; fail-open, banded diff --git a/.agents/skills/augustus/references/mental-models.md b/.agents/skills/augustus/references/mental-models.md index d0d847f..433e1e8 100644 --- a/.agents/skills/augustus/references/mental-models.md +++ b/.agents/skills/augustus/references/mental-models.md @@ -488,6 +488,11 @@ Use these as *existence proofs of a position*. Write your own card. | Phishing / fraud screen | hold vs deliver | SDT criterion on a Noul | blocklist, SPF/DKIM exact (**Hypothesis**) | | Personal ops | cook done / not | "looks done" Noul | thermometer probe | | Org safety | stop the line | sensor Noul | interlock, two-person rule | +| Knowledge / RAG | reason only over kept evidence | retrieve wide → decide → evidence set (**Empirical** as architecture: decision-native-rag-skills; **Hypothesis** as a measured win) | Conflict/temporal/provenance in code; embeddings generate candidates | +| Session memory | next task sees last session's facts | scored recall over a verbatim ledger (**Empirical**: carryforward; 9×3 hint) | Constraints always-keep; fail-open dump; never summarize | +| Application control flow | `if` / `case` on a judgment | `chance`/`pick`/`rate` as language primitives (**Empirical**: hunch; English-as-config) | Fail polarity per action; stub backend | +| Healthcare huddle / recon / inbox | escalate / hold / route | S1 remainder after NEWS2/code (**Empirical** as synthetic report: explore-typesafe-ai; **not clinically validated**) | NEWS2, recon, routing in code; S2 blinded review | +| Intent cascade vs nano/encoder | escalate when unsure | pre-registered kill/go (**Empirical as practice**: jev-baselines-eval **AMBIGUOUS**; cascade sign-flip; encoder-with-labels wins) | Thresholds, serving-path honesty, ECE if you claim calibration | Rejected in every domain: replacing the ledger with a vibe; replacing the interlock with confidence; replacing the essay with a Noul; diff --git a/.agents/skills/augustus/references/methods-catalog.md b/.agents/skills/augustus/references/methods-catalog.md index 7330fe4..52c9e0c 100644 --- a/.agents/skills/augustus/references/methods-catalog.md +++ b/.agents/skills/augustus/references/methods-catalog.md @@ -31,10 +31,12 @@ judgment component is new). | Self-consistency / ensembling of judges | Repeated independent ratings of the same object | N repeats over one state (output tokens free); entropy/disagreement across repeats as the review signal | Aggregation, escalation policy | **Empirical recipe** (self-consistency: nouls cookbook) | | Judge qualification (interrater reliability) | A judge worth gating must be repeatable | Repeated judgments over frozen outputs before trusting either Jev or LLM as judge | Variance stats, agreement metrics | **Empirical recipe** (jev-as-a-judge: 224–279× tighter than GPT judge) | | Neyman–Pearson / selective classification | Decision threshold under error costs | One threshold per action, set on split A, reported on split B; abstention path | Loss model, ROC analysis | **Contract + empirical** (confidence-routing; evaluator script) | -| Value of information (EVPI / EVSI) | Whether another observation is worth its cost | Gather as an enumerated act; pay iff expected decision-loss drop > cost. Wake/resume: skip the LLM turn only if the judge answers and p is low | Cost of the observation; the loss table; skip-limit | **Hypothesis** as a numeric calculator; **Contract** as the placement (`mappings.md` §6). wakegate 21/21 is smoke (`notes.md` §51) | +| Value of information (EVPI / EVSI) | Whether another observation is worth its cost | Gather as an enumerated act; pay iff expected decision-loss drop > cost. Wake/resume: skip the LLM turn only if the judge answers and p is low. Selective memory: score the ledger against the task; dump on failure | Cost of the observation; the loss table; skip-limit; always-keep rules | **Hypothesis** as a numeric calculator; **Contract** as the placement (`mappings.md` §6). wakegate 21/21 is smoke (`notes.md` §51). carryforward 9×3 is a hint (`notes.md` §55) | | Signal detection (Green & Swets) | Evidence variable + criterion | Noul as noisy evidence; t from costs and base rate; ROC/PR on your labels | Operating point, base-rate tracking | **Hypothesis** for non-SWE plots; **Empirical** as moderation *shape* (`mappings.md` §7) | | Reliability calibration (Platt/temperature) | Raw scores → calibrated probabilities | Noul is natively calibrated **in-distribution only**; verify with reliability bins on your own population; re-fit a correction out-of-distribution | Calibration fitting, binning | **Empirical recipe** (ECE 0.0313 in-distribution; 32% OOD collapse — Archer Hume). Atlas: DAIR Emotion dangerous-high (48% / 0.819); DMB S5 ECE 0.246 (`notes.md` §49) | | Frozen-protocol bake-off vs constrained LLMs | Same items, accuracy + ECE + latency + cost + honesty | Decision-model as one contender class, not the score | Protocol, raw logs, baselines | **Empirical as Harbor/jevals practice** (DMB v2; jevals-data CC-BY-4.0 recompute-from-logs; `notes.md` §49) | +| Pre-registered cascade vs nano/frontier/encoder | Kill/go printed; cascade R at a stated accuracy target; error-ranking AUROC ≠ ECE | Confidence as an escalation signal is a *hypothesis to kill* | Pre-reg hash, paired CIs, margin sensitivity, serving-path controls | **Empirical as Harbor/jevals practice (honest negative)** (jev-baselines-eval: both AMBIGUOUS; cascade sign-flip at exact parity; confidence=1.0 on 102/200 incl. 6 wrong; encoder 0.933/9ms; ~2.2× serving-path; errata ×3; `notes.md` §55) | +| Healthcare S1+S2 on synthetic FHIR | Remainder after NEWS2 / recon / routing in code | Typed Noul/Score/Choice; S2 blinded review | Labels first; independent outcome policy | **Empirical as a named report, not clinical validation** (explore-typesafe-ai; 20 cases/scenario; Claude wrote labels; `notes.md` §55) | | Harbor on/off routing | Same task, routing on vs off, hidden verifier | Tool Choice per turn; cheaper unsolved is not a saving | Fresh gateway; perft / checks the agent never sees | **Empirical as a *shape* and one-run signal** (jev-gateway-bench chess-bugfix 36/36 both; 4 vs 6 LLM req; `notes.md` §51). Not a measurement until reps ≥5 | | Survey scoring / psychometrics | Rubric level judgment with defined anchors | Score with concrete level descriptions; probabilities read beside every score | Weighted aggregation, reliability analysis | **Contract** (score docs: split composite judgments) | @@ -51,7 +53,7 @@ judgment component is new). | STPA / STAMP control structure | Sensor reading vs enforced constraint | Judgment as sensor; constraints in policy/code/interlock; STPA table if the sensor lies | The constraint, the actuator, the probe | **Contract** as ownership; **Hypothesis** as domain product (`mappings.md` §8) | | PufferLib / Ocean env contracts | Does this episode look like a known trainer-bug mode? | Cluster failing episodes; never "the policy is correct" | Seeded serial env, Ocean sanity, observed rewards | **Hypothesis** as placement; **Contract** that Ocean is not a comparative baseline (`formal-methods.md` DST trio) | | Routing / dispatch (OR) | Which queue/agent owns this item | Choice + confidence-gated escalation; code owns capacity | Cost matrix, capacity constraints | **Empirical recipe** (intent-routing; LlamaIndex Jev selectors; skillranker) | -| Cascade / prefilter (IR) | Cheap reject before an expensive scorer or LLM | Per-candidate Noul/Score; fail-open on drop, fail-closed on dispatch | Candidate generation, always-keep set, recall keys | **Empirical recipe** (classifying RAG passages; jevprune; git-jev-stage) | +| Cascade / prefilter (IR) | Cheap reject before an expensive scorer or LLM | Per-candidate Noul/Score; fail-open on drop, fail-closed on dispatch. Decision-native: evidence-set after wide retrieve | Candidate generation, always-keep set, recall keys; conflict/provenance | **Empirical recipe** (classifying RAG passages; jevprune; git-jev-stage). **Empirical as architecture** (decision-native-rag-skills; Hypothesis as a measured win, `notes.md` §55) | | Knapsack / portfolio selection | Per-item feature vector from text | Fan-out nouls/scores as features; optimizer in code | Constraint solver, weights | **Hypothesis** (mapping 1 shape) | ## Information theory & signals @@ -59,11 +61,11 @@ judgment component is new). | Method | Judgment-shaped component | Jev substitution | Stays in code | Status | |---|---|---|---|---| | Entropy as uncertainty signal | Measuring "how spread is this belief" | Entropy of returned distributions across repeats or options — computed in code from returned probabilities | All arithmetic | **Empirical recipe** (cookbook pattern) | -| Detector / Neyman filter (context) | Is this artifact relevant to the current task? | One relevance Noul per block before it enters context; stub + recall key. Encoder cousin: retention Choice + span locate, copy verbatim. Indexer cousin: GLiNER extract on the bulk, escalate LLM on the tail. Stdout cousin: Noul per chunk after a hard size/format envelope | Cache, recall, safety keeps; mutation/shell envelope; ≤10k/JSON-diff pass-through; archive dropped spans; do not dump the repo if S1 failed to load | **Empirical recipe** (winnow ≤0.22 hide; compaction 2-noul rule; pi-jev-context hide-not-delete; gliner25-compaction GLiNER2.5 keep_full/keep_evidence/keep_call_only/drop, `notes.md` §50; s1-graphify-indexer degraded-load / no-invent-edges, 10–50× unfilled, `notes.md` §51; jev-pruner Bash stdout prune, `notes.md` §53) | +| Detector / Neyman filter (context) | Is this artifact relevant to the current task? | One relevance Noul per block before it enters context; stub + recall key. Encoder cousin: retention Choice + span locate, copy verbatim. Indexer cousin: GLiNER extract on the bulk, escalate LLM on the tail. Stdout cousin: Noul per chunk after a hard size/format envelope. Session-ledger cousin: score verbatim facts; rules never judged | Cache, recall, safety keeps; mutation/shell envelope; ≤10k/JSON-diff pass-through; archive dropped spans; do not dump the repo if S1 failed to load; fail-open dump of the ledger | **Empirical recipe** (winnow ≤0.22 hide; compaction 2-noul rule; pi-jev-context hide-not-delete; gliner25-compaction GLiNER2.5 keep_full/keep_evidence/keep_call_only/drop, `notes.md` §50; s1-graphify-indexer degraded-load / no-invent-edges, 10–50× unfilled, `notes.md` §51; jev-pruner Bash stdout prune, `notes.md` §53; carryforward, `notes.md` §55) | | Anomaly detection | Does this deviate from expected shape? | Guard nouls + harm Score over {input, output, tool trace} | Baselines, alert thresholds | **Empirical recipe** (guardrails cookbook; pi-jev output judge) | | Allowlist ∩ remainder (code-then-model) | Unlisted / unstructured leftovers after a **proof** | Typed questions only on the unknown tier; admit iff every p < τ | Proven/refused in code; cannot block unless a sandbox sits under | **Empirical recipe** (jevgate 0/59 unsafe unasked held-out; allowlist *proves* read-only verbs; doc-router 1.74× $). Domain-general: `mappings.md` §18 | | Decision-token LoRA (constrained-AR) | Specialize a generator for parallel constrained fields | Loss only on the single decision token; KV broadcast across fields | Schema, candidate tokens, policy | **Empirical recipe** as Foodoo1 200-case / 4-field receipt (fraud_risk 64→95%, overall 85.2→98.8%, ~234 ms); **Hypothesis** as a general recipe. Synthetic; not a financial product. Softmax ≠ Noul | -| Teacher distill of judgments | Copy a hosted decision API onto a small local head | LoRA / frozen-encoder heads trained on teacher answers | Independent gold labels; ECE on *your* cases | **Empirical recipe** as one 70-row run (openjev-lm 92.9%); **Hypothesis** as a general recipe | +| Teacher distill of judgments | Copy a hosted decision API onto a small local head | LoRA / frozen-encoder heads trained on teacher answers | Independent gold labels; ECE on *your* cases | **Empirical recipe** as one 70-row run (openjev-lm 92.9%); student-b n=60 MAE 0.187 / Pearson 0.791 / 90% vs vanilla (HF card unchanged ~17:48); **Hypothesis** as a general recipe | ## Verification & logic @@ -75,7 +77,7 @@ judgment component is new). | Spec vs artifact conformance (model checking *mindset*) | Property holds / violated / unverifiable for a named requirement | One Noul/Score per requirement, batched; violated → named rule back into context (pi-warden / Abide shape). This is **not** TLC/Apalache/GNATprove | Requirement enumeration, enforcement, logging; the **linter** if the rule is lintable; the real checker if you have one | **Empirical recipe** (pi-warden: 6→0 rule breaks, 150 paired runs; jev-pref: YOU define the rule; Abide: productized compile/calibrate/tune/replay, `notes.md` §47; if-ai: plain-English PR check, fail-closed on error, `notes.md` §51). Ownership split: `formal-methods.md` | | AST ∩ semantic lint | Semantic remainder after a parser already extracted units | Typed questions on Tree-sitter targets; do not execute scanned code | Parser, selection, fail-on; `tenbin` owns the lint *skill* | **Empirical as a shape** (jevscan 0.2.0rc4; not a calibration claim; `notes.md` §48) | | Alloy finder vs Apalache / TLC | Which bound, which counterexample, is the property tautological? | Triage instances/CEs; never "this spec looks right" | Analyzer / SMT / explicit-state engine | **Hypothesis** as product; **Contract** as ownership (`formal-methods.md` §2) | -| Type-checking analog | Does this planned call match the schema/operation/target? | Decomposed nouls over {request, schema, trace}; never trust a Jev pass as authorization | Real validation of operation+target in code | **Empirical recipe** (validation.md self-monitoring) | +| Type-checking analog | Does this planned call match the schema/operation/target? | Decomposed nouls over {request, schema, trace}; never trust a Jev pass as authorization | Real validation of operation+target in code | **Empirical recipe** (validation.md self-monitoring). Product: toolgate allow/block/review — Jev is not authorization; timeout stops (`notes.md` §55) | ## Economics & game theory diff --git a/.agents/skills/augustus/references/mixed-architecture.md b/.agents/skills/augustus/references/mixed-architecture.md index 7ee9d4c..40a20d8 100644 --- a/.agents/skills/augustus/references/mixed-architecture.md +++ b/.agents/skills/augustus/references/mixed-architecture.md @@ -184,11 +184,13 @@ not a global virtue: |---|---|---| | Drop a RAG chunk or log line | **Fail open** (keep on error) | A false drop loses evidence; a false keep costs tokens | | Compact / drop a completed tool result | **Fail closed** to keep-full (`gliner25-compaction`) | Compaction is a destructive edit of memory. Uncertain *looks* like keep-on-error from the evidence side; name the *reduction* as the act. Contrast Abide / jevgate fail-open | +| Omit a session-memory fact from the brief | **Fail open** (dump the whole ledger) (`carryforward`) | Scoring failure must not hide a rule. Constraints/corrections always return; Jev never votes on them | | Prune Bash stdout before the LLM | **Fail closed** to original (`jev-pruner`) | Dropping the log is irreversible. ≤10k / JSON-diff-whole-doc prove pass-through; archive/Jev/incomplete-score failure keeps the result. Harbor plugin-eval cannot reach Jev → cannot prune | | Skip waking a sleeping agent | **Fail open** (wake on error / unsure / no key) (`wakegate`) | Skip is the irreversible act. User-message, skip-limit, nothing-to-judge, and p in 0.2–0.5 all wake. Contrast pi-jev-approver fail-closed without a key | | Merge a red CI run | **Fail closed** on `--gate` (`latch`); reporter stays fail-open | False PASS merges a real bug. Missing key never fails Playwright; the gate is a separate step. Judge never says ignore alone | | Plain-English PR check | **Fail closed** on error / empty / low confidence (`if-ai`) | A skipped or timed-out check is not a pass. Threshold is policy, not measured correctness | | Route to a tool / start a side effect | **Fail closed** (don't call) | A wrong tool is an action | +| Execute a proposed tool call | **Fail closed** on block / timeout / guard error (`toolgate`) | Execution is the irreversible act. `review` needs authenticated human approval, not self-approval. Jev is not authorization. Distinct from ndolinschi allow/ask_human/deny *vocab* | | Actuate an observed browser control | **Fail closed** (code validates the node) | Freshness / visibility / disabled / occlusion in code; model never emits selectors (`gliner2-ultrafast`, jev-ultrafast, solari-reflex). `DONE` does not authorize "success". Cua-S1: dry-run default; `execute`/`submit` opt-in; fail-closed unknown checkbox; fill execution fails closed without token `set_value` | | Rerank a retrieved list | Fail open: keep retrieval order (`WiktorB2004/llama-index-jev`, **Empirical recipe** on BEIR nfcorpus: MiniLM 0.340 nDCG@5 → MiniLM+Jev 0.396; rerank fails open, *select* fails closed). Listwise/cross-encoder scores belong here, not on the row above. | Ranking errors are quality; selection errors are control-flow | @@ -209,6 +211,14 @@ Worked placements (2026-09-18 topic:jev hour + prior archive): residual noisy Bash after a hard ≤10k/format envelope; fail-safe original; archive for recovery (`notes.md` §53). Marketplace id still `fast-jev-output`. +- **Session ledger before the next task** — + [carryforward](https://github.com/Dharundp6/jev-carryforward): + verbatim facts; Jev scores which are still live; rules never + judged; fail-open dump (`notes.md` §55). +- **Evidence set before the generator** — + [decision-native-rag-skills](https://github.com/emergency-lee/decision-native-rag-skills): + retrieve wide → decide → evidence set → LLM. Embeddings stay + candidate generators. No universal benchmark (`notes.md` §55). - **Diff hunks before `git add`** — `ibrahemid/git-jev-stage`: one Choice per hunk (`include` / `exclude` / `mixed`); mixed and low-confidence stay unstaged; lines never split; staging is an exact patch after confirm. @@ -483,6 +493,11 @@ decision-design card. Do not clone APIs from READMEs. | S1 extract + escalate-S2 index | GLiNER spans / relations on the bulk | Graph in code; LLM only if backend loaded and low conf; query does not invent edges | s1-graphify-indexer (10–50× unfilled) | | S1 specialists + S2 coordinator | Typed {value, probability} | Coordinator / hysteresis in code | reification-labs/foreman (**description-only** Phoenix scaffold; not the super-jev loop) | | Bounded Pi supervisor | Skills / recovery / review / verify | Shadow default; never generates commands | jevons | +| Judgment as language primitive | `chance` / `pick` / `rate` (Noul / Choice / Score) | English-as-config; stub backend; fail polarity per action (`rescue nil` at save ≠ spam gate) | hunch (Ruby library, not a new language; cousin of probably-lang) | +| Decision-native RAG | Relevance / evidence / freshness / authority Nouls + Score | Evidence-set builder, conflict/temporal logic, provenance; embeddings generate candidates only | decision-native-rag-skills (no bundled harness; Hypothesis as a measured win) | +| Verbatim session recall | Noul "still live for this task?" | JSONL ledger; constraints/corrections always-keep; fail-open dump | carryforward (9×3 hint, not proof) | +| Pre-exec tool gate | allow / block / review | Permissions, arg validation, transaction limits in code; timeout stops | toolgate (72-case synthetic, not independently annotated; Jev ≠ authorization) | +| Healthcare S1 + S2 | NEWS2 remainder / med recon / inbox route | Code owns NEWS2, recon, routing; S2 blinded review | explore-typesafe-ai (synthetic FHIR; not clinically validated) | On-device / Home Assistant / mobile are newly-feasible via the economics inversion, not proven ports of every app. Named placements this hour @@ -513,7 +528,10 @@ shim, CC BY-NC; Jev still leads general text; not Archer Watch decision-model drop is **Watch**. Closed calibrated API vs open weights is a self-eval tradeoff (`research/notes.md` §18, §33, §45). When-to-use axes: `judgment-class.md`. TypeSafe remains the documented *exemplar*, -not the class monopoly. **Local contract drop-in this hour:** +not the class monopoly. **Local CUDA/PyTorch replica this hour:** +[`Mintzs/jevify`](https://github.com/Mintzs/jevify) — Qwen2.5-1.5B +Choice/Score/Noul *shape*; **uncalibrated likelihoods ≠ Noul**; no +LICENSE this pass (`notes.md` §55). **Local contract drop-in this hour:** [`us/jev-local`](https://github.com/us/jev-local) speaks `/v1/systemone`; **default scorer is a deterministic stub** until `JEVLOCAL_SCORER=hf` (`notes.md` §48). **ONNX replica path:** diff --git a/.agents/skills/augustus/references/toolbox-mapping.md b/.agents/skills/augustus/references/toolbox-mapping.md index 952f000..aceb4c2 100644 --- a/.agents/skills/augustus/references/toolbox-mapping.md +++ b/.agents/skills/augustus/references/toolbox-mapping.md @@ -71,17 +71,18 @@ component; keep the rest of the method in code. | Measurement theory: probe vs estimate | Only post-execution probes concede milestones; model estimates never do — "estimation wearing a measurement costume" is the rejection template | **Empirical recipe** (jev-mcts, pi-warden done-check) | | Experimental design: perturbation | Behavioral tests as the stats layer: candidate removal, option-order shuffle, letter-shuffle on screenshot Choice, distractor injection, boundary cases | **Contract-level** (validation.md); letter-shuffle receipt: blackwood-rlcd 0.133 vs Jev 1.13 0.587 on 300 web steps (`notes.md` §46) | | Experimental design: frozen protocol bake-off | Decision-model vs constrained LLMs vs deterministic baselines; accuracy + ECE + latency + cost + honesty; raw logs; recompute | **Empirical as Harbor/jevals practice** (DMB v2; jevals-data CC-BY-4.0 boards + JSONL; `notes.md` §49). Do not merge Banking77 across protocols | +| Experimental design: pre-registered AMBIGUOUS eval | Kill/go printed; cascade margin sensitivity; AUROC ≠ ECE; serving-path ≠ model-speed; same-day errata | **Empirical as Harbor/jevals practice (honest negative)** (jev-baselines-eval; cascade sign-flip; confidence=1.0 theater; encoder-with-labels; `notes.md` §55) | | Experimental design: extractable-from-state axis | Same question with vs without a supporting passage; citation paraphrase vs reversed-meaning | **Empirical as a boundary map** (jev-capability-atlas history suite N=3; `notes.md` §49). Qualitative, not a knowledge-breadth estimate | | Experimental design: combinatorial negative | Cell-wise Choice assembly of a grid vs extractive keep/drop | **Empirical as a negative** (ARC-AGI Direct Jev 4/400; `notes.md` §49) | | Experimental design: collab arms | `llm_autonomous` vs `scripted_plus_jev` vs `llm_plus_jev`; Wilson + McNemar; the decision model is **not** a peer arm | **Empirical as a harness shape** (jev-testbench; bake into jevals/Harbor, `notes.md` §48) | | Experimental design: Harbor on/off routing | Same coding-agent task with routing on vs off; hidden verifier; cheaper unsolved is not a saving | **Empirical as a *shape* and one-run signal** (jev-gateway-bench; `notes.md` §51). Pair CI merge-gate with Harbor + rh-guard | | Discrete math: width vs depth | Fan out in width (parallel ≈ free), pay depth linearly; two-stage only when next options depend on an earlier answer | **Empirical recipe** (fan-out: 12.2× cheaper, 10× faster) | -| Psychology: Kahneman | System 2 generates/proposes (LLM), System 1 discriminates (Jev); never the reverse. S1 keeps control; optional S2 is one-use advice and does not fly. Productized cascade: `conf ≥ τ` → S1 decides else S2 writes; routing fails open / safety fails closed; **routing accuracy unmeasured**; keyword fallback ≠ S1. S1 specialists + S2 coordinator is the same split (reification-labs/foreman is description-only Phoenix scaffold this pass — do not invent an Elixir API). Indexer: S1 GLiNER extract on the bulk, escalate LLM on the tail (10–50× unfilled) | **Empirical recipe** (mcts-agent role split; 2026-09-18 mixed-architecture discourse; jev-reflex-autonomy-lab as a control-loop shape, `notes.md` §46; [dual-process-ai](https://github.com/taro1985/dual-process-ai) as a business/life cascade, `notes.md` §49; s1-graphify-indexer, `notes.md` §51). **Route ≠ memory:** a cheap intent gate skips memory/tool *tours* on easy routes; memory still writes; complex still searches (jev-hermes, `notes.md` §48) | -| IR / cascades | Cheap relevance / irrelevance before an expensive ranker or generator | **Empirical recipe** (RAG cookbook; jevprune; git-jev-stage; LlamaIndex Jev rerank; jev-pruner stdout after size/format envelope, `notes.md` §53) | +| Psychology: Kahneman | System 2 generates/proposes (LLM), System 1 discriminates (Jev); never the reverse. S1 keeps control; optional S2 is one-use advice and does not fly. Productized cascade: `conf ≥ τ` → S1 decides else S2 writes; routing fails open / safety fails closed; **routing accuracy unmeasured**; keyword fallback ≠ S1. S1 specialists + S2 coordinator is the same split (reification-labs/foreman is description-only Phoenix scaffold this pass — do not invent an Elixir API). Indexer: S1 GLiNER extract on the bulk, escalate LLM on the tail (10–50× unfilled). Healthcare: S1 remainder after NEWS2/code, S2 blinded review (explore-typesafe-ai; synthetic; not clinically validated) | **Empirical recipe** (mcts-agent role split; 2026-09-18 mixed-architecture discourse; jev-reflex-autonomy-lab as a control-loop shape, `notes.md` §46; [dual-process-ai](https://github.com/taro1985/dual-process-ai) as a business/life cascade, `notes.md` §49; s1-graphify-indexer, `notes.md` §51; explore-typesafe-ai, `notes.md` §55). **Route ≠ memory:** a cheap intent gate skips memory/tool *tours* on easy routes; memory still writes; complex still searches (jev-hermes, `notes.md` §48) | +| IR / cascades | Cheap relevance / irrelevance before an expensive ranker or generator | **Empirical recipe** (RAG cookbook; jevprune; git-jev-stage; LlamaIndex Jev rerank; jev-pruner stdout after size/format envelope, `notes.md` §53). **Empirical as architecture** (decision-native-rag-skills retrieve-wide→decide→evidence-set; Hypothesis as a measured win, `notes.md` §55) | | IR / extractive keep-drop | Number candidates in code; model selects; copy verbatim; `redecide` thresholds on the log with no new calls. Compaction: same pointer job on tool results (Jev Noul/Score *or* GLiNER encoder). Computer-use: same pointer job on observed a11y/DOM controls (Jev *or* GLiNER2 *or* Cua-S1 option-attention). Stdout: same pointer job on Bash chunks after a hard envelope (Jev Noul) | **Empirical recipe** (testimonial-miner; jev-reviewer pointer-not-generator, `notes.md` §48; gliner25-compaction char-offset + fail-closed keep_full, `notes.md` §50; gliner2-ultrafast observe→score→act, `notes.md` §52; jev-pruner, `notes.md` §53; cua-s1 specialist form, source-only, `notes.md` §54) | | Spec / lint | Project-defined semantic rules as predicates over a diff; linter owns hard rules. AST remainder: Tree-sitter units, then typed questions; do not execute scanned code. Plain-English PR check: one condition + min-confidence; fail-closed on error | **Empirical recipe** (jev-pref contract; Abide productized path — replay 93 sessions, edit precision ~26% / turn ~73% before tune, `notes.md` §47; JevLint file-level Noul; pi-warden; snifftest unsure-band; jevscan AST∩semantic, `tenbin` owns the lint skill, `notes.md` §48; if-ai, `notes.md` §51). jev-marshal is Watch / empty this pass | | Formal methods / DST / safety | Judgment triages counterexamples, failing seeds, and named-rule conformance; proof/MC/DST stay with their tools. Alloy finder ≠ Apalache BMC ≠ Quint run. DST trio: Antithesis hypervisor / Resonate HQ Lean+oracle+SDK (durable async) / PufferLib env+seed. Noul is a sensor, not a discharged PO. Semi-formal diagrams are vocabularies, not enforcers | **Hypothesis as product**, **Contract** as ownership (matching `mappings.md` §8 and `methods-catalog.md`; worked shape pi-warden — `formal-methods.md`, `formal-semi-formal.md`) | -| Decision analysis: VOI | Gather as an enumerated act; pay iff expected decision-loss drop > cost. Wake/resume: skip the LLM turn only if the judge answers and p is low (Horvitz); user-message / skip-limit already answer without a model | **Hypothesis** as calculator (`mappings.md` §6; `mental-models.md`). wakegate 21/21 is smoke (`notes.md` §51) | +| Decision analysis: VOI | Gather as an enumerated act; pay iff expected decision-loss drop > cost. Wake/resume: skip the LLM turn only if the judge answers and p is low (Horvitz); user-message / skip-limit already answer without a model. Selective memory: score a verbatim ledger; dump on failure; never judge the rules | **Hypothesis** as calculator (`mappings.md` §6; `mental-models.md`). wakegate 21/21 is smoke (`notes.md` §51). carryforward 9×3 is a hint (`notes.md` §55) | | Signal detection | Noul as evidence variable; criterion from costs and base rate; ROC/PR on your labels | **Hypothesis** for non-SWE plots (`mappings.md` §7) | | Safety engineering: STPA | Sensor ≠ constraint; table of unsafe control actions if the sensor lies | **Contract** as ownership (`mappings.md` §8; Leveson) | | Bandits / RL | Value from observed rewards only — Jev provides none; rejected without an environment. PufferLib Ocean is a trainer contract, not a baseline | **Rejected** (standing boundary) | diff --git a/.agents/skills/augustus/references/validation.md b/.agents/skills/augustus/references/validation.md index 7329623..d0e609d 100644 --- a/.agents/skills/augustus/references/validation.md +++ b/.agents/skills/augustus/references/validation.md @@ -327,6 +327,9 @@ Rules: | Collab / computer-use product loop | [jev-testbench](https://github.com/ufx7/jev-testbench); [solari-reflex](https://github.com/hitakshiA/solari-reflex); [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast); [cua-s1](https://github.com/trycua/cua/tree/main/libs/cua-s1) | Wilson/McNemar arms; independently checked task time; `DONE` ≠ success; Cua-S1 source-only (metric names, no checkpoint scores) | | Agent routing on vs off | [jev-gateway-bench](https://github.com/vinilana/jev-gateway-bench) | Hidden perft; cost/quality; one-run signal this pass | | Command-output prune (needle/noise) | [jev-pruner](https://github.com/tamaratran/jev-pruner) | Manual `trimOutput` sweep (theirs); plugin eval cannot reach Jev (fail-safe original); Terminal-Bench paired pilot is integration, not a full bench | +| Pre-registered cascade vs nano/frontier/encoder | [jev-baselines-eval](https://github.com/ickma2311/jev-baselines-eval) | Both experiments **AMBIGUOUS**; cascade sign-flip at exact parity; confidence=1.0 theater; encoder 0.933/9ms with labels; serving-path ≠ model-speed; same-day errata ×3 | +| Healthcare S1+S2 (synthetic FHIR) | [explore-typesafe-ai](https://github.com/si618/explore-typesafe-ai) | Labels committed first; 60 requests / 403 judgments; **not clinically validated**; Claude wrote labels | +| Precision PDF (honest negative) | [databricks-jev-pdf-lab](https://github.com/laurentfabre/databricks-jev-pdf-lab) | No quality-equivalent Jev payoff; no OSS license selected | rh-guard is a reward-hack hook, a different surface from jevgate and from Abide (eval-integrity vs allowlist-remainder vs project soft @@ -378,6 +381,35 @@ Jev is down. Pair CI merge-gate (frozen JUnit artifacts × PASS/BLOCK) and rh-guard (eval-integrity). Do not copy npm/ports (`notes.md` §51). +**Pre-registered independent eval (Empirical as Harbor/jevals +*practice*, including the honest negative; 2026-09-18 ~17:48).** +[jev-baselines-eval](https://github.com/ickma2311/jev-baselines-eval) +(MIT): kill/go printed by the scripts; **both AMBIGUOUS**. CLINC150 +Jev 0.870 vs nano 0.795 vs Terra 0.915. Banking77 encoder **0.933 / +9 ms** wins. Cascade Δ +0.265 at 1pp-below-frontier; **at exact +parity the sign flips** (R_jev=1.000) because confidence is exactly +1.0 on 102/200 including 6 wrong. AUROC neither direction; **no +ECE**. Latency ~2.2× of two serving paths, not 40–200×, not +model-speed. Same-day errata three rounds. Do not copy pip +(`notes.md` §55). + +**Healthcare Harbor-shaped receipt (Empirical as that named +report; not clinical validation; 2026-09-18 ~17:48).** +[explore-typesafe-ai](https://github.com/si618/explore-typesafe-ai) +— 100 synthetic Synthea patients; labels committed before any Jev +run; 60 requests to jev-1.13.0; Claude S2 blinded review. Report: +NEWS2 alone under-triaged 10/20, NEWS2+Jev 1/20; 403 judgments / +p50 329 ms / $0.0038. Claude wrote the labels. 20 cases/scenario. +**Not clinically validated.** License not in GitHub API this pass +(`notes.md` §55). + +**Honest-negative PDF lab (Empirical as a negative; no OSS +license).** +[databricks-jev-pdf-lab](https://github.com/laurentfabre/databricks-jev-pdf-lab) +— no quality-equivalent end-to-end Jev payoff. Compact tokens +changed 26/236 recommendations. Public snapshot cannot reproduce +historical accuracy. Typed output is not truth (`notes.md` §55). + **Harbor-adjacent stdout prune (Empirical as README / evals README behavior, not a full Terminal-Bench ranking; 2026-09-18 ~17:15).** [jev-pruner](https://github.com/tamaratran/jev-pruner) ships diff --git a/CHANGELOG.md b/CHANGELOG.md index ed3294e..76fe0e5 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -268,6 +268,34 @@ Pegged against [`typesafe-ai/skills` v0.5.7](https://github.com/typesafe-ai/skil abstention, coverage, wrong actions/targets, unsafe when should abstain). Tests exercise implementation, not checkpoint quality. Watch for a `cua-s1-form-v0` artifact drop. No invented metrics. +- Hourly ~17:48 Boise fold (`research/notes.md` §55): Archer still + Watch. X MCP flap; `since_id` not advanced. Architecture notes, not + a how-to. Local CUDA/PyTorch Choice/Score/Noul replica + ([jevify](https://github.com/Mintzs/jevify); uncalibrated + likelihoods ≠ Noul; no LICENSE this pass; independent of + Distillation). Decision-native RAG + ([decision-native-rag-skills](https://github.com/emergency-lee/decision-native-rag-skills); + retrieve wide → decide → evidence set; no bundled harness; no + universal benchmark). Verbatim session ledger + scored recall + ([carryforward](https://github.com/Dharundp6/jev-carryforward); + rules never judged; fail-open dump; 9×3 hint). Judgment as a + Ruby language primitive ([hunch](https://github.com/carldaws/hunch); + English-as-config; `rescue nil` fail-open at save). Healthcare + Harbor-shaped S1+S2 + ([explore-typesafe-ai](https://github.com/si618/explore-typesafe-ai); + synthetic FHIR; not clinically validated). Pre-registered + independent eval + ([jev-baselines-eval](https://github.com/ickma2311/jev-baselines-eval); + **both AMBIGUOUS**; cascade sign-flip at exact parity; + confidence=1.0 theater; encoder-with-labels wins; serving-path ≠ + model-speed; same-day errata ×3). Student-b light delta only (HF + card unchanged). MED: toolgate (pre-exec allow/block/review; Jev + not authorization), typesafe-screening-mcp (PubMed screening aid), + databricks-jev-pdf-lab (**honest negative**; no OSS license), + yannip1234/codex-jev (extractive compression family; equal + accuracy/lower cost not established), kazuhideoki/jev-search + (recursive *file* search + fzf; **not** superagents-lab web + search). No wrapper. No invented metrics. - Effect-oriented loops (`notes.md` §28, `mappings.md` §19): Ward's ZIO client keeps Jev as the outer Choice and the handler as the effect. Not Effect.ts. GLiNER author: GLiNER2 "like jev" is GLiGuard diff --git a/README.md b/README.md index 1fcbcfc..48c9b6b 100644 --- a/README.md +++ b/README.md @@ -33,7 +33,7 @@ never launder a Noul as a proof. exemplar, not monopoly): open heads (Laya, kev, encoder DeBERTa, LoRA distill), constrained-AR (TypeAR, pcdServer), announced decision-model (Watch), open multimodal RLCD (blackwood-rlcd; not Archer), Laya ONNX port, - contract-compatible local `/v1/systemone` (stub until hf scorer; also kev pointer / von tiny SAN — not replicas), + contract-compatible local `/v1/systemone` (stub until hf scorer; also kev pointer / von tiny SAN — not replicas; **jevify** CUDA/PyTorch packed-logprob cousin — uncalibrated likelihoods ≠ Noul), GLiNER/GLiClass species (locate vs categorize vs local multi-head; GLiNER2.5 extractive compaction as a named job, not a new species; GLiNER code-graph indexer + escalate-S2, 10–50× unfilled; @@ -58,13 +58,17 @@ never launder a Noul as a proof. fail-open wake vs fail-closed merge-gate; Harbor on/off routing; hybrid local decide + remote fill; `DONE` ≠ verified success; evidence-preserving stdout prune (hard envelope then Noul); - specialist S1 computer-use (Cua-S1 form-v0; plan ≠ execute; not TypeSafe Jev) + specialist S1 computer-use (Cua-S1 form-v0; plan ≠ execute; not TypeSafe Jev); + judgment as a language primitive (hunch); decision-native RAG + (retrieve wide → decide → evidence set) - `.agents/skills/augustus/references/applied-mappings.md` — context sieve, - exact-text keep/drop (extractive / pointer-not-generator; char-offset compaction; observed a11y/DOM controls; Bash stdout prune), env triage (OpenSmoke + latch merge-gate), moderation/ranking, skill routing (route ≠ memory) + exact-text keep/drop (extractive / pointer-not-generator; char-offset compaction; observed a11y/DOM controls; Bash stdout prune; verbatim session ledger / carryforward), env triage (OpenSmoke + latch merge-gate), moderation/ranking (decision-native RAG evidence set), skill routing (route ≠ memory) - `.agents/skills/augustus/references/faq.md` — "just classification", stack replacement, Jev vs open head vs encoder vs LoRA vs constrained AR vs kev vs blackwood, wait-for-Archer, missing-other confident-wrong, soft project rules vs linter (Abide), extractive/pointer-not-generator, compaction summarize vs pointer, encoder vs Jev compaction, fail-closed keep_full, shadow-mode rollout, fail-open vs fail-closed wake vs CI gate, observe→score→act backend-agnostic, hybrid local decide + remote fill, DONE ≠ verified success, stdout prune vs session compaction, Cua-S1 vs TypeSafe Jev, plan ≠ execute / dry-run, local drop-in vs stub scorer, route ≠ memory, when-it-holds / extractable-from-state, decision-model vs constrained LLM, dual-process S1/S2, combinatorial grid ≠ extractive, GLiNER vs GLiClass vs CLIP, LLM-as-judge, in-engine vs CLI store, - hard envelope (bitrate / planner), not-another-how-to + hard envelope (bitrate / planner), not-another-how-to, + uncalibrated local likelihoods ≠ Noul, decision-native RAG, cascade + sign-flip / calibration theater, Precision PDF honest negative - `.agents/skills/augustus/references/mappings.md` — classical-method mappings with boundaries, counterexamples, acceptance tests (including Hypothesis cards §6–§19 — promote only with a test that ran) @@ -80,7 +84,10 @@ never launder a Noul as a proof. (metric names, no checkpoint scores; not TypeSafe Jev); jev-testbench collab arms; ARC-AGI Direct Jev as combinatorial-≠-extractive negative; jev-gateway-bench Harbor on/off routing one-run signal; jev-pruner Harbor needle/noise + Terminal-Bench - integration pilot, not a full bench) + integration pilot, not a full bench; jev-baselines-eval pre-registered + **AMBIGUOUS** + cascade sign-flip; explore-typesafe-ai synthetic FHIR + Harbor-shaped, not clinically validated; databricks-jev-pdf-lab honest + negative) - `.agents/skills/augustus/references/boundary-audit.md` — existing-system insertion: fit test, opportunity map, smallest boundary, red flags - `.agents/skills/augustus/scripts/evaluate_decisions.py` — offline evaluator diff --git a/docs/ecosystem.md b/docs/ecosystem.md index 446d8cb..6763f50 100644 --- a/docs/ecosystem.md +++ b/docs/ecosystem.md @@ -17,9 +17,11 @@ weekdays. Jev is the densest public corpus, not the class monopoly. - **lhemerly/mcts-agent** — batched Noul prune + Choice PUCT priors + Score/10 leaf value (no random rollouts). - **carlaiau/jev-reranking** — independent TREC DL2019 benchmark: zero-shot Jev best MAP 0.4748, nDCG@10 0.683 vs monoBERT 0.718; $0.76 per 41k pairs. - **superagents-lab/jev-search** — federated web search: Jev understands intent (query/sources/time-range), lanes fan out concurrently, Jev ranks results; merge by URL + engine agreement + rank. +- **kazuhideoki/jev-search** — recursive *file* search + fzf. Not the federated web product. `notes.md` §55. ### Languages & runtimes - **probably-lang (southpolesteve)** — a programming language whose **loop conditions are Jev feelings**: `while draft feels "like a LinkedIn influencer post" { … }`. Judgment-state recordings give deterministic replay. +- **carldaws/hunch** — Ruby library, not a new language: `almost_certain?` / `pick` / `rate` as chance/Choice/Score. English-as-config. Validations `rescue nil` fail-open at save. `notes.md` §55. - **dannote/jev** — Elixir/OTP: Jev as a peer process; answers are messages; "clause order is the routing, thresholds are guards"; network-free tests. - **jamesward/zio-typesafe-ai** — Effect-oriented (ZIO) client: Jev is the outer Choice; the handler runs the effect. Not Effect.ts and not the Jev HTTP contract. Hypothesis: `mappings.md` §19 (`notes.md` §28). @@ -46,7 +48,7 @@ Per-keystroke launchers (104ms median, sequence-tagged staleness), firehose mode - **arnabgho/rlcd-lite** — GRPO + Brier proper-scoring-rule reward → calibrated decisions; binary reward doesn't calibrate. - **stephanj/parallelConstraintDecoding** — whole JSON schema of booleans/enums in two forward passes (prefill → parallel masked fields). - **Foadsf/jev-for-engineers**, **AbdelStark/jev-benchmarks**, **BrendanH18/jev-lab** — measurement discipline and cost/latency visibility. -- **dayhaysoos/jevals** — local MIT workbench: labeled cases (Noul / Choice / Score), compare runs, WebMCP + agent skill. Empirical acceptance-test surface for Hypothesis mapping cards; complements `evaluate_decisions.py`. Not affiliated with TypeSafe. Pointer: `research/notes.md` §24. Hygiene and the Harbor substrate: `validation.md` Eval & hill-climb (`notes.md` §40). Shared bake-off exemplar: [`pngwn/open-jev-laya-bench`](https://huggingface.co/datasets/pngwn/open-jev-laya-bench) (ECE/NLL/Brier; LLM-as-judge is not the score; `notes.md` §46). Harbor-style frozen protocol vs constrained LLMs: [`nibzard/decision-model-benchmark`](https://github.com/nibzard/decision-model-benchmark) (DMB v2; jev banking 76.3% / spam 93.0% / 256+ cap; p50 264–276 ms; $0.07/1k; `notes.md` §49). Feedstock: [`Jevals/jevals-data`](https://github.com/Jevals/jevals-data) (CC-BY-4.0 boards + JSONL; recompute-from-logs; 2026-09-18 board). Boundary map (not a leaderboard): [`Zaious/jev-capability-atlas`](https://github.com/Zaious/jev-capability-atlas). Combinatorial negative: [`simonmesmith/jev-arc-agi-v1-experiment`](https://github.com/simonmesmith/jev-arc-agi-v1-experiment) (Direct Jev 4/400). +- **dayhaysoos/jevals** — local MIT workbench: labeled cases (Noul / Choice / Score), compare runs, WebMCP + agent skill. Empirical acceptance-test surface for Hypothesis mapping cards; complements `evaluate_decisions.py`. Not affiliated with TypeSafe. Pointer: `research/notes.md` §24. Hygiene and the Harbor substrate: `validation.md` Eval & hill-climb (`notes.md` §40). Shared bake-off exemplar: [`pngwn/open-jev-laya-bench`](https://huggingface.co/datasets/pngwn/open-jev-laya-bench) (ECE/NLL/Brier; LLM-as-judge is not the score; `notes.md` §46). Harbor-style frozen protocol vs constrained LLMs: [`nibzard/decision-model-benchmark`](https://github.com/nibzard/decision-model-benchmark) (DMB v2; jev banking 76.3% / spam 93.0% / 256+ cap; p50 264–276 ms; $0.07/1k; `notes.md` §49). Feedstock: [`Jevals/jevals-data`](https://github.com/Jevals/jevals-data) (CC-BY-4.0 boards + JSONL; recompute-from-logs; 2026-09-18 board). Boundary map (not a leaderboard): [`Zaious/jev-capability-atlas`](https://github.com/Zaious/jev-capability-atlas). Combinatorial negative: [`simonmesmith/jev-arc-agi-v1-experiment`](https://github.com/simonmesmith/jev-arc-agi-v1-experiment) (Direct Jev 4/400). Pre-registered independent eval (honest negative): [`ickma2311/jev-baselines-eval`](https://github.com/ickma2311/jev-baselines-eval) (both AMBIGUOUS; cascade sign-flip; `notes.md` §55). Healthcare Harbor-shaped: [`si618/explore-typesafe-ai`](https://github.com/si618/explore-typesafe-ai) (synthetic FHIR; not clinically validated). - **jeiel85/jevscope** — local-first visual debugger + JSONL regression for Choice/Score/Noul; compare two definitions; policy buckets are JevScope-derived. Sits next to jevals. Pointer: `research/notes.md` §25. ### Local / open heads & GLi\* species @@ -57,6 +59,7 @@ Per-keystroke launchers (104ms median, sequence-tagged staleness), firehose mode - **jaredpalmer/kev** — Qwen2.5-0.5B LoRA + pointer readout; Apache-2.0; Hub [`jaredpalmer/kev-0.5b`](https://huggingface.co/jaredpalmer/kev-0.5b) plus GitHub release tarball. Runnable Archer reconstruction (`POST /v1/systemone`). Public gold, not a Jev teacher. Isolation exact; ID ECE 0.065 (0.031 after T); acc 0.799 / 1,350. NOTA training must confront `"other"` as a wrong alternative (`notes.md` §45 delta). **100★** this pass (light activity delta; `notes.md` §49). Not a knowledge/frontier substitute. - **ikermoel/open-alternative-jev** — packed one-forward logprob System One on open LLMs (HF + vLLM). Apache-2.0. **Not a Jev reproduction.** RACE-H 92.9% @ 4.55 q/s on Qwen3.6-27B 8-bit; interference 6–9%. Space demo. `notes.md` §49. - **wfzyx/von** — 14 MB Needle SAN; local `POST /v1/systemone`; sub-15 ms CPU claim / ~38 ms embed. authored144 needle 52.6% — **not a calibrated Jev replica**. Distinguish from jev-local stub and kev pointer. Do not copy vs-Jev table. `notes.md` §49. +- **Mintzs/jevify** — CUDA/PyTorch packed-logprob cousin on Qwen2.5-1.5B (`ora_decision_engine`). Uncalibrated likelihoods ≠ Noul. Independent of Distillation. No LICENSE this pass. `notes.md` §55. - **BlackwoodAI/blackwood-rlcd** — open multimodal RLCD (image-text-to-text), Jev-compatible shim, CC BY-NC 4.0. Screenshot + marked candidates → Choice. Card: web acc 0.907 vs Jev 1.13 text-only 0.480; letter-shuffle 0.133 vs 0.587; ECE 0.037; ~200 ms H100. Jev still leads general text 0.850 vs 0.786. Not Archer Watch. `notes.md` §46. - **Foodoo1/Qwen3-14B-RLCD-Decision-LoRA** — decision-token QLoRA on Qwen3-14B under parallel constrained decoding. Held-out 200-case / 4-field: fraud_risk 64→95%, overall 85.2→98.8% at ~234 ms. Synthetic; not a financial product. `notes.md` §46. - **zmtomorrow/TypeAR** — constrained autoregressive decoding surface: typed fields on a pretrained open model, no retraining. Not a proper-scoring decision head. `research/notes.md` §32. @@ -191,6 +194,19 @@ Architecture notes, not a Driver / MCP catalog. `notes.md` §54. TypeSafe Jev is - **trycua/cua `libs/cua-s1`** — specialist System One computer-use research. Profile `cua-s1-form-v0` (form-oriented UI). Parent MIT; ~23.3k★ this pass. Byte-level `tinyx` encoder + option-attention head: per observed element fill (from extracted `Label: value`) / check / click / skip. Does not generate values or selectors. Code owns execution order. Plan ≠ execute; dry-run default; `execute` and `submit` independent opt-ins; submit at most one high-confidence Button labeled exactly `Submit` / `Submit Form`; fail-closed on missing checkbox state; fill execution fails closed unless the runtime advertises token-based `set_value`. Source-only: no weights, no checkpoint scores. Offline metric *names* only (accuracy, abstention, coverage, wrong actions/targets, unsafe when should abstain). Tests exercise implementation, not checkpoint quality. Same observe→score-among-candidates→code-acts *job* as jev-ultrafast / gliner2-ultrafast / solari-reflex / laya-mind2web; parallel "System One" name in CUA, not a TypeSafe contract. Watch for a `cua-s1-form-v0` artifact drop. +### Hourly ~17:48 Boise (CUDA replica, decision-native RAG, verbatim recall, Ruby primitive, FHIR Harbor, AMBIGUOUS baselines) + +Architecture notes, not a CUDA/venv / gem / mcp / uv catalog. `notes.md` §55. TypeSafe Jev is the exemplar, not the monopoly. Archer still Watch. X MCP flap; `since_id` not advanced. + +- **Mintzs/jevify** — CUDA/PyTorch parallel Choice/Score/Noul *shape* on Qwen2.5-1.5B (`ora_decision_engine`). CUDA graphs, branch kernels, literal-label scoring. **Uncalibrated likelihoods ≠ Noul.** Independent of Distillation. Default refund workflow is not a validated policy. No LICENSE this pass. +- **emergency-lee/decision-native-rag-skills** — MIT. Retrieve wide → decide → evidence set → conflict resolve → reason only over kept evidence. Provider-agnostic. No bundled harness. No universal benchmark. Core Augustus RAG mental model. +- **Dharundp6/jev-carryforward** — MIT, 1★, npm `carryforward`. Verbatim session ledger; Jev scores recall; rules never judged; fail-open dump. 9×3 hint, not proof. +- **carldaws/hunch** — MIT. Ruby `chance`/`pick`/`rate`; English-as-config; `rescue nil` fail-open at save. Cousin of probably-lang (library, not a new language). +- **si618/explore-typesafe-ai** — FHIR S1 (Jev) + Claude S2 on 100 synthetic Synthea patients. Labels first. 60 requests / 403 judgments. **Not clinically validated.** License not in API this pass. +- **ickma2311/jev-baselines-eval** — MIT. Pre-registered vs nano/frontier/encoder. **Both AMBIGUOUS.** Cascade sign-flip at exact parity; confidence=1.0 theater; encoder 0.933/9ms with labels; serving-path ≠ model-speed; same-day errata ×3. jevals/Harbor practice exemplar. +- **SargeDev/jev-gate-student-b** — light delta only; HF card unchanged (MAE 0.187 / Pearson 0.791 / 90% n=60; fail-open; teacher-copy). +- MED: **fdemir/toolgate** (pre-exec allow/block/review; Jev not authorization; 72-case synthetic); **masa-med-ai/typesafe-screening-mcp** (PubMed include/maybe/exclude; 326 hits ~17s ~$0.014; screening aid); **laurentfabre/databricks-jev-pdf-lab** (honest negative; no OSS license); **yannip1234/codex-jev** (extractive Codex compression; 185→44 estimated tokens is an integration demo; equal accuracy/lower cost not established); **kazuhideoki/jev-search** (recursive *file* search + fzf; **not** superagents-lab federated web search). + See `references/mixed-architecture.md` in the skill. Class-level family choice: `references/judgment-class.md`. Proof vs judgment (Alloy vs Apalache; DST trio Antithesis / Resonate HQ / PufferLib): diff --git a/research/archive/findings.md b/research/archive/findings.md index e15d84f..7c1ee69 100644 --- a/research/archive/findings.md +++ b/research/archive/findings.md @@ -1068,6 +1068,70 @@ One" in CUA research is a parallel name, not a TypeSafe contract; specialist; (cp) source-only drops publish metric *names*, not checkpoint scores — Watch for `cua-s1-form-v0`. +## Batch #39 (2026-09-18 ~17:48 Boise) — CUDA replica, decision-native RAG, verbatim recall, Ruby primitive, FHIR Harbor, AMBIGUOUS baselines + +Note: `research/notes.md` §55. Docs-only. Folded into PR #2. Not a +second PR. Archer still Watch. X MCP flap; `since_id` not advanced. +Archive `234740` not present locally. No invented metrics. Do not +re-fold §50–§54. Student-b species already §33 — light delta only. + +- **Mintzs/jevify (Empirical as README behavior).** Python. Created + 2026-09-18T23:41:21Z; 0★. CUDA/PyTorch parallel Choice/Score/Noul + *shape* on Qwen2.5-1.5B (`ora_decision_engine` / `ora-decision`). + CUDA graphs, branch kernels, literal-label scoring. Independent of + Distillation. **Uncalibrated model likelihoods, not measured + correctness.** Default refund workflow is not a validated policy. + Default `--answer-encoding letters`. **No LICENSE file this pass.** +- **emergency-lee/decision-native-rag-skills (Empirical as + architecture; Hypothesis as a measured win).** MIT. Created + 2026-09-18T23:29:20Z; 0★. Retrieve wide → decide → evidence set → + conflict resolve → reason only over kept evidence. Provider- + agnostic. No bundled Python harness. No universal benchmark. + Offline replay → shadow → canary → A/B. +- **Dharundp6/jev-carryforward (Empirical as README behavior).** MIT, + 1★. npm `carryforward`. Verbatim JSONL ledger; Jev scores recall; + constraints/corrections never judged; fail-open dump. 9×3 hint, not + proof. +- **carldaws/hunch (Empirical as README / example suite).** MIT. Ruby + `chance`/`pick`/`rate`; English-as-config; validations `rescue nil` + fail-open at save. Stub backend. Cousin of probably-lang (library, + not a new language). +- **si618/explore-typesafe-ai (Empirical as named report; not + clinical validation).** Created 2026-09-18T23:48:44Z; 0★; license + not in API. 100 synthetic Synthea; labels first; 60 requests / 403 + judgments to jev-1.13.0. Report: NEWS2 10/20 → 1/20 under-triage; + 98% of 143 med statuses; 7/20 inbox auto-dispatch all correct; + p50 329 ms; $0.0038. Claude wrote labels. 20 cases/scenario. +- **ickma2311/jev-baselines-eval (Empirical as Harbor/jevals + practice, including the honest negative).** MIT. Created + 2026-09-18T22:57:35Z. Pre-registered vs nano/frontier/encoder. + **Both AMBIGUOUS.** CLINC150 Jev 0.870 vs nano 0.795 vs Terra + 0.915. Banking77 encoder **0.933 / 9 ms**. Cascade Δ +0.265 at + 1pp; **sign flips at exact parity** (R_jev=1.000) because + confidence=1.0 on 102/200 incl. 6 wrong. AUROC neither direction; + no ECE. Latency ~2.2× serving-path, not 40–200×. Errata ×3. +- **SargeDev/jev-gate-student-b (light delta).** HF card unchanged: + MAE 0.187 / Pearson 0.791 / 90% n=60; ~59 ms; fail-open; teacher- + copy. +- **MED:** fdemir/toolgate (MIT; allow/block/review; Jev not + authorization; 72-case synthetic); masa-med-ai/typesafe-screening-mcp + (MIT; 326 hits ~17s ~$0.014; screening aid); laurentfabre/ + databricks-jev-pdf-lab (honest negative; no OSS license); + yannip1234/codex-jev (Apache-2.0; 185→44 estimated tokens + integration demo; equal accuracy/lower cost not established); + kazuhideoki/jev-search (file+fzf; **not** superagents-lab web + search; no LICENSE). + +Cross-repo addition: (cq) uncalibrated local likelihoods ≠ Noul; +(cr) retrieve-wide → decide → evidence set is the RAG sandwich; +(cs) verbatim ledger + scored recall, rules never judged; (ct) +judgment as a language primitive (English-as-config); (cu) +confidence=1.0 theater flips cascade sign at exact parity; (cv) +encoder-with-labels still wins; (cw) serving-path ≠ model-speed; +(cx) kazuhideoki/jev-search ≠ superagents-lab/jev-search; (cy) +toolgate product ≠ ndolinschi vocab; (cz) Precision PDF honest +negative is a result. + diff --git a/research/notes.md b/research/notes.md index 7e944e5..3a882b8 100644 --- a/research/notes.md +++ b/research/notes.md @@ -4158,3 +4158,432 @@ Not multimodal pixels. Strong **computer-use composition** signal: perception (a11y/snapshots) → specialist decide → verified act. Form specialist, not pixels-in. Weights TBD — Watch for `cua-s1-form-v0`. Archive + landscape pointer. + +## 55. CUDA replica, decision-native RAG, verbatim recall, Ruby primitive, FHIR Harbor, AMBIGUOUS baselines (2026-09-18 ~17:48 Boise) + +America/Boise ~17:48 = 23:48 UTC. Docs-only fold into open PR #2 +(`cursor/augustus-store-envelope-00b4`). Not a competing PR. Archer +27B drop still **WATCH**. X MCP namespace flap continues; `since_id` +not advanced this pass. Attached archive path +`/workspace/jev-archive/2026-09-18/234740` is **not present locally** — +receipts are live GitHub/HF + the published report site. Identity lock +vs `typesafe-ai` / `tenbin` / `decision-first` holds. No wrapper, no +Windows CUDA/venv / gem / Rails / `uv` / `mcp add` / pip how-to, no +copied ports, thresholds as class constants, or invented metrics. + +Do **not** re-fold §50 GLiNER2.5 compaction, §51 latch/wakegate/ +s1-indexer/clear-head/foreman/jevons, §52 gliner2-ultrafast, §53 +jev-pruner, or §54 Cua-S1. Do not rewrite the student-b *species* +(already §33 / applied-mappings §1 / judgment-class LoRA table). + +Backend-agnostic: this hour is **local inference replica**, +**retrieve-wide → decide → evidence set**, **verbatim ledger + +scored recall**, **judgment as a language primitive**, **Harbor- +shaped healthcare measurement**, and **honest pre-registered +AMBIGUOUS eval** (cascade sign-flip; calibration theater). TypeSafe +Jev is the documented exemplar, not the monopoly. Augustus stays +family-first. + +### HIGH + +1. **[`Mintzs/jevify`](https://github.com/Mintzs/jevify)** + (Python; created 2026-09-18T23:41:21Z; 0★ at capture; **no LICENSE + file this pass — do not invent**). Experimental CUDA/PyTorch engine + for parallel classification, yes/no, and rubric scoring with a + shared context. Default model `Qwen/Qwen2.5-1.5B-Instruct`. Python + package `ora_decision_engine`; CLI `ora-decision`. README: **this + repository is independent of the Distillation project.** Default + `workflow.json` is a four-question **refund rubric, not a validated + policy**. Default `--answer-encoding letters`; `--answer-encoding + labels` is the literal-label comparison path. Single-token answers + keep selected-head scoring; explicit label encoding evaluates + multi-token labels with a cached prompt and batched known + continuations. **These are uncalibrated model likelihoods, not + measured correctness probabilities.** Optimizations named in + README: bounded CUDA graph replay, reused answer-head weights, + optional Triton RMSNorm/SwiGLU/RoPE, short-branch kernel, grouped + similar question lengths. Historical tensors live under ignored + `outputs/`; fresh checkout skips those integration tests. + + **Four load-bearing mental models:** + + 1. **Open local packed-logprob / CUDA replica class.** Same *job* + as [`ikermoel/open-alternative-jev`](https://github.com/ikermoel/open-alternative-jev) + (packed one-forward on an open LLM; not a Jev reproduction; + `notes.md` §49): shared context, score allowed continuations, + no generated prose. Family is constrained-AR / logprob surface, + **not** trained decision-only (Laya / kev / Archer Watch) and + **not** a teacher-copy LoRA (openjev-lm / student-b). + 2. **Uncalibrated likelihood ≠ Noul.** Softmax over A/B/C or + `false`/`true` is not a proper-scoring head. Do not threshold + it as P(permit) or as calibrated abstention. Temperature / + ECE on *your* labels if you use it. + 3. **Default refund workflow is not a policy.** Configure + questions separately from inputs; do not ship their example + as production refund logic. + 4. **Do not copy the Windows CUDA/venv how-to.** + + **Placement.** Judgment-class constrained-AR / packed-logprob + cousin (`judgment-class.md` when-to-use). Pillar: search/control + (one substituted classifier step). Hole: replace-one-classifier- + step / perceive. Family: constrained-AR surface on an open LLM. + Fail polarity: **not established** — likelihoods are not + decision scores. Eval path: historical tensors not in a fresh + clone. **Empirical** as README behavior. **Hypothesis** that + CUDA graphs / branch kernels transfer to *your* GPU. Cards: + `judgment-class.md` (primary); `faq.md`; `mixed-architecture.md`. + No wrapper. + +2. **[`emergency-lee/decision-native-rag-skills`](https://github.com/emergency-lee/decision-native-rag-skills)** + (MIT; HTML+skills; created 2026-09-18T23:29:20Z; 0★). Agent Skills + for migrating, evaluating, and designing RAG around a + **decision-native evidence pipeline** rather than fixed Top-K. + Tagline: retrieve broadly → decide explicitly → build an evidence + set → resolve conflicts → reason only over what matters. + **Provider-agnostic.** Jev / OpenJev are cheap semantic operators, + not a required SDK. Three skills: `rag-migrate` / `rag-evaluate` / + `rag-design`. **No bundled Python harness** — generate the smallest + fit-for-purpose harness inside the target project. Does **not** + claim a universal benchmark. Falsifiable hypothesis: at comparable + answer quality and safety, wide retrieval + explicit evidence + decisions can improve evidence recall and cut irrelevant/redundant + context vs fixed Top-K, inside an acceptable latency/cost envelope. + Default migration *gates* (starting targets, not promises): + required-evidence recall improve-or-hold; delivered precision + improve-or-tolerance; redundancy and unresolved contradiction + reduce; unsupported claims and provenance must not regress; p95 + and cost within SLO or explicit trade-off; live user/task success + before full rollout. Eval stages: offline frozen replay → shadow + → canary → A/B. A team should not ship because an offline LLM + judge prefers it. + + **Four load-bearing mental models (core Augustus RAG):** + + 1. **Embeddings stay candidate generators.** Similarity / rerank + is not relevance, sufficiency, redundancy, conflict, time, or + authority. Stop asking one Top-K cutoff to solve all six. + 2. **Retrieve wide → decide → evidence set → LLM.** Reason only + over kept evidence. Generation is downstream of an explicit + keep set. Same family as classifying-RAG-passages cookbook + + mappings §4, promoted from "rerank the shortlist" to + **evidence-set construction**. + 3. **Skills, not a measured harness.** No bundled Python, no + private corpus, no universal number. The hypothesis is + falsifiable on *your* system. + 4. **Do not copy skill files as a product.** + + **Placement.** Retrieval + bounded semantic reranking + (`mappings.md` §4) + applied moderation/ranking + (`applied-mappings.md` §4) + mixed architecture (code owns + evidence-set / conflict / provenance; model scores candidates). + Pillar: VOI (expand only if the evidence set is insufficient) + + MCDA (relevance / freshness / authority as named features). + Hole: rank / sieve / gather. Family: closed decision API *or* + any cheap semantic operator (backend-agnostic). Fail-open on + drop of a candidate (false drop loses evidence). Eval path: + rag-evaluate four stages; **Hypothesis** until a target-system + test runs. **Empirical** as README architecture. Cards: + `mappings.md` §4 (primary); `applied-mappings.md` §4; + `mixed-architecture.md`; `mental-models.md`; `faq.md`. No + wrapper. + +3. **[`Dharundp6/jev-carryforward`](https://github.com/Dharundp6/jev-carryforward)** + (MIT; TypeScript; created 2026-09-18T23:04:58Z; 1★; npm + `carryforward`). MCP session memory: `record` saves a fact + verbatim the moment it happens; `recall` scores non-rule entries + with Jev against the current task. **Nothing is summarised. + Nothing is deleted.** Kinds: `constraint` / `correction` / + `decision` / `measurement` / `thread`. Constraints and + corrections **always return in full** — Jev never votes on a + rule you set. Decisions / measurements / threads need a `ref`. + Provenance: `measured` / `decided` / `told` / `inferred`. + Thresholds (theirs, exported constants, not class constants): + p ≥ 0.60 full entry; 0.30–0.60 one line; below omit from the + brief (still on disk). Fail-open: no key / no task / scorer down + / rate-limited → **whole list** plus a line saying why. + Asker is swappable (`Asker.ask(state, questions)`). JSONL + append-only at `~/.carryforward/.jsonl`. Nine entries × + three tasks is a **hint, not proof**; no accuracy claim until a + proper test. Tests use a fake scorer; never the network. + + **Four load-bearing mental models:** + + 1. **Verbatim ledger, scored recall.** Pointer-not-generator on + *memory*: the model never rewrites the note. Sorting happens + at read, not write. Family with pi-jev-compaction / + testimonial-miner (select, copy, do not summarize). + 2. **Rules never judged.** Constraints/corrections are always- + keep in code. Same sandwich as jevgate's allowlist *proves* + / remainder judged — here the remainder is "is this still + live for the task?" + 3. **VOI / selective memory.** Pay for a scored brief iff it + beats dumping the whole ledger (fail-open dump is the safe + default). Horvitz: skip the irrelevant, never skip the rule. + 4. **Do not copy `claude mcp add` / SessionStart hooks.** + + **Placement.** Context sieve (`applied-mappings.md` §1) + VOI + (`mappings.md` §6). Pillar: VOI + selective classification. + Hole: sieve / gather. Family: closed decision API (Noul + "still live?"). Fail-open on scoring failure. Eval path: none + published (9×3 hint). **Empirical** as README behavior. + **Hypothesis** that scored recall beats dump-or-summary on + *your* session. Cards: `applied-mappings.md` §1 (primary); + `mappings.md` §6; `mixed-architecture.md`; + `agent-self-assessment.md`. No wrapper. + +4. **[`carldaws/hunch`](https://github.com/carldaws/hunch)** + (MIT; Ruby; created 2026-09-18T23:08:32Z; 0★). Probabilistic + control flow for Ruby: `if` / `case` / `<=>` for facts; Hunch + for judgment calls. English is the configuration. `chance` → + Noul (`almost_certain?` / `likely?` / `probable?` / named + levels); `pick` → Choice; `rate` → Score. Batch `Hunch.decide` + over one `given:`. Rails examples (theirs): validations, inbound + email routing, error triage, job retries, enum coercion, + comment moderation. **`rescue nil` on validations is deliberate + fail-open at save** — fail closed instead where it matters + (spam gate). Stub backend for tests. Same interface ≠ same + guarantees for a future LLM backend (Jev calibrated + typed + + milliseconds; an LLM backend is estimates, slower, dearer). + Cousin of [`southpolesteve/probably`](https://github.com/southpolesteve/probably) + (language whose *loop conditions* are Jev feelings) — Hunch is + a library in Ruby, not a new language. + + **Four load-bearing mental models:** + + 1. **Judgment as a language primitive.** `almost_certain?` / + `pick` / `rate` are control-flow, not a prompt. English-as- + config. Same instinct as probably-lang, one layer down. + 2. **Fail polarity is per action.** Validation fail-open + (`rescue nil`); spam *gate* should fail closed. Name the + act, not the slogan. + 3. **Stub is a backend.** Tests never need the network. + 4. **Do not copy gem / Rails how-to.** + + **Placement.** Decision circuits (`mappings.md` §3) + mixed + architecture. Pillar: EU / selective classification. Hole: + gate / route / replace-one-classifier-step. Family: closed + decision API. Fail polarity per call-site. Eval path: example + app tests against live model (theirs); stub for CI. + **Empirical** as README / example suite. Cards: + `mappings.md` §3 (primary); `mixed-architecture.md`; `faq.md`. + No wrapper. + +5. **[`si618/explore-typesafe-ai`](https://github.com/si618/explore-typesafe-ai)** + (Python; created 2026-09-18T23:48:44Z; 0★; **license not in + GitHub API this pass — do not invent**). FHIR clinical System + One (Jev) + Claude System Two on **100 synthetic Synthea** + patients. Report: + [si618.github.io/explore-typesafe-ai](https://si618.github.io/explore-typesafe-ai). + Three scenarios: NEWS2 huddle (Noul/Score/Choice); discharge + med recon (Choice fan-out, Noul, Score); post-discharge inbox + (Choice/Score/Noul + confidence gate). Labels committed + **before** any Jev run. 60 requests to `jev-1.13.0`. **Not + clinically validated.** Claude wrote reference labels, not + clinicians. 20 cases per scenario — wide uncertainty. + + **Report headline (theirs; not re-run):** NEWS2 alone under- + triaged 10/20; NEWS2 + Jev under-triaged 1/20; new-confusion + Noul 20/20. Discharge: 98% of 143 medication statuses; allergy + check 100%; duplicate/interaction checks weak (multi-hop) and + mostly escalate. Inbox: 7/20 auto-dispatched, all correctly; + every misroute caught by the confidence gate; prompt injection + did not steer routing. 65 of 403 judgments (16%) escalated to + blinded Claude Sonnet 5. Cost/speed: 403 judgments / 60 + requests; p50 329 ms/request; **$0.0038** total. + + **Four load-bearing mental models:** + + 1. **Harbor-shaped healthcare measurement.** Frozen synthetic + cohort, labels first, independent S2 review packet, code + owns NEWS2 / recon / routing. Capability demonstration, + not clinical safety evidence. + 2. **Code stays in charge.** Jev supplies inputs code cannot + compute (note meaning, brand names, new vs baseline + confusion). Thresholds re-policy without a new prompt. + 3. **Multi-hop over a list is still jagged.** Duplicate / + interaction checks escalate — decompose or don't ask. + 4. **Do not copy `uv` how-to.** Not a medical device. + + **Placement.** Validation Harbor (`validation.md`) + mixed + architecture (S1 decide / S2 review / code policy). Pillar: + SDT (under-triage cost >> over-triage) + Leveson + (sensor ≠ constraint). Hole: triage / gate / perceive. + Family: closed decision API. Fail-closed on actuation + (escalate / hold); **not clinically validated**. Eval path: + published report + committed labels. **Empirical** as that + named report. **Hypothesis** that the same split transfers + to real FHIR. Cards: `validation.md` (primary); + `mental-models.md`; `mixed-architecture.md`. No wrapper. + +6. **[`ickma2311/jev-baselines-eval`](https://github.com/ickma2311/jev-baselines-eval)** + (MIT; Python; created 2026-09-18T22:57:35Z; 0★). Pre-registered + independent eval of TypeSafe Jev vs nano-class LLM + (`gpt-5.4-nano`), frontier (`GPT-5.6 Terra`), and a supervised + encoder (`bge-small-en-v1.5` + logistic regression, 10,003 + Banking77 train, 9 ms laptop). Not affiliated; ~$1 API paid by + the author. **Both experiments returned AMBIGUOUS.** Same-day + errata, **three rounds** (calibration language, B0 escalation + numbers, encoder-vs-frontier arithmetic, missing cross-fit + accuracies, **threshold-margin sensitivity that flips the sign + of the headline cascade**, wrong parity explanation, latency + framing). Reviews in `reviews/` (GPT-6 Astra via Codex CLI); + author verified every quantitative finding from `results/`. + + **Numbers (theirs; recomputable from published JSONL):** + + - CLINC150 zero-shot n=200: Jev **0.870** vs nano **0.795** + (paired +7.5pp, 95% CI [+3.0, +12.5]) vs Terra **0.915**. + - Banking77 paired n=208: encoder **0.933** [0.899, 0.966] / + **9 ms** wins; vs Jev 0.832, paired encoder **+10.1pp + [+5.3, +15.4]**; encoder vs Terra **+5.8pp [+2.4, +9.6]**. + B0's own pre-registered verdict was also AMBIGUOUS. + - Cascade (B1 primary): at A_Terra − **1pp** (0.905), R_jev + **0.220** vs R_nano 0.485, Δ **+0.265**, CI [−0.530, +0.595] + → **AMBIGUOUS**. At **exact parity** R_jev **1.000** vs + nano 0.730 (Δ **−0.270**) — **sign flips**. Mechanism: Jev + confidence **exactly 1.0 on 102/200 items, 6 of which are + wrong** (only 1 of those 6 is one Terra gets right). No + threshold that *keeps any Jev answer* reaches parity + (t=1.0 → 0.490 escalation at 0.910; t=1.01 escalates + everything). Read the 1pp row as "cheap to get *close*", + never "at equal accuracy". + - Error-ranking AUROC: CLINC150 Jev 0.734 vs nano 0.816 + (paired CI includes zero); Banking77 reverse. **Neither + direction established.** This is *error ranking*, **not + ECE**. No ECE/reliability diagram in this report. + - Latency: recorded median call duration **~2.2×** shorter + for the Jev *configuration* than nano on the same 30 items + (0.42 s vs 0.92 s) — **not** the vendor 40–200×, **not** + isolated model inference speed, **serving-path not + model-speed**. Throughput under Vercel free-tier rate + limit is a different number (200-item run ~3.5 h). Same- + gateway control named and **not run**. + - Deviations disclosed: B0 n 300→208 (Terra 0.875 retained + vs 0.804 omitted); encoder added after B0 pre-reg; B1 run + despite B0's "ambiguous would not expand" stopping rule. + + **Four load-bearing mental models (jevals / Harbor practice + exemplar this hour):** + + 1. **Honest negative + pre-registration.** Kill/go printed + by the analysis scripts, including the one that failed. + AMBIGUOUS is a result. + 2. **Calibration theater.** Confidence = 1.0 on 102/200 + including 6 wrong. AUROC is not ECE. Do not say "better + calibrated" from error-ranking. A cascade at 1pp-below- + frontier is not a cascade at equal accuracy. + 3. **Encoder with labels still wins.** 10k labeled Banking77 + → 0.933 / 9 ms / $0. Test that baseline before paying + per call. Different information regime, not a like-for- + like model bake-off. + 4. **Serving-path ≠ model-speed.** 2.2× is two client-and- + service configurations. Do not invent 40–200× from this + repo. Do not copy pip. + + **Placement.** Validation Harbor / jevals (`validation.md` + primary). Pillar: SDT + calibration. Hole: measure / hill- + climb. Family: bake-off, not a product. **Empirical** as that + named report (AMBIGUOUS + errata). Cards: `validation.md`; + `faq.md`; `methods-catalog.md`; `toolbox-mapping.md`. No + wrapper. + +7. **[`SargeDev/jev-gate-student-b`](https://huggingface.co/SargeDev/jev-gate-student-b) + + [`jev-distill-corpus`](https://huggingface.co/datasets/SargeDev/jev-distill-corpus)** + — **light delta only.** HF card unchanged this pass vs §33: + LoRA r=16 α=32 on Qwen2.5-0.5B; P(relevant) from yes/no + logits; held-out n=60 MAE **0.187** / Pearson **0.791** / + agreement **90.0%** vs vanilla 0.536 / −0.067 / 38.3%; + ~59 ms RTX 3060; gate at 0.5; **fail-open on errors**; + teacher-copy, not independent gold; 148,160-row corpus. + Distillation of a System One *memory gate* remains the + species. Do not rewrite the LoRA table. Do not copy the + usage snippet. + +### STRONG MED (brief) + +- **[`fdemir/toolgate`](https://github.com/fdemir/toolgate)** + (MIT; TypeScript; created 2026-09-18T23:21:59Z; 0★). Pre-exec + tool gate: `allow` / `block` / `review` before execution. + Guard error or timeout **stops** (error distinct from a model + decision) — fail-closed on the *execution* act. Jev is a + probabilistic check, **not authorization**; keep permissions, + argument validation, and transaction limits. 72-case synthetic + starter dataset **not independently human-annotated**. Demos + prove execution wiring, not model accuracy. **Product**, not + the ndolinschi *vocabulary* already in + `agent-self-assessment.md` (that family used allow / ask_human + / deny). This repo's labels are allow / block / review. + `onReview` must obtain authenticated human approval, not ask + the agent to approve itself. Do not copy pnpm how-to. + Placement: `mappings.md` §18 + agent pre-action gate. + +- **[`masa-med-ai/typesafe-screening-mcp`](https://github.com/masa-med-ai/typesafe-screening-mcp)** + (MIT; Python; created 2026-09-18T23:45:47Z; 0★). PubMed + title/abstract screening MCP: `include` / `maybe` / `exclude`. + One Jev request per article (match Noul + relevance Score + + criterion Nouls); decision rule in code, sensitivity-first + (unmet inclusion never auto-excludes). Abstracts never enter + the LLM conversation. One real run (theirs): **326 hits ~ + 17 s ~ $0.014**. Thresholds **not calibrated** on labelled + data. Screening aid, not a systematic-review replacement. + Do not send patient/confidential text. Do not copy `uv` / + keychain how-to. + +- **[`laurentfabre/databricks-jev-pdf-lab`](https://github.com/laurentfabre/databricks-jev-pdf-lab)** + (Python; created 2026-09-18T23:47:15Z; 0★; **no OSS license + selected — public visibility is not a license**). Honest + negative: **no quality-equivalent, end-to-end Jev payoff + demonstrated** for Precision-Mode PDF extraction. Compact + metadata requests: 32.48% fewer input tokens but **26/236 + recommendations changed** (not equivalent-policy). Bounded + verifier 3/5 flags / 0/3 false alarms on four correlated + inspected cases — not calibrated acceptance. Selective-parse + rehearsal retained all 236 pages. Typed output is not truth. + Public snapshot cannot independently reproduce historical + accuracy. Do not treat this as a production router. + +- **[`yannip1234/codex-jev`](https://github.com/yannip1234/codex-jev)** + (Apache-2.0 via upstream Codex; Rust/Swift; created + 2026-09-18T23:49:00Z; 0★). Codex extractive compression + family: custom engine + desktop bridge + native macOS client. + Kept passages copied from source; API failure / timeout / + uncertainty / insufficient savings **preserve original**. + Manually sent official-app message reduced **~185 → 44 + estimated tokens**, `COMPACTION_OK` — **integration demo**, + not complete desktop compatibility. **Equal task accuracy and + lower total cost have not been established.** Savings are + estimated, not tokenizer-exact billing. Family with + fast-jev-compaction / jev-pruner / gliner25-compaction + (pointer, not summarizer). Do not copy Xcode/Rust build. + +- **[`kazuhideoki/jev-search`](https://github.com/kazuhideoki/jev-search)** + (Python; created 2026-09-18T23:43:52Z; 0★; **no LICENSE file + this pass**). Recursive semantic **file** search + fzf: + ripgrep enumerate → Jev match probability → fzf select. + **Not** [`superagents-lab/jev-search`](https://github.com/superagents-lab/jev-search) + (federated *web* search; Jev as query-understanding head and + result-ranking tail). File score is max over overlapping + chunks — **not** a calibrated whole-file probability; long + files may be favored. Failures/unevaluated chunks are not + treated as 0%. `--dry-run` needs no key. Do not copy `.env` + how-to. + +### Omni / Jev-omni / Archer + +Still **WATCH**. No Hub weights. X MCP flap; `since_id` not +advanced. This hour does not wait. Local CUDA replica (jevify) +is an open *inference class*, not that drop. Healthcare Harbor +(explore-typesafe-ai) is synthetic FHIR, not AU residency +weights. + +### Cross-links + +Cards: `judgment-class.md` (jevify uncalibrated replica; +student-b light); `applied-mappings.md` §1 (carryforward), +§4 (decision-native RAG); `mappings.md` §3 (hunch), §4 (RAG + +file-search vs web-search), §6 (carryforward VOI), §18 +(toolgate); `mixed-architecture.md` (fail table + gallery); +`validation.md` (jev-baselines-eval AMBIGUOUS + errata; +explore-typesafe-ai; pdf-lab negative); `faq.md`; +`mental-models.md`; `methods-catalog.md`; `toolbox-mapping.md`; +`agent-self-assessment.md`. No wrapper. diff --git a/research/refresh-log.md b/research/refresh-log.md index 2d96899..71b5745 100644 --- a/research/refresh-log.md +++ b/research/refresh-log.md @@ -606,6 +606,35 @@ README. - notes.md §54; sources.json; findings.md batch #38. No wrapper. +## 2026-09-18 23:54 UTC — CUDA replica / decision-native RAG / verbatim recall / Ruby primitive / FHIR Harbor / AMBIGUOUS baselines (~17:48 Boise) + +- Folded into open PR #2 (`cursor/augustus-store-envelope-00b4`). + Docs-only. Not a competing PR. Archer 27B drop still **WATCH**. + X MCP namespace flap; `since_id` **not** advanced. Archive + `/workspace/jev-archive/2026-09-18/234740` not present locally. + No invented metrics. No wrapper. Do not re-fold §50–§54. +- HIGH: [Mintzs/jevify](https://github.com/Mintzs/jevify) CUDA/PyTorch + Choice/Score/Noul *shape* on Qwen2.5-1.5B; uncalibrated likelihoods + ≠ Noul; no LICENSE. [decision-native-rag-skills](https://github.com/emergency-lee/decision-native-rag-skills) + retrieve-wide → decide → evidence set; no harness; no universal + benchmark. [jev-carryforward](https://github.com/Dharundp6/jev-carryforward) + verbatim ledger + scored recall; rules never judged; 9×3 hint. + [hunch](https://github.com/carldaws/hunch) Ruby language primitive. + [explore-typesafe-ai](https://github.com/si618/explore-typesafe-ai) + synthetic FHIR Harbor-shaped; not clinically validated. + [jev-baselines-eval](https://github.com/ickma2311/jev-baselines-eval) + **both AMBIGUOUS**; cascade sign-flip; confidence=1.0 theater; + encoder-with-labels; serving-path ≠ model-speed; errata ×3. + Student-b light delta only (HF card unchanged). +- MED: toolgate, typesafe-screening-mcp, databricks-jev-pdf-lab + (honest negative, no OSS license), yannip1234/codex-jev, + kazuhideoki/jev-search (**not** superagents-lab web search). +- Cards: SKILL.md, judgment-class, applied-mappings §1/§4, + mappings §3/§4/§6/§18, mixed-architecture, validation, faq, + mental-models, methods-catalog, toolbox, agent-self-assessment, + ecosystem, CHANGELOG, README. +- notes.md §55; sources.json; findings.md batch #39. No wrapper. + diff --git a/research/sources.json b/research/sources.json index ba20aa4..4a4592c 100644 --- a/research/sources.json +++ b/research/sources.json @@ -1,6 +1,6 @@ { "refresh_cadence": "hourly", - "retrieved": "2026-09-18T23:21Z", + "retrieved": "2026-09-18T23:54Z", "sources": [ { "kind": "docs", @@ -1604,6 +1604,78 @@ "title": "trycua/cua libs/cua-s1", "url": "https://github.com/trycua/cua/tree/main/libs/cua-s1", "note": "Parent MIT; ~23.3k stars this pass (23339 live; updated 2026-09-18T23:21:08Z). Specialist System One computer-use research (form-v0 profile). Source-only: no weights, no checkpoint scores. tinyx byte encoder + option-attention head chooses fill/check/click/skip per observed element; does not generate values or selectors. Plan \u2260 execute; dry-run default; execute/submit independent opt-ins; fail-closed unknown checkbox / fill without advertised token set_value. Not TypeSafe Jev \u2014 parallel System One naming in CUA. Same observe\u2192score-among-candidates\u2192code-acts job as jev-ultrafast / gliner2-ultrafast / solari-reflex / laya-mind2web. Watch for cua-s1-form-v0 artifact. notes.md \u00a754." + }, + { + "kind": "github", + "title": "Mintzs/jevify", + "url": "https://github.com/Mintzs/jevify", + "note": "Python. Created 2026-09-18T23:41:21Z; 0 stars. CUDA/PyTorch parallel Choice/Score/Noul shape on Qwen2.5-1.5B (ora_decision_engine / ora-decision). CUDA graphs, branch kernels, literal-label scoring. Uncalibrated model likelihoods, not measured correctness. Independent of Distillation. Default refund workflow is not a validated policy. Default --answer-encoding letters. No LICENSE file this pass. notes.md \u00a755." + }, + { + "kind": "github", + "title": "emergency-lee/decision-native-rag-skills", + "url": "https://github.com/emergency-lee/decision-native-rag-skills", + "note": "MIT. Created 2026-09-18T23:29:20Z; 0 stars. Retrieve wide \u2192 decide \u2192 evidence set \u2192 conflict resolve \u2192 reason only over kept evidence. Provider-agnostic. No bundled Python harness. No universal benchmark. Offline replay \u2192 shadow \u2192 canary \u2192 A/B. notes.md \u00a755." + }, + { + "kind": "github", + "title": "Dharundp6/jev-carryforward", + "url": "https://github.com/Dharundp6/jev-carryforward", + "note": "MIT. Created 2026-09-18T23:04:58Z; 1 star. npm carryforward. Verbatim JSONL session ledger; Jev scores recall; constraints/corrections never judged; fail-open dump. 9\u00d73 hint, not proof. notes.md \u00a755." + }, + { + "kind": "github", + "title": "carldaws/hunch", + "url": "https://github.com/carldaws/hunch", + "note": "MIT. Created 2026-09-18T23:08:32Z; 0 stars. Ruby chance/pick/rate as Noul/Choice/Score; English-as-config; validations rescue nil fail-open at save. Stub backend. Cousin of probably-lang (library, not a new language). notes.md \u00a755." + }, + { + "kind": "github", + "title": "si618/explore-typesafe-ai", + "url": "https://github.com/si618/explore-typesafe-ai", + "note": "Created 2026-09-18T23:48:44Z; 0 stars; license not in GitHub API. FHIR S1 (Jev) + Claude S2 on 100 synthetic Synthea patients. Labels committed before any Jev run. 60 requests / 403 judgments to jev-1.13.0. Report: https://si618.github.io/explore-typesafe-ai. Not clinically validated. notes.md \u00a755." + }, + { + "kind": "web", + "title": "explore-typesafe-ai report", + "url": "https://si618.github.io/explore-typesafe-ai", + "note": "Headline (theirs): NEWS2 10/20 \u2192 1/20 under-triage; 98% of 143 med statuses; 7/20 inbox auto-dispatch all correct; 65/403 (16%) S2 escalate; p50 329 ms; $0.0038. Claude wrote labels. 20 cases/scenario. notes.md \u00a755." + }, + { + "kind": "github", + "title": "ickma2311/jev-baselines-eval", + "url": "https://github.com/ickma2311/jev-baselines-eval", + "note": "MIT. Created 2026-09-18T22:57:35Z; 0 stars. Pre-registered independent eval vs nano/frontier/encoder. Both experiments AMBIGUOUS. CLINC150 Jev 0.870 vs nano 0.795 vs Terra 0.915. Banking77 encoder 0.933/9ms. Cascade \u0394 +0.265 at 1pp; sign flips at exact parity (R_jev=1.000) because confidence=1.0 on 102/200 incl. 6 wrong. AUROC neither direction; no ECE. Latency ~2.2\u00d7 serving-path, not 40\u2013200\u00d7. Same-day errata three rounds. notes.md \u00a755." + }, + { + "kind": "github", + "title": "fdemir/toolgate", + "url": "https://github.com/fdemir/toolgate", + "note": "MIT. Created 2026-09-18T23:21:59Z; 0 stars. Pre-exec allow/block/review. Guard error/timeout stops. Jev is not authorization. 72-case synthetic, not independently annotated. Distinct from ndolinschi allow/ask_human/deny vocab. notes.md \u00a755." + }, + { + "kind": "github", + "title": "masa-med-ai/typesafe-screening-mcp", + "url": "https://github.com/masa-med-ai/typesafe-screening-mcp", + "note": "MIT. Created 2026-09-18T23:45:47Z; 0 stars. PubMed title/abstract screening MCP: include/maybe/exclude. One run (theirs): 326 hits ~17s ~$0.014. Thresholds not calibrated. Screening aid, not an SR replacement. notes.md \u00a755." + }, + { + "kind": "github", + "title": "laurentfabre/databricks-jev-pdf-lab", + "url": "https://github.com/laurentfabre/databricks-jev-pdf-lab", + "note": "Created 2026-09-18T23:47:15Z; 0 stars. No OSS license selected. Honest negative: no quality-equivalent end-to-end Jev payoff for Precision PDF. Compact tokens changed 26/236 recommendations. Typed output is not truth. notes.md \u00a755." + }, + { + "kind": "github", + "title": "yannip1234/codex-jev", + "url": "https://github.com/yannip1234/codex-jev", + "note": "Apache-2.0 (upstream Codex). Created 2026-09-18T23:49:00Z; 0 stars. Extractive Codex compression family. COMPACTION_OK 185\u219244 estimated tokens is an integration demo. Equal accuracy/lower cost not established. notes.md \u00a755." + }, + { + "kind": "github", + "title": "kazuhideoki/jev-search", + "url": "https://github.com/kazuhideoki/jev-search", + "note": "Created 2026-09-18T23:43:52Z; 0 stars; no LICENSE file this pass. Recursive semantic file search + fzf. Not superagents-lab/jev-search (federated web). Max-over-chunks is not a calibrated whole-file probability. notes.md \u00a755." } ] } From 59aee41a425266d1fe35c6824dce9a33c614bb9a Mon Sep 17 00:00:00 2001 From: Cursor Agent Date: Sat, 19 Sep 2026 00:47:20 +0000 Subject: [PATCH 14/43] Fold classify-first MCP and living applied-mappings atlas MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Docs-only: jev-sift classify-first (retrieve-wide → decide → evidence-set on agent I/O) and jevable.com class patterns (not a 342-title dump). No invented metrics; Archer still Watch. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com> --- .agents/skills/augustus/SKILL.md | 14 +- .../references/agent-self-assessment.md | 13 +- .../augustus/references/applied-mappings.md | 40 +++ .agents/skills/augustus/references/faq.md | 38 ++- .../skills/augustus/references/mappings.md | 40 ++- .../augustus/references/mental-models.md | 6 +- .../augustus/references/methods-catalog.md | 10 +- .../augustus/references/mixed-architecture.md | 19 +- .../augustus/references/question-design.md | 2 +- .../augustus/references/toolbox-mapping.md | 6 +- CHANGELOG.md | 13 + README.md | 7 +- docs/ecosystem.md | 10 +- research/archive/findings.md | 48 ++++ research/notes.md | 262 ++++++++++++++++++ research/refresh-log.md | 22 ++ research/sources.json | 14 +- 17 files changed, 536 insertions(+), 28 deletions(-) diff --git a/.agents/skills/augustus/SKILL.md b/.agents/skills/augustus/SKILL.md index b8c3ac2..10c196b 100644 --- a/.agents/skills/augustus/SKILL.md +++ b/.agents/skills/augustus/SKILL.md @@ -1,6 +1,6 @@ --- name: augustus -description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill, open multimodal RLCD / blackwood), announced open decision-model (Watch — still not landed), constrained-AR (TypeAR, pcdServer, decision-token LoRA), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"wait for Archer vs open multimodal RLCD\", \"screenshot/DOM candidates → typed Choice\", \"eval path\", \"jevals\", \"Harbor taskset\", \"shared bake-off ECE/NLL/Brier\", \"LLM-as-judge is not the System One score\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge (allowlist *proves*; fail-open cannot block), \"missing other → confident wrong Choice\", \"lint the Jev request\", \"training confronts Choice other / none-of-the-above\", \"S1 reflex keeps control / optional S2 one-use advice\", \"soft AGENTS.md rules vs linter\" (Abide / jev-pref), \"edit-phase vs turn-phase observation window\", banded confidence fail-open preference lint, \"extractive selection / pointer-not-generator\", \"encoder GLiNER compaction vs Jev Score compaction (same job; pointer not summarizer)\", \"fail-closed keep_full under mutation envelope\", \"CI flaky-vs-real merge gate\", \"fail-open VOI wake/resume (Horvitz)\", \"claim/evidence Stop integrity\", \"S1 extract + escalate-S2 indexer\", \"Harbor on/off routing\", \"policy-as-judgment PR marshal\", \"shadow-mode compaction rollout\", \"Jev Ultrafast vs GLiNER Ultrafast (observe-score-act backend-agnostic)\", \"hybrid local decide + remote fill\", \"DONE ≠ verified success\", \"observed a11y/DOM candidates vs screenshot multimodal\", \"evidence-preserving stdout prune (not summarize)\", \"hard token/format envelope then soft Noul\", \"fail-safe keep original on prune failure\", \"stdout prune vs session compaction\", \"specialist S1 computer-use (Cua-S1 form-v0; not TypeSafe Jev)\", \"plan ≠ execute / dry-run default\", \"observed-element option head (fill/check/click/skip)\", \"local /v1/systemone drop-in (stub until hf scorer)\", \"dataframe-native semantic columns\", \"route≠memory\", \"advisory sidecar receipts\", \"structure induction over bags\", \"AST ∩ semantic lint\", \"extractable-from-state / retrieve first\", \"decision-model vs constrained-LLM bake-off\", \"dual-process S1 decide / S2 generate\", \"combinatorial grid ≠ extractive\", TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, \"uncalibrated local likelihoods vs Noul / CUDA replica\", \"decision-native RAG retrieve wide then decide then evidence set\", \"verbatim session ledger / scored recall\", \"judgment as language primitive / English-as-config\", \"pre-registered AMBIGUOUS eval / cascade sign-flip\", \"healthcare Harbor-shaped S1+S2\", \"pre-exec tool gate allow/block/review\", or \"formally verify with Jev\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." +description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill, open multimodal RLCD / blackwood), announced open decision-model (Watch — still not landed), constrained-AR (TypeAR, pcdServer, decision-token LoRA), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"wait for Archer vs open multimodal RLCD\", \"screenshot/DOM candidates → typed Choice\", \"eval path\", \"jevals\", \"Harbor taskset\", \"shared bake-off ECE/NLL/Brier\", \"LLM-as-judge is not the System One score\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge (allowlist *proves*; fail-open cannot block), \"missing other → confident wrong Choice\", \"lint the Jev request\", \"training confronts Choice other / none-of-the-above\", \"S1 reflex keeps control / optional S2 one-use advice\", \"soft AGENTS.md rules vs linter\" (Abide / jev-pref), \"edit-phase vs turn-phase observation window\", banded confidence fail-open preference lint, \"extractive selection / pointer-not-generator\", \"encoder GLiNER compaction vs Jev Score compaction (same job; pointer not summarizer)\", \"fail-closed keep_full under mutation envelope\", \"CI flaky-vs-real merge gate\", \"fail-open VOI wake/resume (Horvitz)\", \"claim/evidence Stop integrity\", \"S1 extract + escalate-S2 indexer\", \"Harbor on/off routing\", \"policy-as-judgment PR marshal\", \"shadow-mode compaction rollout\", \"Jev Ultrafast vs GLiNER Ultrafast (observe-score-act backend-agnostic)\", \"hybrid local decide + remote fill\", \"DONE ≠ verified success\", \"observed a11y/DOM candidates vs screenshot multimodal\", \"evidence-preserving stdout prune (not summarize)\", \"hard token/format envelope then soft Noul\", \"fail-safe keep original on prune failure\", \"stdout prune vs session compaction\", \"specialist S1 computer-use (Cua-S1 form-v0; not TypeSafe Jev)\", \"plan ≠ execute / dry-run default\", \"observed-element option head (fill/check/click/skip)\", \"local /v1/systemone drop-in (stub until hf scorer)\", \"dataframe-native semantic columns\", \"route≠memory\", \"advisory sidecar receipts\", \"structure induction over bags\", \"AST ∩ semantic lint\", \"extractable-from-state / retrieve first\", \"decision-model vs constrained-LLM bake-off\", \"dual-process S1 decide / S2 generate\", \"combinatorial grid ≠ extractive\", TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, \"uncalibrated local likelihoods vs Noul / CUDA replica\", \"decision-native RAG retrieve wide then decide then evidence set\", \"classify-first MCP / read selectively\", \"living applied-mappings atlas / class patterns not a 342 hit list\", \"draft-gate silence as safer / heartbeat\", \"robotics text-state vs pixels\", \"verbatim session ledger / scored recall\", \"judgment as language primitive / English-as-config\", \"pre-registered AMBIGUOUS eval / cascade sign-flip\", \"healthcare Harbor-shaped S1+S2\", \"pre-exec tool gate allow/block/review\", or \"formally verify with Jev\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." license: MIT metadata: version: 0.3.0 @@ -54,7 +54,7 @@ classical method you already trust, substitute it, classify the win "paraphrase brittleness", "allowlist then judge", "TOCTOU-of-Noul", "Jev inside the database / sqlite-jev", "Jev picks bitrate / join order / the model", "wait for Archer", "lint the request / missing - other", "training confronts Choice other / none-of-the-above", "soft AGENTS.md rules vs the linter", "screenshot Choice / omni System One", "extractive quotes / pointer not generator", "compaction summarize vs pointer", "encoder vs Jev compaction backend", "shadow-mode compaction rollout", "CI flaky-vs-real merge gate", "fail-open VOI wake/resume", "claim vs session evidence", "S1 indexer escalate-S2", "Harbor on/off routing", "fail-open vs fail-closed wake vs CI gate", "encoder vs Jev computer-use backend", "hybrid local decide + remote fill", "DONE vs verified success", "stdout prune vs session compaction", "hard envelope then Noul prune", "Cua-S1 vs TypeSafe Jev", "plan vs execute dry-run", "specialist computer-use vs general agent", "local drop-in vs stub scorer", "route vs memory", "when does it hold / extractable from state", "decision model vs constrained LLM", "dual-process S1/S2", "combinatorial grid vs extractive", "uncalibrated local likelihoods", "decision-native RAG", "verbatim ledger vs summary", "judgment as language primitive", or "cascade sign-flip / calibration theater": + other", "training confronts Choice other / none-of-the-above", "soft AGENTS.md rules vs the linter", "screenshot Choice / omni System One", "extractive quotes / pointer not generator", "compaction summarize vs pointer", "encoder vs Jev compaction backend", "shadow-mode compaction rollout", "CI flaky-vs-real merge gate", "fail-open VOI wake/resume", "claim vs session evidence", "S1 indexer escalate-S2", "Harbor on/off routing", "fail-open vs fail-closed wake vs CI gate", "encoder vs Jev computer-use backend", "hybrid local decide + remote fill", "DONE vs verified success", "stdout prune vs session compaction", "hard envelope then Noul prune", "Cua-S1 vs TypeSafe Jev", "plan vs execute dry-run", "specialist computer-use vs general agent", "local drop-in vs stub scorer", "route vs memory", "when does it hold / extractable from state", "decision model vs constrained LLM", "dual-process S1/S2", "combinatorial grid vs extractive", "uncalibrated local likelihoods", "decision-native RAG", "classify-first / read selectively", "living applied-mappings atlas / class patterns", "silence as safer / draft-gate heartbeat", "robotics text-state vs pixels", "verbatim ledger vs summary", "judgment as language primitive", or "cascade sign-flip / calibration theater": read `references/faq.md`, then `references/mental-models.md`, then `references/mixed-architecture.md`, then @@ -118,8 +118,8 @@ classical method you already trust, substitute it, classify the win | Open weights vs constrained decoding vs proprietary API | Three open paths: encoder open-jev / AR constrained decode (TypeAR + pcdServer; decision-token LoRA; packed one-forward logprob on open LLMs; **CUDA/PyTorch local replica** jevify — uncalibrated likelihoods ≠ Noul) / trained decision-only (Laya + ONNX port, Nimble, kev, **blackwood-rlcd** multimodal now, Archer Watch still Watch). Local `/v1/systemone` surfaces: jev-local (stub until `hf`), kev (trained pointer), von (tiny SAN — not a replica). Softmax over allowed tokens ≠ Noul | `references/judgment-class.md` (when-to-use table) | | Entropy as allocator (low / medium / high) | Typed low+medium decisions → System One marginals; high-entropy synthesis → frontier decoder. Product rhetoric, not a meter. **Hypothesis** | `references/judgment-class.md` | | Formal / semi-formal (proof vs judgment) | Sensor vs constraint vs searchlight; Alloy vs Apalache; DST trio; TOCTOU-of-Noul, AI×FM | `references/formal-methods.md` (one-screen: `references/formal-semi-formal.md`) | -| Mixed architecture (judgment model + LLM) | Provider judges, LLM writes, code owns control; not a stack replacement. Advisory sidecar never changes host routing. Dual-process: S1 decides, S2 generates (routing accuracy unmeasured); S1 specialists + S2 coordinator is the same split (description-only greenfield this hour). Internals ≠ FSM; placement is a **component node**. Browser-use strength = DOM-as-text + speculative fan-out, not vision. Fail polarity is per act: skip-wake fail-open vs merge-gate BLOCK fail-closed. Hybrid local decide + remote fill; `DONE` ≠ verified success. Specialist computer-use: plan ≠ execute, dry-run default (Cua-S1; not TypeSafe Jev). Judgment as a language primitive (Ruby `almost_certain?`/`pick`/`rate`). Decision-native RAG: retrieve wide → decide → evidence set → LLM | `references/mixed-architecture.md` | -| Context sieve | Relevance Noul per block; always-keep set in code; stub + recall key. Encoder cousin: GLiNER2.5 retention Choice + char-offset spans (gliner25-compaction); fail-closed keep_full; shadowMode default. Stdout cousin: jev-pruner (Jev Noul after hard ≤10k/JSON-diff envelope; fail-safe original; archive). Session-ledger cousin: carryforward (verbatim facts; Jev scores recall; rules never judged; fail-open dump) | `references/applied-mappings.md#1-context-sieve` | +| Mixed architecture (judgment model + LLM) | Provider judges, LLM writes, code owns control; not a stack replacement. Advisory sidecar never changes host routing. Dual-process: S1 decides, S2 generates (routing accuracy unmeasured); S1 specialists + S2 coordinator is the same split (description-only greenfield this hour). Internals ≠ FSM; placement is a **component node**. Browser-use strength = DOM-as-text + speculative fan-out, not vision. Fail polarity is per act: skip-wake fail-open vs merge-gate BLOCK fail-closed. Hybrid local decide + remote fill; `DONE` ≠ verified success. Specialist computer-use: plan ≠ execute, dry-run default (Cua-S1; not TypeSafe Jev). Judgment as a language primitive (Ruby `almost_certain?`/`pick`/`rate`). Decision-native RAG: retrieve wide → decide → evidence set → LLM. Classify-first MCP (topology A): content to the judge without entering main agent context first. Generative UI: model decides, compiler emits. Draft-gate silence ≠ safer (heartbeat). Living class-pattern atlas (not a 342-title dump) | `references/mixed-architecture.md` | +| Context sieve | Relevance Noul per block; always-keep set in code; stub + recall key. Encoder cousin: GLiNER2.5 retention Choice + char-offset spans (gliner25-compaction); fail-closed keep_full; shadowMode default. Stdout cousin: jev-pruner (Jev Noul after hard ≤10k/JSON-diff envelope; fail-safe original; archive). Session-ledger cousin: carryforward (verbatim facts; Jev scores recall; rules never judged; fail-open dump). Classify-first MCP cousin: jev-sift (batch path/url/text → Jev without entering main agent context; uncertain/errors/truncation ≠ irrelevant) | `references/applied-mappings.md#1-context-sieve` | | Exact-text keep / drop | Choice include/exclude/mixed over candidates code already holds. Extractive quotes / pointer-not-generator (model never writes the excerpt; char-offset compaction same species). Observed a11y/DOM controls: score among them; code clicks (Jev or GLiNER2 or Cua-S1 option-attention) | `references/applied-mappings.md#2-exact-text-keep--drop` | | Environment / harness triage | Scan every step for env failure; LLM autopsy only on flags. Merge-gate cousin: cluster in code, judge labels cause, policy owns PASS/BLOCK (latch; judge never says ignore alone) | `references/applied-mappings.md#3-environment--harness-triage` | | Moderation and ranking | Hold-before-publish vs graded rerank; fail policy per action | `references/applied-mappings.md#4-moderation-and-ranking` | @@ -131,13 +131,13 @@ classical method you already trust, substitute it, classify the win | Feature engineering / multi-criteria analysis | Nouls + Score distributions as named features, weights in code | `references/mappings.md#1-semantic-judgments--features-and-explicit-utility` | | Selective classification / decision theory | Thresholds from action costs, abstention paths | `references/mappings.md#2-probabilistic-judgments--cost-sensitive-decisions` | | Decision tables / circuits / state machines | Judgment predicates, code owns transitions. Language primitive: Ruby `chance`/`pick`/`rate` as control flow (hunch; English-as-config; fail polarity per action) | `references/mappings.md#3-semantic-predicates--decision-circuits` | -| Retrieve + expensive relevance fn | Bounded rerank of a retrieved shortlist. Decision-native RAG: retrieve wide → decide explicitly → evidence set → conflict resolve → reason only over kept evidence (embeddings stay candidate generators; no universal benchmark) | `references/mappings.md#4-retrieval--bounded-semantic-reranking` (independent TREC DL2019 benchmark: Jev zero-shot best MAP 0.4748, nDCG@10 0.683 vs tuned monoBERT 0.718 — competitive, not dominant) | +| Retrieve + expensive relevance fn | Bounded rerank of a retrieved shortlist. Decision-native RAG: retrieve wide → decide explicitly → evidence set → conflict resolve → reason only over kept evidence (embeddings stay candidate generators; no universal benchmark). Classify-first MCP: same sandwich on agent I/O (path/url/text → judge; main LLM opens survivors) | `references/mappings.md#4-retrieval--bounded-semantic-reranking` (independent TREC DL2019 benchmark: Jev zero-shot best MAP 0.4748, nDCG@10 0.683 vs tuned monoBERT 0.718 — competitive, not dominant) | | Store as semantic index (SQL / SQLite / zoxide / dataframe) | Cheap exact predicates first; typed questions on the remainder. In-engine extension (sqlite-jev) vs CLI rewrite (jevql) vs path index (joxide) vs dataframe columns (jevpandas / jevframe) | `references/mappings.md#4-retrieval--bounded-semantic-reranking` | | Soft judgment inside a hard envelope | Model may only match the deterministic policy or be more conservative (bitrate ABR; query-planner override-when-confident; compaction mutations/shell operators → keep_full; stdout prune: ≤10k/JSON-diff-whole-doc untouched, then Noul; Cua-S1: plan≠execute, dry-run, fail-closed checkbox/fill; pre-exec toolgate: allow/block/review — Jev is not authorization; guard error/timeout stops) | `references/mappings.md#12-runtime-assurance-sandwich-hypothesis`; `references/mappings.md#18-structural-prove--soft-remainder-hypothesis-as-domain-general-empirical-as-named-shapes` | -| Value of information / gather as an act | Pay for another observation only if EV(decision) improves more than cost; abstain from calling *any* model when a regex already answers (meta-VOI). Fail-open wake/resume: skip the LLM turn only if the judge answers and p(wake) is low (Horvitz). Selective memory: verbatim ledger + scored recall (carryforward; rules never judged; 9×3 is a hint) | `references/mappings.md#6-value-of-information--gather-as-an-enumerated-act` (**Hypothesis** until a labeled act/outcome log; 149-row receipt is Empirical as a shape; wakegate 21/21 is smoke) | +| Value of information / gather as an act | Pay for another observation only if EV(decision) improves more than cost; abstain from calling *any* model when a regex already answers (meta-VOI). Fail-open wake/resume: skip the LLM turn only if the judge answers and p(wake) is low (Horvitz). Selective memory: verbatim ledger + scored recall (carryforward; rules never judged; 9×3 is a hint). Classify-first read: pay for a full agent open iff relevance (or typed question) says it might change the act (jev-sift; errors/truncation ≠ irrelevant) | `references/mappings.md#6-value-of-information--gather-as-an-enumerated-act` (**Hypothesis** until a labeled act/outcome log; 149-row receipt is Empirical as a shape; wakegate 21/21 is smoke) | | Signal detection / ROC | Criterion and operating point from costs and base rate, not accuracy | `references/mappings.md#7-signal-detection--criterion-not-accuracy` (**Hypothesis** for non-SWE plots) | | Org / safety control structure | Sensor ≠ constraint (Leveson); STPA if the sensor lies | `references/mappings.md#8-control-structure--sensor--constraint-leveson` | -| Search / control loops (any domain) | Algorithm stays yours; judgment substitutes one classifier step | `references/mappings.md#9-search--control-loops--one-substituted-classifier-step` | +| Search / control loops (any domain) | Algorithm stays yours; judgment substitutes one classifier step. Robotics text-state (geometry-as-text, not pixels); do not replace A* / a solver with a Noul | `references/mappings.md#9-search--control-loops--one-substituted-classifier-step` | | Spec property pipeline | Rank candidate props; checker owns validity | `references/mappings.md#10-spec-property-pipeline-hypothesis` (**Hypothesis**) | | Alloy instance loop | Cluster CEXs; Analyzer owns in-scope truth | `references/mappings.md#11-alloy-instance-loop-hypothesis` (**Hypothesis**) | | Runtime assurance sandwich | Abstain → RV/monitor → act | `references/mappings.md#12-runtime-assurance-sandwich-hypothesis` (**Hypothesis**) | diff --git a/.agents/skills/augustus/references/agent-self-assessment.md b/.agents/skills/augustus/references/agent-self-assessment.md index 839a180..d4d7720 100644 --- a/.agents/skills/augustus/references/agent-self-assessment.md +++ b/.agents/skills/augustus/references/agent-self-assessment.md @@ -38,7 +38,11 @@ relying: foreman, pi-jev, pi-warden, winnow, fast-jev-compaction, jev-judgment. — loop `DONE` is termination, not verified success; apps inspect the actual result (`notes.md` §52). 4. **Stuck-detector**: three failures with the same strategy → ask for a - new hypothesis, not another retry. + new hypothesis, not another retry. **Silence is not safer:** a draft + gate that treats a missing Jev answer as "don't send" holds forever. + Missing verdict needs a fail-open / heartbeat — not a block and not a + pass. Contrast Abide `<0.5` silence (the *edit proceeds*). Showcase + class pattern on [jevable.com](https://jevable.com/) (`notes.md` §56). 5. **Supervision during long runs** (foreman): separate concurrent loop estimates `meaningful_progress`, `implementation_complete`, `tests_sufficient`, `worker_stuck`, `work_off_track`, @@ -102,6 +106,11 @@ relying: foreman, pi-jev, pi-warden, winnow, fast-jev-compaction, jev-judgment. and corrections always return; fail-open dump if the scorer is down. Nine entries × three tasks is a hint, not proof (`notes.md` §55). Do not copy mcp add. + Classify-first cousin: + [jev-sift](https://github.com/kbhuw/jev-sift) — batch path/url/text + → Jev **before** the main agent reads; uncertain/errors/truncation + ≠ irrelevant. Transport tests ≠ accuracy. No LICENSE this pass + (`notes.md` §56). Do not copy plugin how-to. ## Non-negotiable boundaries @@ -124,6 +133,8 @@ relying: foreman, pi-jev, pi-warden, winnow, fast-jev-compaction, jev-judgment. stops on block / review-without-approval / guard error (`notes.md` §55). Session-memory *omit* fails open (dump the ledger): [carryforward](https://github.com/Dharundp6/jev-carryforward). + Draft-gate *silence* is the opposite mistake: treating no-answer as + a hold. Missing verdict needs a heartbeat (`notes.md` §56). - Cache identical judgments (~120s) and deduplicate sibling calls into one in-flight request. - pi-warden measured cost makes continuous guarding viable: ~$0.00004 and diff --git a/.agents/skills/augustus/references/applied-mappings.md b/.agents/skills/augustus/references/applied-mappings.md index 7f9f409..35b4e7b 100644 --- a/.agents/skills/augustus/references/applied-mappings.md +++ b/.agents/skills/augustus/references/applied-mappings.md @@ -75,6 +75,21 @@ corrections **always return in full** (Jev never votes on a rule). Fail-open: no key → whole list. Thresholds 0.60 full / 0.30–0.60 one line are *theirs*. Nine entries × three tasks is a **hint, not proof** (`notes.md` §55). Do not copy `mcp add`. +**Classify-first MCP, same family, different job (Empirical as +README / schema, 2026-09-19 ~00:38):** +[jev-sift](https://github.com/kbhuw/jev-sift) — batch path / public +URL / inline text (or a tool description) → Jev relevance or 1–8 +typed questions **before** the main agent reads. Content goes to +the judge without entering main agent context first (paths/URLs). +Uncertain → closer look; errors and truncation ≠ irrelevant. Hard +envelope (theirs): 50 items, 60k char, 2 MB / 20 s, public-IP only, +no JS/cookies/login, PDFs unsupported. Transport tests ≠ accuracy. +No LICENSE this pass. Same retrieve-wide → decide → evidence-set +family as decision-native-rag-skills. Cousins: typesafe-screening-mcp, +kazuhideoki/jev-search, jev-pruner (after Bash), carryforward (ledger +you already hold). Not jev-routing (host adapter). Topology A MCP +(LLM outer loop). Do not copy plugin / `mcpServers` / key-file +how-to (`notes.md` §56). Local teacher-copy for the same hole: [`SargeDev/jev-gate-student-b`](https://huggingface.co/SargeDev/jev-gate-student-b) (Qwen2.5-0.5B LoRA; P(relevant) from yes/no logits; 148,160-row @@ -153,6 +168,16 @@ screenshot). Specialist-form cousin, **not TypeSafe Jev:** option-attention among observed elements (fill/check/click/skip); code owns execution order; dry-run default; source-only (`notes.md` §54). +**Score-among-observed atlas (Empirical as public showcase class +pattern, 2026-09-19 ~00:38):** +[jevable.com](https://jevable.com/) — candidates already on the +page (a11y/DOM, ads, on-screen posts); the model scores; **code** +clicks / filters. Not a 342-title dump. Cross-link: jev-ultrafast / +gliner2-ultrafast / solari-reflex / cua-s1 / laya-mind2web. Your +Signal: score posts already on screen, apply rules locally — same +judge-once / re-policy family as Near Here. Do not merge Flights +7 s / $0.0039 with gliner2-ultrafast 12.20 s; computer-use "100×" +is a **claim** (`notes.md` §56). **DOM-as-text + fan-out (Empirical as atlas browser-use *shape*):** a screenshot task translated into a structured DOM snapshot as `state`, then speculative questions over numbered candidates — not vision @@ -258,6 +283,21 @@ generators. Provider-agnostic; no bundled Python harness; **no universal benchmark**. Default migration gates are starting targets. Offline replay → shadow → canary → A/B (`notes.md` §55). Do not ship because an LLM judge prefers it. +**Classify-first agent I/O of the same sandwich (Empirical as +README, 2026-09-19 ~00:38):** +[jev-sift](https://github.com/kbhuw/jev-sift) — file lists and +public URLs stay candidate generators; the judge sees content; the +main LLM opens only items worth a closer look. Inline text the +agent already read cannot recover that cost. Uncertain/errors/ +truncation ≠ irrelevant. Mocks ≠ accuracy (`notes.md` §56). +**Living applied-mappings atlas (Empirical as showcase; 342 is +*their* count):** +[jevable.com](https://jevable.com/) — class patterns (intent +columns, score-among-observed, VOI gates, generative UI decide, +robotics text-state, draft-gate fail modes), not a hit list. +JSON-LD first page is 36; `pageSize` 36. No public API this pass. +Maker clocks stay claims unless already a named receipt +(`notes.md` §56). **Recursive file search (MED; distinguish from federated web):** [kazuhideoki/jev-search](https://github.com/kazuhideoki/jev-search) scores local files then fzf — **not** diff --git a/.agents/skills/augustus/references/faq.md b/.agents/skills/augustus/references/faq.md index 59164af..26c5225 100644 --- a/.agents/skills/augustus/references/faq.md +++ b/.agents/skills/augustus/references/faq.md @@ -481,7 +481,35 @@ retrieve wide → decide explicitly → evidence set → resolve conflicts generators. Provider-agnostic; no bundled harness; **no universal benchmark**. Default migration gates are starting targets, not promises. Do not ship because an LLM judge prefers it. -`mappings.md` §4; `notes.md` §55. +[jev-sift](https://github.com/kbhuw/jev-sift) is the same sandwich +on **agent I/O**: classify first, read selectively; content to Jev +without entering main agent context first (paths/URLs). Uncertain / +errors / truncation ≠ irrelevant. Transport tests ≠ accuracy. +`mappings.md` §4; `notes.md` §55, §56. + +## Dump files into context, or classify first? + +Classify first when the items are **not** already in the main agent +context. [jev-sift](https://github.com/kbhuw/jev-sift): batch path / +public URL / inline text → Jev; the main LLM opens survivors. +Uncertain → closer look; errors and truncation ≠ irrelevant. Inline +text the agent already read cannot recover that cost. Hard envelope +in code (50 / 60k / 2MB / public-IP). Transport tests ≠ accuracy. +Same family as decision-native RAG. Topology A MCP — not +jev-routing (host adapter). Do not copy plugin how-to. +`applied-mappings.md` §1; `notes.md` §56. + +## Is jevable.com a 342-title census? + +No. [jevable.com](https://jevable.com/) is a living **applied-mappings +atlas**: extract class patterns (intent columns, score-among-observed, +VOI gates, generative UI decide, robotics text-state, draft-gate fail +modes). The site claims **342** curated projects this pass; homepage +JSON-LD lists **36** (first page / `pageSize` 36). We did not enumerate +titles. Maker clocks stay **claims** unless already a named receipt. +Cross-link exemplars already in notes; do not dump a hit list. Not a +model. Not multimodal substrate. Archer still Watch. +`applied-mappings.md`; `notes.md` §56. ## Did Jev beat nano as an escalation gate? @@ -585,8 +613,12 @@ Tool *execution* fails closed on block/timeout: [toolgate](https://github.com/fdemir/toolgate) (Jev is not authorization). Ruby validations in [hunch](https://github.com/carldaws/hunch) `rescue nil` at save — -spam gates should not. Same sandwich, opposite authorized act. -`notes.md` §50, §51, §53, §55. +spam gates should not. Draft-gate *silence* is the same rule: +missing verdict is not a block and not a pass — fail-open / +heartbeat, do not hold forever ([jevable.com](https://jevable.com/) +class pattern). Contrast Abide `<0.5` silence (the *edit proceeds*). +Same sandwich, opposite authorized act. +`notes.md` §50, §51, §53, §55, §56. ## Is observe→score→act Jev-only? diff --git a/.agents/skills/augustus/references/mappings.md b/.agents/skills/augustus/references/mappings.md index a8f2401..1c32f67 100644 --- a/.agents/skills/augustus/references/mappings.md +++ b/.agents/skills/augustus/references/mappings.md @@ -52,6 +52,14 @@ as firehose sliders. **Beyond SWE Nouls; price and deadline exact); apartment shortlist (commute/light/ noise Scores; rent exact); hiring scorecard (evidence Nouls; labor-law vetoes in policy). Full gallery: `mental-models.md` §MCDA. +**Intent columns / catalog MCDA (Empirical as *shape*; clocks are +claims, 2026-09-19 ~00:38):** +[jevable.com](https://jevable.com/) class pattern: a heading +("Urgency") scores each row. Same hole as jevpandas / jevframe +(`mappings.md` §4) and dabit3 spreadsheet JUDGE/SCORE/CHOOSE. +Snack multi-criteria at catalog scale is a **maker claim** (3,000 / +28 s / $0.11) unless independently re-run. Weights and vetoes stay +in code (`notes.md` §56). **Counterexample** (from **Contract** Score docs): levels 0,1,2 with distributions `[0,1,0]` vs `[0.5,0,0.5]` both score 1.0 with radically different extreme-outcome risk — always read probabilities beside the @@ -197,7 +205,8 @@ PyPI; pandas **and** Polars `.jev`; full `p__` columns; no silent renormalize; one row per request). Same hole, two surfaces. Row contents leave the store (same residency warning as AU health). Do not copy SQL, env, or CLI flags. -`notes.md` §42, §44, §46, §48. +`notes.md` §42, §44, §46, §48. Intent-column / snack MCDA *shape*: +`mappings.md` §1; `notes.md` §56. **Decision-native evidence set (Empirical as architecture; Hypothesis as a measured win, 2026-09-18 ~17:48):** @@ -216,6 +225,16 @@ Local-file cousin: [kazuhideoki/jev-search](https://github.com/kazuhideoki/jev-search) (recursive files + fzf) — **not** superagents-lab/jev-search (federated web). Max-over-chunks ≠ calibrated whole-file p. +**Classify-first MCP (Empirical as README / schema, 2026-09-19 +~00:38):** +[jev-sift](https://github.com/kbhuw/jev-sift) is the same sandwich +on agent I/O: retrieve-wide (paths / public URLs / inline text) → +decide (relevance or 1–8 typed questions) → the main LLM opens +only the evidence set. Content never enters main agent context +first (paths/URLs). Hard envelope in code. Transport tests ≠ +accuracy. No LICENSE this pass. Cousin of typesafe-screening-mcp +(abstracts never enter the LLM conversation). Not jev-routing +(host adapter). `notes.md` §56. ## 5. Hierarchy → bounded heuristic search @@ -341,6 +360,14 @@ task; constraints/corrections always return (never judged). Fail- open dump if the scorer is down. Pay for a scored brief iff it beats dumping the whole file. No accuracy claim until a proper test (`notes.md` §55). Do not copy mcp add. +**Classify-first read (Empirical as README; Hypothesis as a +measured win, 2026-09-19 ~00:38):** +[jev-sift](https://github.com/kbhuw/jev-sift) — pay for a full +agent open iff the relevance (or typed question) says it might +change the act. Uncertain → closer look. Errors and truncation are +**not** evidence of irrelevance. Webpage fetch still costs +bandwidth; this saves the *agent's* read, not the download. +Mocks ≠ accuracy (`notes.md` §56). **Beyond SWE (Hypothesis until you log act/outcome pairs):** full PDF vs abstract; customer call vs CRM fields that already fail a hard rule (credit limit is exact); blood test vs @@ -482,6 +509,17 @@ same substituted-classifier *job* on a specialist form contract (option-attention; plan ≠ execute; not TypeSafe Jev; source-only, `notes.md` §54). +**Robotics text-state, same job different body (Empirical as +showcase class pattern, 2026-09-19 ~00:38):** +MuJoCo robot-arm on [jevable.com](https://jevable.com/): Jev does +not accept images; simplified geometry and contacts **as text**; +two-call split (what to do, then how to move). MOSS: Jev picks the +target; the robot picks up. Cousins: [jev-drone](https://github.com/RomanSlack/jev-drone) +(code at 500/50 Hz, Jev advisory 2.5 Hz); Doom JSON, not pixels. +Drawing-pixel-parallel is a **claim** — contrast MuJoCo honesty. +Do not replace A* or a Sudoku solver with a Noul. Archer still +Watch (`notes.md` §56). + **Structure induction over a bag (Empirical as a *shape*, 2026-09-18):** [`Joymfl/dag-jev`](https://github.com/Joymfl/dag-jev) — unordered items in, pairwise "does i depend on j?" judgments, DAG in `petgraph`. Code diff --git a/.agents/skills/augustus/references/mental-models.md b/.agents/skills/augustus/references/mental-models.md index 433e1e8..29dd488 100644 --- a/.agents/skills/augustus/references/mental-models.md +++ b/.agents/skills/augustus/references/mental-models.md @@ -488,7 +488,11 @@ Use these as *existence proofs of a position*. Write your own card. | Phishing / fraud screen | hold vs deliver | SDT criterion on a Noul | blocklist, SPF/DKIM exact (**Hypothesis**) | | Personal ops | cook done / not | "looks done" Noul | thermometer probe | | Org safety | stop the line | sensor Noul | interlock, two-person rule | -| Knowledge / RAG | reason only over kept evidence | retrieve wide → decide → evidence set (**Empirical** as architecture: decision-native-rag-skills; **Hypothesis** as a measured win) | Conflict/temporal/provenance in code; embeddings generate candidates | +| Knowledge / RAG | reason only over kept evidence | retrieve wide → decide → evidence set (**Empirical** as architecture: decision-native-rag-skills; classify-first MCP cousin: jev-sift; **Hypothesis** as a measured win) | Conflict/temporal/provenance in code; embeddings / file lists generate candidates; errors/truncation ≠ irrelevant | +| Agent I/O | classify first, read selectively | batch path/url/text → relevance or typed questions (**Empirical** as README: jev-sift; topology A MCP) | Hard envelope (50 / 60k / 2MB / public-IP); main LLM opens survivors | +| Spreadsheet / catalog | named semantic columns | heading scores each row (**Empirical** as *shape*: jevpandas / jevframe; jevable intent columns). Snack MCDA clocks are **claims** | Weights, vetoes, exact fields in code | +| Robotics / control | observe → decide → act on a body | Choice on **geometry-as-text**, not pixels (**Empirical** as showcase: MuJoCo / MOSS; cousins jev-drone, Doom JSON) | Kinematics / Hz in code; two-call split; do not replace A*. Drawing-pixel claim ≠ Archer | +| Draft quality gate | kill drafts that break rules | quality Noul/Score (**Empirical** as fail *mode*: silence treated as safer) | Fail-open / heartbeat on missing verdict; contrast Abide `<0.5` (edit proceeds) | | Session memory | next task sees last session's facts | scored recall over a verbatim ledger (**Empirical**: carryforward; 9×3 hint) | Constraints always-keep; fail-open dump; never summarize | | Application control flow | `if` / `case` on a judgment | `chance`/`pick`/`rate` as language primitives (**Empirical**: hunch; English-as-config) | Fail polarity per action; stub backend | | Healthcare huddle / recon / inbox | escalate / hold / route | S1 remainder after NEWS2/code (**Empirical** as synthetic report: explore-typesafe-ai; **not clinically validated**) | NEWS2, recon, routing in code; S2 blinded review | diff --git a/.agents/skills/augustus/references/methods-catalog.md b/.agents/skills/augustus/references/methods-catalog.md index 52c9e0c..ced5d86 100644 --- a/.agents/skills/augustus/references/methods-catalog.md +++ b/.agents/skills/augustus/references/methods-catalog.md @@ -31,7 +31,7 @@ judgment component is new). | Self-consistency / ensembling of judges | Repeated independent ratings of the same object | N repeats over one state (output tokens free); entropy/disagreement across repeats as the review signal | Aggregation, escalation policy | **Empirical recipe** (self-consistency: nouls cookbook) | | Judge qualification (interrater reliability) | A judge worth gating must be repeatable | Repeated judgments over frozen outputs before trusting either Jev or LLM as judge | Variance stats, agreement metrics | **Empirical recipe** (jev-as-a-judge: 224–279× tighter than GPT judge) | | Neyman–Pearson / selective classification | Decision threshold under error costs | One threshold per action, set on split A, reported on split B; abstention path | Loss model, ROC analysis | **Contract + empirical** (confidence-routing; evaluator script) | -| Value of information (EVPI / EVSI) | Whether another observation is worth its cost | Gather as an enumerated act; pay iff expected decision-loss drop > cost. Wake/resume: skip the LLM turn only if the judge answers and p is low. Selective memory: score the ledger against the task; dump on failure | Cost of the observation; the loss table; skip-limit; always-keep rules | **Hypothesis** as a numeric calculator; **Contract** as the placement (`mappings.md` §6). wakegate 21/21 is smoke (`notes.md` §51). carryforward 9×3 is a hint (`notes.md` §55) | +| Value of information (EVPI / EVSI) | Whether another observation is worth its cost | Gather as an enumerated act; pay iff expected decision-loss drop > cost. Wake/resume: skip the LLM turn only if the judge answers and p is low. Selective memory: score the ledger against the task; dump on failure. Classify-first: pay for a full agent open iff relevance says it might change the act | Cost of the observation; the loss table; skip-limit; always-keep rules; hard I/O envelope | **Hypothesis** as a numeric calculator; **Contract** as the placement (`mappings.md` §6). wakegate 21/21 is smoke (`notes.md` §51). carryforward 9×3 is a hint (`notes.md` §55). jev-sift mocks ≠ accuracy (`notes.md` §56) | | Signal detection (Green & Swets) | Evidence variable + criterion | Noul as noisy evidence; t from costs and base rate; ROC/PR on your labels | Operating point, base-rate tracking | **Hypothesis** for non-SWE plots; **Empirical** as moderation *shape* (`mappings.md` §7) | | Reliability calibration (Platt/temperature) | Raw scores → calibrated probabilities | Noul is natively calibrated **in-distribution only**; verify with reliability bins on your own population; re-fit a correction out-of-distribution | Calibration fitting, binning | **Empirical recipe** (ECE 0.0313 in-distribution; 32% OOD collapse — Archer Hume). Atlas: DAIR Emotion dangerous-high (48% / 0.819); DMB S5 ECE 0.246 (`notes.md` §49) | | Frozen-protocol bake-off vs constrained LLMs | Same items, accuracy + ECE + latency + cost + honesty | Decision-model as one contender class, not the score | Protocol, raw logs, baselines | **Empirical as Harbor/jevals practice** (DMB v2; jevals-data CC-BY-4.0 recompute-from-logs; `notes.md` §49) | @@ -48,12 +48,12 @@ judgment component is new). | Beam search over taxonomies | Which branches deserve expansion | Choice distributions as branch priority; keep K paths where ambiguity is early | Frontier, budget, final selection | **Empirical recipe** (beam K=3 cookbook) | | Structure induction over a bag | Pairwise "does i depend on j?" (or Choice over order) | One judgment per pair; DAG / scheduler in code | Topology, cycles, execution | **Empirical as a shape** (dag-jev experiment; empty README; no metrics, `notes.md` §48) | | Collab-arm product loop | Scripted legal set vs LLM-propose vs unconstrained | Choice over legal actions; stop on low p rather than guess | Legality, Wilson/McNemar, ceiling flags | **Empirical as a harness shape** (jev-testbench; bake into jevals/Harbor, `notes.md` §48) | -| Computer-use observe → score → act | Which observed control matches the current requirement | Score / Choice / option-attention among a11y/DOM candidates (Jev *or* GLiNER2 *or* Laya *or* Cua-S1); code clicks | Observation, freshness, dates; never generate selectors or values; independent outcome check (`DONE` ≠ success); plan ≠ execute | **Empirical recipe** as shipped loops (jev-ultrafast / solari-reflex Jev; gliner2-ultrafast GLiNER2; laya-mind2web DOM-index Laya); **Empirical as README/MODEL_CARD** for Cua-S1 (source-only, not TypeSafe Jev, `notes.md` §54); contrast blackwood-rlcd screenshot, `notes.md` §48, §52 | +| Computer-use observe → score → act | Which observed control matches the current requirement | Score / Choice / option-attention among a11y/DOM candidates (Jev *or* GLiNER2 *or* Laya *or* Cua-S1); code clicks. Robotics cousin: Choice on geometry-as-text, not pixels | Observation, freshness, dates; never generate selectors or values; independent outcome check (`DONE` ≠ success); plan ≠ execute; kinematics / Hz in code | **Empirical recipe** as shipped loops (jev-ultrafast / solari-reflex Jev; gliner2-ultrafast GLiNER2; laya-mind2web DOM-index Laya); **Empirical as README/MODEL_CARD** for Cua-S1 (source-only, not TypeSafe Jev, `notes.md` §54); contrast blackwood-rlcd screenshot, `notes.md` §48, §52. **Empirical as showcase** MuJoCo / MOSS text-state (`notes.md` §56); drawing-pixel claim ≠ Archer | | Screening / Wald sequential tests | Pass / fail / keep-looking per candidate | One Noul gate per candidate in one batched request; budget in code | Sequential rule, stop boundaries | **Hypothesis** | | STPA / STAMP control structure | Sensor reading vs enforced constraint | Judgment as sensor; constraints in policy/code/interlock; STPA table if the sensor lies | The constraint, the actuator, the probe | **Contract** as ownership; **Hypothesis** as domain product (`mappings.md` §8) | | PufferLib / Ocean env contracts | Does this episode look like a known trainer-bug mode? | Cluster failing episodes; never "the policy is correct" | Seeded serial env, Ocean sanity, observed rewards | **Hypothesis** as placement; **Contract** that Ocean is not a comparative baseline (`formal-methods.md` DST trio) | | Routing / dispatch (OR) | Which queue/agent owns this item | Choice + confidence-gated escalation; code owns capacity | Cost matrix, capacity constraints | **Empirical recipe** (intent-routing; LlamaIndex Jev selectors; skillranker) | -| Cascade / prefilter (IR) | Cheap reject before an expensive scorer or LLM | Per-candidate Noul/Score; fail-open on drop, fail-closed on dispatch. Decision-native: evidence-set after wide retrieve | Candidate generation, always-keep set, recall keys; conflict/provenance | **Empirical recipe** (classifying RAG passages; jevprune; git-jev-stage). **Empirical as architecture** (decision-native-rag-skills; Hypothesis as a measured win, `notes.md` §55) | +| Cascade / prefilter (IR) | Cheap reject before an expensive scorer or LLM | Per-candidate Noul/Score; fail-open on drop, fail-closed on dispatch. Decision-native: evidence-set after wide retrieve. Classify-first MCP: content to the judge without entering main agent context first | Candidate generation, always-keep set, recall keys; conflict/provenance; hard I/O envelope | **Empirical recipe** (classifying RAG passages; jevprune; git-jev-stage). **Empirical as architecture** (decision-native-rag-skills; Hypothesis as a measured win, `notes.md` §55). **Empirical as README** (jev-sift; mocks ≠ accuracy, `notes.md` §56) | | Knapsack / portfolio selection | Per-item feature vector from text | Fan-out nouls/scores as features; optimizer in code | Constraint solver, weights | **Hypothesis** (mapping 1 shape) | ## Information theory & signals @@ -61,7 +61,7 @@ judgment component is new). | Method | Judgment-shaped component | Jev substitution | Stays in code | Status | |---|---|---|---|---| | Entropy as uncertainty signal | Measuring "how spread is this belief" | Entropy of returned distributions across repeats or options — computed in code from returned probabilities | All arithmetic | **Empirical recipe** (cookbook pattern) | -| Detector / Neyman filter (context) | Is this artifact relevant to the current task? | One relevance Noul per block before it enters context; stub + recall key. Encoder cousin: retention Choice + span locate, copy verbatim. Indexer cousin: GLiNER extract on the bulk, escalate LLM on the tail. Stdout cousin: Noul per chunk after a hard size/format envelope. Session-ledger cousin: score verbatim facts; rules never judged | Cache, recall, safety keeps; mutation/shell envelope; ≤10k/JSON-diff pass-through; archive dropped spans; do not dump the repo if S1 failed to load; fail-open dump of the ledger | **Empirical recipe** (winnow ≤0.22 hide; compaction 2-noul rule; pi-jev-context hide-not-delete; gliner25-compaction GLiNER2.5 keep_full/keep_evidence/keep_call_only/drop, `notes.md` §50; s1-graphify-indexer degraded-load / no-invent-edges, 10–50× unfilled, `notes.md` §51; jev-pruner Bash stdout prune, `notes.md` §53; carryforward, `notes.md` §55) | +| Detector / Neyman filter (context) | Is this artifact relevant to the current task? | One relevance Noul per block before it enters context; stub + recall key. Encoder cousin: retention Choice + span locate, copy verbatim. Indexer cousin: GLiNER extract on the bulk, escalate LLM on the tail. Stdout cousin: Noul per chunk after a hard size/format envelope. Session-ledger cousin: score verbatim facts; rules never judged. Classify-first cousin: batch path/url/text to the judge before the main agent reads | Cache, recall, safety keeps; mutation/shell envelope; ≤10k/JSON-diff pass-through; archive dropped spans; do not dump the repo if S1 failed to load; fail-open dump of the ledger; 50 / 60k / 2MB / public-IP envelope | **Empirical recipe** (winnow ≤0.22 hide; compaction 2-noul rule; pi-jev-context hide-not-delete; gliner25-compaction GLiNER2.5 keep_full/keep_evidence/keep_call_only/drop, `notes.md` §50; s1-graphify-indexer degraded-load / no-invent-edges, 10–50× unfilled, `notes.md` §51; jev-pruner Bash stdout prune, `notes.md` §53; carryforward, `notes.md` §55; jev-sift classify-first, `notes.md` §56) | | Anomaly detection | Does this deviate from expected shape? | Guard nouls + harm Score over {input, output, tool trace} | Baselines, alert thresholds | **Empirical recipe** (guardrails cookbook; pi-jev output judge) | | Allowlist ∩ remainder (code-then-model) | Unlisted / unstructured leftovers after a **proof** | Typed questions only on the unknown tier; admit iff every p < τ | Proven/refused in code; cannot block unless a sandbox sits under | **Empirical recipe** (jevgate 0/59 unsafe unasked held-out; allowlist *proves* read-only verbs; doc-router 1.74× $). Domain-general: `mappings.md` §18 | | Decision-token LoRA (constrained-AR) | Specialize a generator for parallel constrained fields | Loss only on the single decision token; KV broadcast across fields | Schema, candidate tokens, policy | **Empirical recipe** as Foodoo1 200-case / 4-field receipt (fraud_risk 64→95%, overall 85.2→98.8%, ~234 ms); **Hypothesis** as a general recipe. Synthetic; not a financial product. Softmax ≠ Noul | @@ -83,7 +83,7 @@ judgment component is new). | Method | Judgment-shaped component | Jev substitution | Stays in code | Status | |---|---|---|---|---| -| Multi-criteria decision analysis | Attribute scores per alternative | Composite scores with weights in code (re-weight without re-inference) | Weight policy, Pareto views | **Empirical recipe** (mapping 1) | +| Multi-criteria decision analysis | Attribute scores per alternative | Composite scores with weights in code (re-weight without re-inference). Intent columns: a heading scores each row | Weight policy, Pareto views, exact fields | **Empirical recipe** (mapping 1). Showcase *shape*: jevable intent columns; snack MCDA clocks are **claims** (`notes.md` §56) | | Mechanism/game response (adversarial state) | Opponent-intent / bluff / risk read on a state | Choice over reads + risk Score; policy in code; 2.5Hz-style advisory rate | Strategy solvers, exploitative math | **Empirical recipe** (jev-trader, game agents) | | Auction/market event classification | Is this signal material? Direction? | Choice over event classes + urgency nouls; execution in code | Order execution, risk limits | **Empirical recipe** (jev-trader 81ms/block) | | Assignment / scheduling (OR) | Soft affinity per pair | Score/Noul as a cost *feature*; solver owns capacity/legality | ILP/heuristic, fairness | **Hypothesis** (`mappings.md` §15) | diff --git a/.agents/skills/augustus/references/mixed-architecture.md b/.agents/skills/augustus/references/mixed-architecture.md index 40a20d8..abd9b2d 100644 --- a/.agents/skills/augustus/references/mixed-architecture.md +++ b/.agents/skills/augustus/references/mixed-architecture.md @@ -185,6 +185,8 @@ not a global virtue: | Drop a RAG chunk or log line | **Fail open** (keep on error) | A false drop loses evidence; a false keep costs tokens | | Compact / drop a completed tool result | **Fail closed** to keep-full (`gliner25-compaction`) | Compaction is a destructive edit of memory. Uncertain *looks* like keep-on-error from the evidence side; name the *reduction* as the act. Contrast Abide / jevgate fail-open | | Omit a session-memory fact from the brief | **Fail open** (dump the whole ledger) (`carryforward`) | Scoring failure must not hide a rule. Constraints/corrections always return; Jev never votes on them | +| Open a file / URL into agent context | **Fail open** (closer look on error / truncation / uncertain) (`jev-sift`) | False drop loses evidence. Errors and truncation are not irrelevance. Transport failure ≠ "irrelevant" | +| Withhold a draft because the judge is silent | **Fail open** (heartbeat / proceed or escalate; do not hold forever) | Missing verdict is not a block and not a pass. Contrast Abide `<0.5` silence (the *edit proceeds*) | | Prune Bash stdout before the LLM | **Fail closed** to original (`jev-pruner`) | Dropping the log is irreversible. ≤10k / JSON-diff-whole-doc prove pass-through; archive/Jev/incomplete-score failure keeps the result. Harbor plugin-eval cannot reach Jev → cannot prune | | Skip waking a sleeping agent | **Fail open** (wake on error / unsure / no key) (`wakegate`) | Skip is the irreversible act. User-message, skip-limit, nothing-to-judge, and p in 0.2–0.5 all wake. Contrast pi-jev-approver fail-closed without a key | | Merge a red CI run | **Fail closed** on `--gate` (`latch`); reporter stays fail-open | False PASS merges a real bug. Missing key never fails Playwright; the gate is a separate step. Judge never says ignore alone | @@ -219,6 +221,12 @@ Worked placements (2026-09-18 topic:jev hour + prior archive): [decision-native-rag-skills](https://github.com/emergency-lee/decision-native-rag-skills): retrieve wide → decide → evidence set → LLM. Embeddings stay candidate generators. No universal benchmark (`notes.md` §55). +- **Files / URLs before the main agent reads** — + [jev-sift](https://github.com/kbhuw/jev-sift): classify first, + read selectively. Batch path/url/text → Jev; the main LLM opens + survivors. Uncertain/errors/truncation ≠ irrelevant. Topology A + MCP (LLM outer loop). Transport tests ≠ accuracy. No LICENSE this + pass (`notes.md` §56). Do not copy plugin how-to. - **Diff hunks before `git add`** — `ibrahemid/git-jev-stage`: one Choice per hunk (`include` / `exclude` / `mixed`); mixed and low-confidence stay unstaged; lines never split; staging is an exact patch after confirm. @@ -318,7 +326,11 @@ a *closed* tool catalog (the host executes; the calculator does the math). "First general-purpose System One agent" is a claim. [`nekowasabi/jev-routing`](https://github.com/nekowasabi/jev-routing) is a host adapter in front of an existing coding CLI (not topology B, not -MCP): it peels the catalog *before* the generator sees it. **Does not:** the decision model as the planner — neither inventing tools +MCP): it peels the catalog *before* the generator sees it. +[`kbhuw/jev-sift`](https://github.com/kbhuw/jev-sift) **is** topology A +MCP: the LLM still owns the outer loop; Jev is a tool that classifies +paths/URLs/text before the agent reads them (`notes.md` §56). Do not +merge with jev-routing. **Does not:** the decision model as the planner — neither inventing tools nor picking its own next tool in a loop (standing red flag, above and in `boundary-audit.md`); skipping schemas so the model "just knows"; treating a workflow AST as a proof. The outer loop stays with the LLM or @@ -495,6 +507,11 @@ decision-design card. Do not clone APIs from READMEs. | Bounded Pi supervisor | Skills / recovery / review / verify | Shadow default; never generates commands | jevons | | Judgment as language primitive | `chance` / `pick` / `rate` (Noul / Choice / Score) | English-as-config; stub backend; fail polarity per action (`rescue nil` at save ≠ spam gate) | hunch (Ruby library, not a new language; cousin of probably-lang) | | Decision-native RAG | Relevance / evidence / freshness / authority Nouls + Score | Evidence-set builder, conflict/temporal logic, provenance; embeddings generate candidates only | decision-native-rag-skills (no bundled harness; Hypothesis as a measured win) | +| Classify-first agent I/O | Relevance / typed questions on path/url/text | Hard envelope (50 / 60k / 2MB / public-IP); main LLM opens survivors; uncertain/errors/truncation ≠ irrelevant | jev-sift (MCP topology A; mocks ≠ accuracy; no LICENSE this pass) | +| Generative UI decide | Intent / layout Choice | Zod + deterministic compiler; model cannot add components | json-render + jev-agentworld-web-simulator | +| Robotics text-state | Choice on geometry-as-text | Code owns kinematics / Hz; two-call split; not pixels | MuJoCo showcase; jev-drone; Doom JSON. Drawing-pixel claim is not Archer | +| Draft-gate heartbeat | Quality Noul / Score | Fail-open / heartbeat on silence; missing verdict ≠ hold forever | jevable draft-gate fail mode; contrast Abide `<0.5` (edit proceeds) | +| Living class-pattern atlas | (not a model) | Cross-link exemplars; do not dump 342 titles | jevable.com (342 is *their* count; JSON-LD first page 36) | | Verbatim session recall | Noul "still live for this task?" | JSONL ledger; constraints/corrections always-keep; fail-open dump | carryforward (9×3 hint, not proof) | | Pre-exec tool gate | allow / block / review | Permissions, arg validation, transaction limits in code; timeout stops | toolgate (72-case synthetic, not independently annotated; Jev ≠ authorization) | | Healthcare S1 + S2 | NEWS2 remainder / med recon / inbox route | Code owns NEWS2, recon, routing; S2 blinded review | explore-typesafe-ai (synthetic FHIR; not clinically validated) | diff --git a/.agents/skills/augustus/references/question-design.md b/.agents/skills/augustus/references/question-design.md index a90b8fa..2d3ea17 100644 --- a/.agents/skills/augustus/references/question-design.md +++ b/.agents/skills/augustus/references/question-design.md @@ -44,7 +44,7 @@ request, and treat a stale pin as a prior, never a setting. ## Criteria shape - Criteria are an extension of the instruction and must ask the same thing, in the same direction (a Noul whose `true` side describes "no" performs worse). -- Choice options: contrastive `what` / `not_for` / short concrete `examples` (instances, not descriptions of instances). Add an `other` / `none-of-the-above` when the list may not cover inputs. Skipping that hatch is not a style nit: the model will pick a listed option at confidence 1.00, and no downstream gate will see a problem (`notes.md` §46). Request-shape lint (wellposed / `tenbin`) puts the hatch on the offered set; **training must confront it as a wrong alternative too**, with varied wording, or the model learns "this wording ⇒ pick it" ([kev](https://github.com/jaredpalmer/kev) first-run shortcut; dedicated `none_of_the_above` eval; `notes.md` §45 delta). +- Choice options: contrastive `what` / `not_for` / short concrete `examples` (instances, not descriptions of instances). Add an `other` / `none-of-the-above` when the list may not cover inputs. Skipping that hatch is not a style nit: the model will pick a listed option at confidence 1.00, and no downstream gate will see a problem (`notes.md` §46). Request-shape lint (wellposed / `tenbin`) puts the hatch on the offered set; **training must confront it as a wrong alternative too**, with varied wording, or the model learns "this wording ⇒ pick it" ([kev](https://github.com/jaredpalmer/kev) first-run shortcut; dedicated `none_of_the_above` eval; `notes.md` §45 delta). Corpus-scale cousin (maker claim, not re-run): SEO internal-link audit **584** placed / **139** refused because nothing honestly fit — Choice-with-`other` at catalog scale (`notes.md` §45–§46, §56). - Score levels (2–10): describe **situations**, one dimension each, each standing alone (Jev sees neither the level's number nor its neighbors — "worse than previous" means nothing). No numerals. Levels may be objects `{"summary", "signals"}`. Give a rare extreme its own level when code treats it differently. - Composite scoring: one Score per dimension, normalize by `len(criteria)-1`, weight and combine in code. Change policy by changing weights — never by rewriting questions. - Taxonomy walk: one Choice per tree level, walk in code; each option's value is its subtree (direct children + sample leaves); follow several branches when probabilities are close. diff --git a/.agents/skills/augustus/references/toolbox-mapping.md b/.agents/skills/augustus/references/toolbox-mapping.md index aceb4c2..9d07313 100644 --- a/.agents/skills/augustus/references/toolbox-mapping.md +++ b/.agents/skills/augustus/references/toolbox-mapping.md @@ -78,11 +78,11 @@ component; keep the rest of the method in code. | Experimental design: Harbor on/off routing | Same coding-agent task with routing on vs off; hidden verifier; cheaper unsolved is not a saving | **Empirical as a *shape* and one-run signal** (jev-gateway-bench; `notes.md` §51). Pair CI merge-gate with Harbor + rh-guard | | Discrete math: width vs depth | Fan out in width (parallel ≈ free), pay depth linearly; two-stage only when next options depend on an earlier answer | **Empirical recipe** (fan-out: 12.2× cheaper, 10× faster) | | Psychology: Kahneman | System 2 generates/proposes (LLM), System 1 discriminates (Jev); never the reverse. S1 keeps control; optional S2 is one-use advice and does not fly. Productized cascade: `conf ≥ τ` → S1 decides else S2 writes; routing fails open / safety fails closed; **routing accuracy unmeasured**; keyword fallback ≠ S1. S1 specialists + S2 coordinator is the same split (reification-labs/foreman is description-only Phoenix scaffold this pass — do not invent an Elixir API). Indexer: S1 GLiNER extract on the bulk, escalate LLM on the tail (10–50× unfilled). Healthcare: S1 remainder after NEWS2/code, S2 blinded review (explore-typesafe-ai; synthetic; not clinically validated) | **Empirical recipe** (mcts-agent role split; 2026-09-18 mixed-architecture discourse; jev-reflex-autonomy-lab as a control-loop shape, `notes.md` §46; [dual-process-ai](https://github.com/taro1985/dual-process-ai) as a business/life cascade, `notes.md` §49; s1-graphify-indexer, `notes.md` §51; explore-typesafe-ai, `notes.md` §55). **Route ≠ memory:** a cheap intent gate skips memory/tool *tours* on easy routes; memory still writes; complex still searches (jev-hermes, `notes.md` §48) | -| IR / cascades | Cheap relevance / irrelevance before an expensive ranker or generator | **Empirical recipe** (RAG cookbook; jevprune; git-jev-stage; LlamaIndex Jev rerank; jev-pruner stdout after size/format envelope, `notes.md` §53). **Empirical as architecture** (decision-native-rag-skills retrieve-wide→decide→evidence-set; Hypothesis as a measured win, `notes.md` §55) | -| IR / extractive keep-drop | Number candidates in code; model selects; copy verbatim; `redecide` thresholds on the log with no new calls. Compaction: same pointer job on tool results (Jev Noul/Score *or* GLiNER encoder). Computer-use: same pointer job on observed a11y/DOM controls (Jev *or* GLiNER2 *or* Cua-S1 option-attention). Stdout: same pointer job on Bash chunks after a hard envelope (Jev Noul) | **Empirical recipe** (testimonial-miner; jev-reviewer pointer-not-generator, `notes.md` §48; gliner25-compaction char-offset + fail-closed keep_full, `notes.md` §50; gliner2-ultrafast observe→score→act, `notes.md` §52; jev-pruner, `notes.md` §53; cua-s1 specialist form, source-only, `notes.md` §54) | +| IR / cascades | Cheap relevance / irrelevance before an expensive ranker or generator | **Empirical recipe** (RAG cookbook; jevprune; git-jev-stage; LlamaIndex Jev rerank; jev-pruner stdout after size/format envelope, `notes.md` §53). **Empirical as architecture** (decision-native-rag-skills retrieve-wide→decide→evidence-set; Hypothesis as a measured win, `notes.md` §55). **Empirical as README** (jev-sift classify-first MCP; mocks ≠ accuracy; errors/truncation ≠ irrelevant, `notes.md` §56) | +| IR / extractive keep-drop | Number candidates in code; model selects; copy verbatim; `redecide` thresholds on the log with no new calls. Compaction: same pointer job on tool results (Jev Noul/Score *or* GLiNER encoder). Computer-use: same pointer job on observed a11y/DOM controls (Jev *or* GLiNER2 *or* Cua-S1 option-attention). Stdout: same pointer job on Bash chunks after a hard envelope (Jev Noul). Showcase: score-among-observed (ads, on-screen posts) | **Empirical recipe** (testimonial-miner; jev-reviewer pointer-not-generator, `notes.md` §48; gliner25-compaction char-offset + fail-closed keep_full, `notes.md` §50; gliner2-ultrafast observe→score→act, `notes.md` §52; jev-pruner, `notes.md` §53; cua-s1 specialist form, source-only, `notes.md` §54). Atlas class pattern: Your Signal / Near Here (`notes.md` §56) | | Spec / lint | Project-defined semantic rules as predicates over a diff; linter owns hard rules. AST remainder: Tree-sitter units, then typed questions; do not execute scanned code. Plain-English PR check: one condition + min-confidence; fail-closed on error | **Empirical recipe** (jev-pref contract; Abide productized path — replay 93 sessions, edit precision ~26% / turn ~73% before tune, `notes.md` §47; JevLint file-level Noul; pi-warden; snifftest unsure-band; jevscan AST∩semantic, `tenbin` owns the lint skill, `notes.md` §48; if-ai, `notes.md` §51). jev-marshal is Watch / empty this pass | | Formal methods / DST / safety | Judgment triages counterexamples, failing seeds, and named-rule conformance; proof/MC/DST stay with their tools. Alloy finder ≠ Apalache BMC ≠ Quint run. DST trio: Antithesis hypervisor / Resonate HQ Lean+oracle+SDK (durable async) / PufferLib env+seed. Noul is a sensor, not a discharged PO. Semi-formal diagrams are vocabularies, not enforcers | **Hypothesis as product**, **Contract** as ownership (matching `mappings.md` §8 and `methods-catalog.md`; worked shape pi-warden — `formal-methods.md`, `formal-semi-formal.md`) | -| Decision analysis: VOI | Gather as an enumerated act; pay iff expected decision-loss drop > cost. Wake/resume: skip the LLM turn only if the judge answers and p is low (Horvitz); user-message / skip-limit already answer without a model. Selective memory: score a verbatim ledger; dump on failure; never judge the rules | **Hypothesis** as calculator (`mappings.md` §6; `mental-models.md`). wakegate 21/21 is smoke (`notes.md` §51). carryforward 9×3 is a hint (`notes.md` §55) | +| Decision analysis: VOI | Gather as an enumerated act; pay iff expected decision-loss drop > cost. Wake/resume: skip the LLM turn only if the judge answers and p is low (Horvitz); user-message / skip-limit already answer without a model. Selective memory: score a verbatim ledger; dump on failure; never judge the rules. Classify-first: pay for a full agent open iff relevance might change the act | **Hypothesis** as calculator (`mappings.md` §6; `mental-models.md`). wakegate 21/21 is smoke (`notes.md` §51). carryforward 9×3 is a hint (`notes.md` §55). jev-sift mocks ≠ accuracy (`notes.md` §56) | | Signal detection | Noul as evidence variable; criterion from costs and base rate; ROC/PR on your labels | **Hypothesis** for non-SWE plots (`mappings.md` §7) | | Safety engineering: STPA | Sensor ≠ constraint; table of unsafe control actions if the sensor lies | **Contract** as ownership (`mappings.md` §8; Leveson) | | Bandits / RL | Value from observed rewards only — Jev provides none; rejected without an environment. PufferLib Ocean is a trainer contract, not a baseline | **Rejected** (standing boundary) | diff --git a/CHANGELOG.md b/CHANGELOG.md index 76fe0e5..f561f46 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -296,6 +296,19 @@ Pegged against [`typesafe-ai/skills` v0.5.7](https://github.com/typesafe-ai/skil accuracy/lower cost not established), kazuhideoki/jev-search (recursive *file* search + fzf; **not** superagents-lab web search). No wrapper. No invented metrics. +- Hourly ~18:38 Boise 2026-09-18 / 00:38 UTC 2026-09-19 fold + (`research/notes.md` §56): Archer still Watch. Architecture + notes, not a plugin / showcase catalog. Classify-first MCP + ([jev-sift](https://github.com/kbhuw/jev-sift); batch path/url/text + → Jev without entering main agent context first; 50 / 60k / 2MB / + public-IP envelope; mocks ≠ accuracy; no LICENSE this pass; + topology A MCP, not jev-routing). Living applied-mappings atlas + ([jevable.com](https://jevable.com/); claimed 342 vs JSON-LD first + page 36; class patterns — intent columns, score-among-observed, + VOI gates, generative UI decide, robotics text-state, draft-gate + silence ≠ safer — not a 342-title hit list). Maker clocks stay + claims unless already a named receipt. No wrapper. No invented + metrics. - Effect-oriented loops (`notes.md` §28, `mappings.md` §19): Ward's ZIO client keeps Jev as the outer Choice and the handler as the effect. Not Effect.ts. GLiNER author: GLiNER2 "like jev" is GLiGuard diff --git a/README.md b/README.md index 48c9b6b..72399c1 100644 --- a/README.md +++ b/README.md @@ -60,14 +60,15 @@ never launder a Noul as a proof. evidence-preserving stdout prune (hard envelope then Noul); specialist S1 computer-use (Cua-S1 form-v0; plan ≠ execute; not TypeSafe Jev); judgment as a language primitive (hunch); decision-native RAG - (retrieve wide → decide → evidence set) + (retrieve wide → decide → evidence set); classify-first MCP + (jev-sift); draft-gate heartbeat; living class-pattern atlas - `.agents/skills/augustus/references/applied-mappings.md` — context sieve, - exact-text keep/drop (extractive / pointer-not-generator; char-offset compaction; observed a11y/DOM controls; Bash stdout prune; verbatim session ledger / carryforward), env triage (OpenSmoke + latch merge-gate), moderation/ranking (decision-native RAG evidence set), skill routing (route ≠ memory) + exact-text keep/drop (extractive / pointer-not-generator; char-offset compaction; observed a11y/DOM controls; Bash stdout prune; verbatim session ledger / carryforward; classify-first MCP / jev-sift), env triage (OpenSmoke + latch merge-gate), moderation/ranking (decision-native RAG evidence set; living class-pattern atlas), skill routing (route ≠ memory) - `.agents/skills/augustus/references/faq.md` — "just classification", stack replacement, Jev vs open head vs encoder vs LoRA vs constrained AR vs kev vs blackwood, wait-for-Archer, missing-other confident-wrong, soft project rules vs linter (Abide), extractive/pointer-not-generator, compaction summarize vs pointer, encoder vs Jev compaction, fail-closed keep_full, shadow-mode rollout, fail-open vs fail-closed wake vs CI gate, observe→score→act backend-agnostic, hybrid local decide + remote fill, DONE ≠ verified success, stdout prune vs session compaction, Cua-S1 vs TypeSafe Jev, plan ≠ execute / dry-run, local drop-in vs stub scorer, route ≠ memory, when-it-holds / extractable-from-state, decision-model vs constrained LLM, dual-process S1/S2, combinatorial grid ≠ extractive, GLiNER vs GLiClass vs CLIP, LLM-as-judge, in-engine vs CLI store, hard envelope (bitrate / planner), not-another-how-to, - uncalibrated local likelihoods ≠ Noul, decision-native RAG, cascade + uncalibrated local likelihoods ≠ Noul, decision-native RAG, classify-first MCP, living applied-mappings atlas / class patterns, draft-gate silence ≠ safer, robotics text-state vs pixels, cascade sign-flip / calibration theater, Precision PDF honest negative - `.agents/skills/augustus/references/mappings.md` — classical-method mappings with boundaries, counterexamples, acceptance tests (including diff --git a/docs/ecosystem.md b/docs/ecosystem.md index 6763f50..29f2362 100644 --- a/docs/ecosystem.md +++ b/docs/ecosystem.md @@ -18,6 +18,7 @@ weekdays. Jev is the densest public corpus, not the class monopoly. - **carlaiau/jev-reranking** — independent TREC DL2019 benchmark: zero-shot Jev best MAP 0.4748, nDCG@10 0.683 vs monoBERT 0.718; $0.76 per 41k pairs. - **superagents-lab/jev-search** — federated web search: Jev understands intent (query/sources/time-range), lanes fan out concurrently, Jev ranks results; merge by URL + engine agreement + rank. - **kazuhideoki/jev-search** — recursive *file* search + fzf. Not the federated web product. `notes.md` §55. +- **kbhuw/jev-sift** — classify-first MCP: batch path/url/text → Jev before the main agent reads. Topology A, not a host adapter. `notes.md` §56. ### Languages & runtimes - **probably-lang (southpolesteve)** — a programming language whose **loop conditions are Jev feelings**: `while draft feels "like a LinkedIn influencer post" { … }`. Judgment-state recordings give deterministic replay. @@ -199,7 +200,7 @@ Architecture notes, not a Driver / MCP catalog. `notes.md` §54. TypeSafe Jev is Architecture notes, not a CUDA/venv / gem / mcp / uv catalog. `notes.md` §55. TypeSafe Jev is the exemplar, not the monopoly. Archer still Watch. X MCP flap; `since_id` not advanced. - **Mintzs/jevify** — CUDA/PyTorch parallel Choice/Score/Noul *shape* on Qwen2.5-1.5B (`ora_decision_engine`). CUDA graphs, branch kernels, literal-label scoring. **Uncalibrated likelihoods ≠ Noul.** Independent of Distillation. Default refund workflow is not a validated policy. No LICENSE this pass. -- **emergency-lee/decision-native-rag-skills** — MIT. Retrieve wide → decide → evidence set → conflict resolve → reason only over kept evidence. Provider-agnostic. No bundled harness. No universal benchmark. Core Augustus RAG mental model. +- **emergency-lee/decision-native-rag-skills** — MIT. Retrieve wide → decide → evidence set → conflict resolve → reason only over kept evidence. Provider-agnostic. No bundled harness. No universal benchmark. Core Augustus RAG mental model. Classify-first MCP cousin: **kbhuw/jev-sift** (`notes.md` §56). - **Dharundp6/jev-carryforward** — MIT, 1★, npm `carryforward`. Verbatim session ledger; Jev scores recall; rules never judged; fail-open dump. 9×3 hint, not proof. - **carldaws/hunch** — MIT. Ruby `chance`/`pick`/`rate`; English-as-config; `rescue nil` fail-open at save. Cousin of probably-lang (library, not a new language). - **si618/explore-typesafe-ai** — FHIR S1 (Jev) + Claude S2 on 100 synthetic Synthea patients. Labels first. 60 requests / 403 judgments. **Not clinically validated.** License not in API this pass. @@ -207,6 +208,13 @@ Architecture notes, not a CUDA/venv / gem / mcp / uv catalog. `notes.md` §55. T - **SargeDev/jev-gate-student-b** — light delta only; HF card unchanged (MAE 0.187 / Pearson 0.791 / 90% n=60; fail-open; teacher-copy). - MED: **fdemir/toolgate** (pre-exec allow/block/review; Jev not authorization; 72-case synthetic); **masa-med-ai/typesafe-screening-mcp** (PubMed include/maybe/exclude; 326 hits ~17s ~$0.014; screening aid); **laurentfabre/databricks-jev-pdf-lab** (honest negative; no OSS license); **yannip1234/codex-jev** (extractive Codex compression; 185→44 estimated tokens is an integration demo; equal accuracy/lower cost not established); **kazuhideoki/jev-search** (recursive *file* search + fzf; **not** superagents-lab federated web search). +### Hourly ~18:38 Boise 2026-09-18 / 00:38 UTC 2026-09-19 (classify-first MCP + living applied-mappings atlas) + +Architecture notes, not a plugin / showcase catalog. `notes.md` §56. TypeSafe Jev is the exemplar, not the monopoly. Archer still Watch. No invented metrics. No 342-title dump. + +- **kbhuw/jev-sift** — classify first, read selectively. Batch path / public URL / inline text → Jev relevance or 1–8 typed questions. Content to Jev without entering main agent context first. Envelope (theirs): 50 items, 60k char, 2 MB / 20 s, public-IP only, no JS/cookies/login, PDFs unsupported. Uncertain/errors/truncation ≠ irrelevant. Transport tests (mocks) ≠ accuracy. No LICENSE this pass. Same retrieve-wide → decide → evidence-set family as decision-native-rag-skills. Topology A MCP; **not** nekowasabi/jev-routing (host adapter). +- **jevable.com** — living applied-mappings atlas. Claimed **342** curated projects; JSON-LD first page **36**. Categories: Agents, Browser extensions, Creative tools, Data & research, Developer tools, Experiments, Finance, Games, Marketing, Productivity, Robotics. No public API this pass. Class patterns: intent columns, score-among-observed, VOI gates, generative UI decide, robotics text-state, draft-gate silence ≠ safer. Maker clocks stay claims unless already a named receipt. + See `references/mixed-architecture.md` in the skill. Class-level family choice: `references/judgment-class.md`. Proof vs judgment (Alloy vs Apalache; DST trio Antithesis / Resonate HQ / PufferLib): diff --git a/research/archive/findings.md b/research/archive/findings.md index 7c1ee69..d67dea5 100644 --- a/research/archive/findings.md +++ b/research/archive/findings.md @@ -1132,6 +1132,54 @@ encoder-with-labels still wins; (cw) serving-path ≠ model-speed; toolgate product ≠ ndolinschi vocab; (cz) Precision PDF honest negative is a result. +## Batch #40 (2026-09-19 ~00:38 UTC / ~18:38 Boise 2026-09-18) — classify-first MCP + living applied-mappings atlas + +Note: `research/notes.md` §56. Docs-only. Folded into PR #2. Not a +second PR. Archer still Watch. No invented metrics. Do not re-fold +§50–§55. + +- **kbhuw/jev-sift (Empirical as README / schema; Hypothesis as a + measured win).** JavaScript. Created 2026-09-18T00:13:31Z; 10★ + this pass; **no LICENSE file this pass.** Plugin + `0.2.0+codex.20260918200547`; package 0.2.0; author Kush + Bhuwalka. Classify first, read selectively: batch path / public + URL / inline text → Jev relevance or 1–8 typed questions. Direct + `POST /v1/systemone` `jev-latest`. Envelope (theirs): 50 items, + 60k char, 2 MB / 20 s, public-IP only, 3 redirects, no + JS/cookies/login, PDFs unsupported. Uncertain/errors/truncation ≠ + irrelevant. Transport tests (mocks) ≠ accuracy. Same + retrieve-wide → decide → evidence-set family as + decision-native-rag-skills. Topology A MCP; not jev-routing + (host adapter). Cousins: typesafe-screening-mcp, + kazuhideoki/jev-search, jev-pruner, carryforward. +- **jevable.com (Empirical as public showcase; 342 is *their* + count).** Independent curated atlas (Nikunj / `@nikunj` in + JSON-LD). HTTP 200 Railway. Claimed **342**; JSON-LD first page + **36**; `pageSize` 36. Categories: Agents, Browser extensions, + Creative tools, Data & research, Developer tools, Experiments, + Finance, Games, Marketing, Productivity, Robotics. No public API + this pass. Class patterns, not a 342-title dump: (1) intent + columns → jevpandas/jevframe / dabit3 formulas; (2) + score-among-observed → jev-ultrafast / gliner2-ultrafast / + solari / cua-s1 + Your Signal/Near Here (do not merge 7s/$0.0039 + with 12.20s; 100× is a claim); (3) VOI gates → tamara + compaction / jev-pruner / gliner25 / routeKit / Gmail embeddings- + first / jev-sift; (4) generative UI decide → json-render + + jev-agentworld-web-simulator; (5) robotics text-state MuJoCo + geometry-as-text / MOSS / jev-drone / Doom JSON (drawing-pixel + claim ≠ Archer); (6) draft-gate silence-as-safer needs fail-open + / heartbeat vs Abide `<0.5` (edit proceeds). Confirm-don't- + invent: Higgsfield claim, jev-trader, Cambium, SEO 584/139 + `other`, snacks 3000/28s/$0.11 claim, ai-cli/hunch, Manhattan/ + Sudoku ≠ replace A*. + +Cross-repo addition: (da) classify-first MCP is retrieve-wide → +decide → evidence-set on agent I/O; (db) errors/truncation ≠ +irrelevant; (dc) topology A MCP ≠ host-adapter routing; (dd) +living atlas extracts class patterns, not a 342-row dump; (de) +draft-gate silence ≠ safer (heartbeat); (df) robotics text-state ≠ +pixels; (dg) 342 is their count / JSON-LD 36 is page 1. + diff --git a/research/notes.md b/research/notes.md index 3a882b8..c297f66 100644 --- a/research/notes.md +++ b/research/notes.md @@ -4587,3 +4587,265 @@ file-search vs web-search), §6 (carryforward VOI), §18 explore-typesafe-ai; pdf-lab negative); `faq.md`; `mental-models.md`; `methods-catalog.md`; `toolbox-mapping.md`; `agent-self-assessment.md`. No wrapper. + +## 56. Classify-first MCP + living applied-mappings atlas (2026-09-19 ~00:38 UTC / ~18:38 Boise) + +America/Boise ~18:38 = 2026-09-19T00:38Z. Docs-only fold into +open PR #2 (`cursor/augustus-store-envelope-00b4`). Not a +competing PR. Archer 27B drop still **WATCH**. Identity lock vs +`typesafe-ai` / `tenbin` / `decision-first` holds. No wrapper, +no MCP/npm how-to, no copied ports or key-file paths. No +invented metrics. Do not re-fold §50–§55. + +Two HIGH **usage / architecture** signals: a portable classify- +first MCP (same retrieve-wide → decide → evidence-set family as +decision-native RAG), and a living applied-mappings *atlas* +(class patterns from a curated showcase — not a 342-title hit +list). TypeSafe Jev is the documented exemplar, not the +monopoly. Augustus stays family-first. + +### HIGH + +1. **[`kbhuw/jev-sift`](https://github.com/kbhuw/jev-sift)** + (JavaScript; created 2026-09-18T00:13:31Z; 10★ this pass; + **no LICENSE file this pass — do not invent**; GitHub + `license` null). Portable agent plugin + stdio MCP: + **classify first, read selectively.** Tagline: let Jev decide + what the agent should look at next. Pass a query and a batch + of file paths, public webpage URLs, or inline text; the tool + loads content, sends it **directly to Jev**, and returns + compact relevance probabilities. The main agent only opens + items worth a closer look. Tool descriptions work too: score + a supplied description without executing the tool or + predicting an unseen result. Direct TypeSafe + `POST /v1/systemone` with `jev-latest`. Key via `JEV_API_KEY` + / `TYPESAFE_API_KEY` or a private key file (README names + `~/.config/jev-sift/api-key`; do not copy the path). Plugin + `0.2.0+codex.20260918200547`; package `0.2.0`; author Kush + Bhuwalka. Checked-in `dist/server.mjs` includes dependencies. + Core `classify` is framework-independent (inject `evaluate` / + `readText` / `readUrl`). Earlier generic prototype used + `CLASSIFY_*` / chat-completions — those settings are gone. + + **README / schema envelope (theirs; not a how-to):** up to + **50** items; **1–8** typed questions (boolean → Noul; + Choice 2–12 options; Score 2–10 levels) *or* a `query` + shorthand that returns `answers.relevant.probability`. + Exactly one of `text` / `path` / `url` per item; unique ids. + Concurrency **1–8**. File and extracted page text capped at + **60,000** JavaScript characters and flagged if truncated. + Web: **2 MB** / **20 s**, public-IP only (including redirect + targets; DNS pinned to the connection), HTTP(S) ports 80/443, + up to **three** redirects, no JavaScript, no login, no + browser cookies, PDFs / private-network / other binaries + unsupported. A successful fetch of login/challenge HTML is + not the intended article. Webpage fetches receive no Jev + credentials. Plugin does not persist source content or + results. Results retain input order and include per-item + errors, resolved source URLs, truncation flags, model id, + summed input-token usage. **Uncertain items should get a + closer look; errors and truncation are not evidence that an + item is irrelevant.** Tests cover mapping, ordering, + failures, cancellation, truncation, file boundaries, public- + URL validation, redirects, HTML extraction, and an isolated + bundled MCP exchange — **mocks, not an accuracy benchmark.** + Live smoke verifies connectivity only. Model quality on *your* + task still needs evaluation. + + **Four load-bearing mental models:** + + 1. **Retrieve-wide → decide → evidence-set (agent I/O).** + Same family as + [decision-native-rag-skills](https://github.com/emergency-lee/decision-native-rag-skills) + (`notes.md` §55): content goes to the judge **without + entering main agent context first** (paths/URLs). Inline + text the agent already read cannot recover that cost. The + main LLM reasons only over items worth a closer look. + Embeddings/file lists stay candidate generators. + 2. **VOI / context economics.** Uncertain → closer look. + Errors/truncation ≠ irrelevant. Pay for a full read iff + the relevance (or typed question) says it might change the + act. Webpage fetch still costs bandwidth — this saves the + *agent's* read, not the download. + 3. **Hard envelope on I/O.** 60k char, 2 MB / 20 s, public-IP + only, no JS/cookies/login, PDFs unsupported. Same sandwich + family as jev-pruner (size/format then Noul) and bitrate- + advisor (soft propose, code clamps). + 4. **Do not copy marketplace / `mcpServers` / key-file how- + to.** Transport tests ≠ accuracy. + + **Cousins, do not merge.** + [typesafe-screening-mcp](https://github.com/masa-med-ai/typesafe-screening-mcp) + — abstracts never enter the LLM conversation; include/maybe/ + exclude in code (`notes.md` §55). + [kazuhideoki/jev-search](https://github.com/kazuhideoki/jev-search) + — recursive *file* search + fzf, not this MCP (`notes.md` + §55). [jev-pruner](https://github.com/tamaratran/jev-pruner) + — prune Bash *after* it ran; this tool screens *before* the + agent reads. [carryforward](https://github.com/Dharundp6/jev-carryforward) + — scored recall over a ledger you already hold. + [jev-routing](https://github.com/nekowasabi/jev-routing) is a + **host adapter, not MCP**. Dual-orchestration topology A + (Jev-as-tool); the LLM still owns the outer loop. + + **Placement.** Context sieve (`applied-mappings.md` §1) + + retrieval (`mappings.md` §4) + mixed architecture (MCP as + topology A). Pillar: VOI + selective classification. Hole: + sieve / rank / gather. Family: closed decision API. Fail-open + on "open this file" (false drop loses evidence); truncation/ + error ≠ irrelevant. Eval path: none published (transport + tests). **Empirical** as README / schema behavior. + **Hypothesis** that classify-first beats dump-into-context on + *your* agent. Cards: `applied-mappings.md` §1 (primary); + `mappings.md` §4 / §6; `mixed-architecture.md`; `faq.md`. + No wrapper. + +2. **[jevable.com](https://jevable.com/)** — "Discover what + people build with Jev." Independent curated showcase (creator + Nikunj / `@nikunj` in site JSON-LD). **Claim (theirs, this + pass):** **342** curated projects (``, + `#result-count` sr-only, board-data `"total":342`). Homepage + JSON-LD `ItemList.numberOfItems` is **36** (first page / + featured). Board-data `"pageSize":36`, `"nextOffset":36`, + `"sort":"curated"`. Categories in the filter: Agents, Browser + extensions, Creative tools, Data & research, Developer tools, + Experiments, Finance, Games, Marketing, Productivity, + Robotics. HTTP 200 this pass (Railway). **No public API + discovered this pass.** Watch: refresh the claimed count and + category list; do not treat 342 as an Augustus census. + + **What it is for Augustus.** Living **applied-mappings + corpus**: how people apply judgment tools. Primary usage / + application atlas — **not** a 342-title hit list, **not** a + multimodal substrate, **not** a model. Prefer **class + patterns**. Maker demos are claims unless already measured in + notes. Cross-link exemplars already folded; do not invent + repos or clocks. + + **Class patterns (load-bearing; not a gallery dump):** + + 1. **Intent columns.** Spreadsheets recalculate numbers, not + meaning. Type a heading ("Urgency"); each row is scored + (~100 ms is **their** demo claim). Same *hole* as + dataframe semantic columns ([jevpandas](https://github.com/yalindogusahin/jevpandas) + / [jevframe](https://github.com/ktaletsk/jevframe), + `notes.md` §46 / §48) and the launch-week + `dabit3/jev-experiments` JUDGE/SCORE/CHOOSE formulas + (`docs/ecosystem.md`). Predictive app launcher (heading / + keystroke → intent, not alias/fuzzy/habit) is the same + pattern on a catalog. Pillar: MCDA. Hole: perceive / + rank. Weights and vetoes stay in code. + 2. **Score-among-observed.** Candidates already on the page + (a11y/DOM, action space, on-screen posts); the model + scores; **code** clicks / filters / removes. Showcase: + Browser Use Ultrafast (Flights **7 s / $0.0039** is the + *same* demo already in notes as ~7.1 s — do not merge + clocks with gliner2-ultrafast 12.20 s); ad blocker + (DOM element → ad/non-ad); Notte (new action space every + step); computer-use "100×" is a **claim**. Already + folded: jev-ultrafast / gliner2-ultrafast / solari-reflex + / cua-s1 / laya-mind2web (`notes.md` §4, §48, §52, §54). + **Your Signal** (Fabio Angela): score posts *already on + screen*, apply rules locally, reversible, BYOK, no + telemetry — same judge-once / re-policy family as Near + Here firehose (`applied-mappings.md` §4). + 3. **VOI gates.** Instant compaction (tamara: score tool + calls, drop irrelevant — **same job** as + [fast-jev-compaction](https://github.com/tamaratran/fast-jev-compaction) + / [jev-pruner](https://github.com/tamaratran/jev-pruner) / + [gliner25-compaction](https://github.com/m-newhauser/gliner25-compaction); + pointer, not summarizer). Prompt-difficulty classifier + before send (fast-mode offer) is a **route** gate, cousin + of [routeKit](https://github.com/rajdhakad9826/routeKit) + (`notes.md` §33) — Jev does not pick the LLM; code + offers. Gmail intent search: embeddings pull first, then + judge — decision-native RAG on mail (`notes.md` §55). + [jev-sift](https://github.com/kbhuw/jev-sift) (this + section) is the MCP of the same VOI: classify first. + 4. **Generative UI decide.** json-render + Jev: *your* + components, actions, design system; the model decides; + render is milliseconds. Cousin already folded: + [jev-agentworld-web-simulator](https://github.com/knowlet/jev-agentworld-web-simulator) + — Jev Choice for intent / layout; generator writes + documents; Zod + **deterministic** compiler emits json- + render spec; model cannot add components (`notes.md` + §48). Decision for control, generator for content. + 5. **Robotics text-state (not pixels).** MuJoCo robot-arm: + Jev does not accept images; it gets simplified geometry + and contacts **as text**; two-call split (what to do, + then how to move). MOSS: Jev picks the target; the + robot picks up the litter. Same observe→decide→act + *job* as computer-use, different body. Cousins: + [jev-drone](https://github.com/RomanSlack/jev-drone) + (500/50 Hz code, Jev advisory 2.5 Hz); Doom demo fed + structured JSON, not raw pixels (`notes.md` §1 / §4). + **Flag:** "drawing, one decision at a time" *claims* + pixel-parallel prediction — contrast the MuJoCo honesty + (text-state). Perception-then-judgment vs shared + multimodal (`notes.md` §39); Archer still Watch. + 6. **Draft-gate fail mode: silence as "safer".** Jev sat + between GPT and the user, killing drafts that broke + rules. Then Jev did not answer. The agent treated + **silence as safer** and stopped sending anything. Took a + second agent to unstick. **Your checker needs a + fail-open / heartbeat** when the judge is down — + missing verdict is not a block and is not a pass. Name + the irreversible act: *withholding the draft* is fail- + closed-by-absence. Contrast Abide `<0.5` silence (linter + stays quiet; the *edit proceeds*, `notes.md` §47) and + carryforward dump / wakegate wake-on-error. Stuck- + detector / done-check (`agent-self-assessment.md`) must + not treat no-answer as "hold forever." + + **Other patterns already in notes (confirm, don't invent):** + Higgsfield auto-routing is a **claim** (`notes.md` §44; + routeKit hole). Trading / signals → decisions: + [jev-trader](https://github.com/jarrodwatts/jev-trader). + Cambium first-class provider: keep in code what can be in + code (mixed-architecture slogan, not a new family). SEO + internal-link audit (maker claim: 45.1 s, 586 pages, **584** + links placed, **139** refused because nothing honestly fit, + $0.21) is Choice-with-`other` at corpus scale — wellposed / + kev NOTA (`notes.md` §45–§46); not re-run. Snack MCDA + (maker claim: 3,000 kids' snacks, multiple criteria, 28 s, + $0.11) is mapping §1 at catalog scale. ai-cli (yes/no / + choose / score from the shell) is a language-primitive + cousin of [hunch](https://github.com/carldaws/hunch) + (`notes.md` §55). Manhattan pathfinding / Sudoku playground: + **algorithm stays yours**; do not replace A* or a solver + with a Noul (`mappings.md` §9; ARC-AGI combinatorial ≠ + extractive, `notes.md` §49). + + **Jev-omni.** Archive the site snapshot as a community usage + atlas. Not multimodal substrate. Flag demos that claim + pixels vs text-state (drawing vs MuJoCo). Refresh count / + categories on hourly watch if useful. + + **Placement.** Applied-mappings atlas (this file + + `applied-mappings.md` / `mixed-architecture.md` gallery), + not a new species. **Empirical** as the public showcase + (342 is *their* count; we did not enumerate titles). + Maker clocks stay **claims** unless already a named receipt. + Cards: `applied-mappings.md`; `mappings.md` §1 / §4 / §6 / + §9; `mixed-architecture.md`; `faq.md`; `mental-models.md`; + `agent-self-assessment.md`; `question-design.md`. No + wrapper. No 342-row dump. + +### Omni / Jev-omni / Archer + +Still **WATCH**. No Hub weights. Showcase drawing-pixel claim +is **not** that drop. MuJoCo text-state is the honest robotics +posture until a multimodal decide ships (blackwood-rlcd is +screenshot-in, not arm-in). + +### Cross-links + +Cards: `applied-mappings.md` §1 (jev-sift classify-first), +§2 (score-among-observed atlas), §4 (intent search / Your +Signal); `mappings.md` §1 (intent columns / snack MCDA), §4 +(RAG family), §6 (VOI gates), §9 (robotics text-state; do not +replace A*); `mixed-architecture.md` (topology A MCP; generative +UI decide; draft-gate heartbeat); `faq.md`; `mental-models.md`; +`question-design.md` (SEO 139-refused as `other`); +`agent-self-assessment.md` (silence ≠ safer); `methods-catalog.md`; +`toolbox-mapping.md`. No wrapper. diff --git a/research/refresh-log.md b/research/refresh-log.md index 71b5745..1013fe8 100644 --- a/research/refresh-log.md +++ b/research/refresh-log.md @@ -635,6 +635,28 @@ ecosystem, CHANGELOG, README. - notes.md §55; sources.json; findings.md batch #39. No wrapper. +## 2026-09-19 00:38 UTC — classify-first MCP + living applied-mappings atlas (~18:38 Boise 2026-09-18) + +- Folded into open PR #2 (`cursor/augustus-store-envelope-00b4`). + Docs-only. Not a competing PR. Archer 27B drop still **WATCH**. + No invented metrics. No wrapper. Do not re-fold §50–§55. +- HIGH: [kbhuw/jev-sift](https://github.com/kbhuw/jev-sift) + classify-first MCP/plugin; batch path/url/text → Jev; content + without entering main agent context first; 50 / 60k / 2MB / + public-IP envelope; mocks ≠ accuracy; no LICENSE. Same family as + decision-native-rag-skills. Topology A MCP, not jev-routing. + [jevable.com](https://jevable.com/) living applied-mappings atlas: + claimed 342 vs JSON-LD first page 36; class patterns (intent + columns, score-among-observed, VOI gates, generative UI decide, + robotics text-state, draft-gate silence ≠ safer). Not a hit list. + Maker clocks stay claims unless already a named receipt. +- Cards: SKILL.md, applied-mappings §1/§2/§4, mappings §1/§4/§6/§9, + mixed-architecture (topology A; prefilter polarity; gallery), + faq, mental-models, methods-catalog, toolbox, + agent-self-assessment, question-design, ecosystem, CHANGELOG, + README. +- notes.md §56; sources.json; findings.md batch #40. No wrapper. + diff --git a/research/sources.json b/research/sources.json index 4a4592c..81ae88d 100644 --- a/research/sources.json +++ b/research/sources.json @@ -1,6 +1,6 @@ { "refresh_cadence": "hourly", - "retrieved": "2026-09-18T23:54Z", + "retrieved": "2026-09-19T00:38Z", "sources": [ { "kind": "docs", @@ -1676,6 +1676,18 @@ "title": "kazuhideoki/jev-search", "url": "https://github.com/kazuhideoki/jev-search", "note": "Created 2026-09-18T23:43:52Z; 0 stars; no LICENSE file this pass. Recursive semantic file search + fzf. Not superagents-lab/jev-search (federated web). Max-over-chunks is not a calibrated whole-file probability. notes.md \u00a755." + }, + { + "kind": "github", + "title": "kbhuw/jev-sift", + "url": "https://github.com/kbhuw/jev-sift", + "note": "JavaScript. Created 2026-09-18T00:13:31Z; 10 stars this pass; license null / no LICENSE file. Plugin 0.2.0+codex.20260918200547; package 0.2.0; author Kush Bhuwalka. Classify first, read selectively: batch path/url/text \u2192 Jev relevance or 1\u20138 typed questions. Direct POST /v1/systemone jev-latest. Envelope (theirs): 50 items, 60k char, 2 MB / 20 s, public-IP only, 3 redirects, no JS/cookies/login, PDFs unsupported. Uncertain/errors/truncation \u2260 irrelevant. Transport tests (mocks) \u2260 accuracy. Same retrieve-wide\u2192decide\u2192evidence-set family as decision-native-rag-skills. Topology A MCP, not jev-routing host adapter. notes.md \u00a756." + }, + { + "kind": "web", + "title": "jevable.com", + "url": "https://jevable.com/", + "note": "Independent curated showcase (creator Nikunj / @nikunj in JSON-LD). HTTP 200 Railway this pass. Claim (theirs): 342 curated projects (meta description, #result-count, board-data total). JSON-LD ItemList.numberOfItems 36 (first page); pageSize 36; nextOffset 36; sort curated. Categories: Agents, Browser extensions, Creative tools, Data & research, Developer tools, Experiments, Finance, Games, Marketing, Productivity, Robotics. No public API discovered this pass. Treat as living applied-mappings atlas: class patterns, not a 342-title dump. notes.md \u00a756." } ] } From 1cd1f10bb1ddb35490c96f6d8df49ac9370657d8 Mon Sep 17 00:00:00 2001 From: Cursor Agent Date: Sat, 19 Sep 2026 00:53:33 +0000 Subject: [PATCH 15/43] Fold Stagehand experimental Jev pick-and-copy harness MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Docs-only: draft stack #2951–#2955 as harness productization of observe→score-among-candidates→code-acts. Extract judge + pick-and-copy uses their 37/75 ~0.5s vs 4.37s card; pick is a fast path, not a replacement. No invented metrics; Archer still Watch. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com> --- .agents/skills/augustus/SKILL.md | 16 +- .../references/agent-self-assessment.md | 4 + .../augustus/references/applied-mappings.md | 16 ++ .agents/skills/augustus/references/faq.md | 17 +- .../augustus/references/judgment-class.md | 15 +- .../skills/augustus/references/mappings.md | 18 ++ .../augustus/references/mental-models.md | 5 +- .../augustus/references/methods-catalog.md | 4 +- .../augustus/references/mixed-architecture.md | 14 +- .../augustus/references/question-design.md | 2 +- .../augustus/references/toolbox-mapping.md | 2 +- .../skills/augustus/references/validation.md | 13 +- CHANGELOG.md | 12 ++ README.md | 8 +- docs/ecosystem.md | 8 +- research/archive/findings.md | 29 +++ research/notes.md | 196 ++++++++++++++++++ research/refresh-log.md | 20 ++ research/sources.json | 38 +++- 19 files changed, 411 insertions(+), 26 deletions(-) diff --git a/.agents/skills/augustus/SKILL.md b/.agents/skills/augustus/SKILL.md index 10c196b..7620527 100644 --- a/.agents/skills/augustus/SKILL.md +++ b/.agents/skills/augustus/SKILL.md @@ -1,6 +1,6 @@ --- name: augustus -description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill, open multimodal RLCD / blackwood), announced open decision-model (Watch — still not landed), constrained-AR (TypeAR, pcdServer, decision-token LoRA), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"wait for Archer vs open multimodal RLCD\", \"screenshot/DOM candidates → typed Choice\", \"eval path\", \"jevals\", \"Harbor taskset\", \"shared bake-off ECE/NLL/Brier\", \"LLM-as-judge is not the System One score\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge (allowlist *proves*; fail-open cannot block), \"missing other → confident wrong Choice\", \"lint the Jev request\", \"training confronts Choice other / none-of-the-above\", \"S1 reflex keeps control / optional S2 one-use advice\", \"soft AGENTS.md rules vs linter\" (Abide / jev-pref), \"edit-phase vs turn-phase observation window\", banded confidence fail-open preference lint, \"extractive selection / pointer-not-generator\", \"encoder GLiNER compaction vs Jev Score compaction (same job; pointer not summarizer)\", \"fail-closed keep_full under mutation envelope\", \"CI flaky-vs-real merge gate\", \"fail-open VOI wake/resume (Horvitz)\", \"claim/evidence Stop integrity\", \"S1 extract + escalate-S2 indexer\", \"Harbor on/off routing\", \"policy-as-judgment PR marshal\", \"shadow-mode compaction rollout\", \"Jev Ultrafast vs GLiNER Ultrafast (observe-score-act backend-agnostic)\", \"hybrid local decide + remote fill\", \"DONE ≠ verified success\", \"observed a11y/DOM candidates vs screenshot multimodal\", \"evidence-preserving stdout prune (not summarize)\", \"hard token/format envelope then soft Noul\", \"fail-safe keep original on prune failure\", \"stdout prune vs session compaction\", \"specialist S1 computer-use (Cua-S1 form-v0; not TypeSafe Jev)\", \"plan ≠ execute / dry-run default\", \"observed-element option head (fill/check/click/skip)\", \"local /v1/systemone drop-in (stub until hf scorer)\", \"dataframe-native semantic columns\", \"route≠memory\", \"advisory sidecar receipts\", \"structure induction over bags\", \"AST ∩ semantic lint\", \"extractable-from-state / retrieve first\", \"decision-model vs constrained-LLM bake-off\", \"dual-process S1 decide / S2 generate\", \"combinatorial grid ≠ extractive\", TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, \"uncalibrated local likelihoods vs Noul / CUDA replica\", \"decision-native RAG retrieve wide then decide then evidence set\", \"classify-first MCP / read selectively\", \"living applied-mappings atlas / class patterns not a 342 hit list\", \"draft-gate silence as safer / heartbeat\", \"robotics text-state vs pixels\", \"verbatim session ledger / scored recall\", \"judgment as language primitive / English-as-config\", \"pre-registered AMBIGUOUS eval / cascade sign-flip\", \"healthcare Harbor-shaped S1+S2\", \"pre-exec tool gate allow/block/review\", or \"formally verify with Jev\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." +description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill, open multimodal RLCD / blackwood), announced open decision-model (Watch — still not landed), constrained-AR (TypeAR, pcdServer, decision-token LoRA), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"wait for Archer vs open multimodal RLCD\", \"screenshot/DOM candidates → typed Choice\", \"eval path\", \"jevals\", \"Harbor taskset\", \"shared bake-off ECE/NLL/Brier\", \"LLM-as-judge is not the System One score\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge (allowlist *proves*; fail-open cannot block), \"missing other → confident wrong Choice\", \"lint the Jev request\", \"training confronts Choice other / none-of-the-above\", \"S1 reflex keeps control / optional S2 one-use advice\", \"soft AGENTS.md rules vs linter\" (Abide / jev-pref), \"edit-phase vs turn-phase observation window\", banded confidence fail-open preference lint, \"extractive selection / pointer-not-generator\", \"encoder GLiNER compaction vs Jev Score compaction (same job; pointer not summarizer)\", \"fail-closed keep_full under mutation envelope\", \"CI flaky-vs-real merge gate\", \"fail-open VOI wake/resume (Horvitz)\", \"claim/evidence Stop integrity\", \"S1 extract + escalate-S2 indexer\", \"Harbor on/off routing\", \"policy-as-judgment PR marshal\", \"shadow-mode compaction rollout\", \"Jev Ultrafast vs GLiNER Ultrafast (observe-score-act backend-agnostic)\", \"hybrid local decide + remote fill\", \"DONE ≠ verified success\", \"observed a11y/DOM candidates vs screenshot multimodal\", \"evidence-preserving stdout prune (not summarize)\", \"hard token/format envelope then soft Noul\", \"fail-safe keep original on prune failure\", \"stdout prune vs session compaction\", \"specialist S1 computer-use (Cua-S1 form-v0; not TypeSafe Jev)\", \"plan ≠ execute / dry-run default\", \"observed-element option head (fill/check/click/skip)\", \"local /v1/systemone drop-in (stub until hf scorer)\", \"dataframe-native semantic columns\", \"route≠memory\", \"advisory sidecar receipts\", \"structure induction over bags\", \"AST ∩ semantic lint\", \"extractable-from-state / retrieve first\", \"decision-model vs constrained-LLM bake-off\", \"dual-process S1 decide / S2 generate\", \"combinatorial grid ≠ extractive\", TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, \"uncalibrated local likelihoods vs Noul / CUDA replica\", \"decision-native RAG retrieve wide then decide then evidence set\", \"classify-first MCP / read selectively\", \"living applied-mappings atlas / class patterns not a 342 hit list\", \"draft-gate silence as safer / heartbeat\", \"robotics text-state vs pixels\", \"verbatim session ledger / scored recall\", \"judgment as language primitive / English-as-config\", \"pre-registered AMBIGUOUS eval / cascade sign-flip\", \"healthcare Harbor-shaped S1+S2\", \"pre-exec tool gate allow/block/review\", \"Stagehand extract pick-and-copy / judge\", \"harness observe-score-act productization\", or \"formally verify with Jev\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." license: MIT metadata: version: 0.3.0 @@ -54,7 +54,7 @@ classical method you already trust, substitute it, classify the win "paraphrase brittleness", "allowlist then judge", "TOCTOU-of-Noul", "Jev inside the database / sqlite-jev", "Jev picks bitrate / join order / the model", "wait for Archer", "lint the request / missing - other", "training confronts Choice other / none-of-the-above", "soft AGENTS.md rules vs the linter", "screenshot Choice / omni System One", "extractive quotes / pointer not generator", "compaction summarize vs pointer", "encoder vs Jev compaction backend", "shadow-mode compaction rollout", "CI flaky-vs-real merge gate", "fail-open VOI wake/resume", "claim vs session evidence", "S1 indexer escalate-S2", "Harbor on/off routing", "fail-open vs fail-closed wake vs CI gate", "encoder vs Jev computer-use backend", "hybrid local decide + remote fill", "DONE vs verified success", "stdout prune vs session compaction", "hard envelope then Noul prune", "Cua-S1 vs TypeSafe Jev", "plan vs execute dry-run", "specialist computer-use vs general agent", "local drop-in vs stub scorer", "route vs memory", "when does it hold / extractable from state", "decision model vs constrained LLM", "dual-process S1/S2", "combinatorial grid vs extractive", "uncalibrated local likelihoods", "decision-native RAG", "classify-first / read selectively", "living applied-mappings atlas / class patterns", "silence as safer / draft-gate heartbeat", "robotics text-state vs pixels", "verbatim ledger vs summary", "judgment as language primitive", or "cascade sign-flip / calibration theater": + other", "training confronts Choice other / none-of-the-above", "soft AGENTS.md rules vs the linter", "screenshot Choice / omni System One", "extractive quotes / pointer not generator", "compaction summarize vs pointer", "encoder vs Jev compaction backend", "shadow-mode compaction rollout", "CI flaky-vs-real merge gate", "fail-open VOI wake/resume", "claim vs session evidence", "S1 indexer escalate-S2", "Harbor on/off routing", "fail-open vs fail-closed wake vs CI gate", "encoder vs Jev computer-use backend", "hybrid local decide + remote fill", "DONE vs verified success", "stdout prune vs session compaction", "hard envelope then Noul prune", "Cua-S1 vs TypeSafe Jev", "plan vs execute dry-run", "specialist computer-use vs general agent", "local drop-in vs stub scorer", "route vs memory", "when does it hold / extractable from state", "decision model vs constrained LLM", "dual-process S1/S2", "combinatorial grid vs extractive", "uncalibrated local likelihoods", "decision-native RAG", "classify-first / read selectively", "living applied-mappings atlas / class patterns", "silence as safer / draft-gate heartbeat", "robotics text-state vs pixels", "verbatim ledger vs summary", "judgment as language primitive", "Stagehand extract pick-and-copy", "harness observe-score-act vs demo loop", or "cascade sign-flip / calibration theater": read `references/faq.md`, then `references/mental-models.md`, then `references/mixed-architecture.md`, then @@ -114,13 +114,13 @@ classical method you already trust, substitute it, classify the win | Familiar method | Judgment shape | Detail | |---|---|---| | Mental models across domains (not SWE-only) | EU, abstention, VOI, MCDA, SDT, search/control, Leveson, NATM/Norman/snap-fit; **extractable-from-state boundary map** (self-contained vs needs outside knowledge) | `references/mental-models.md` | -| Judgment-model class (Jev is exemplar, not monopoly) | Species: decide / locate (GLiNER) / categorize (GLiClass) / rank / perceive; open heads include encoder DeBERTa, LoRA distill, openjev-lm, kev. Compaction job is backend-agnostic (Jev Score/Noul vs GLiNER2.5 encoder). Indexer cousin: GLiNER extract + escalate-S2 (10–50× unfilled). Computer-use observe→score-among-candidates→code-acts is backend-agnostic (Jev Ultrafast ↔ GLiNER2 Ultrafast ↔ Cua-S1 specialist; Cua-S1 is not TypeSafe Jev) | `references/judgment-class.md` | +| Judgment-model class (Jev is exemplar, not monopoly) | Species: decide / locate (GLiNER) / categorize (GLiClass) / rank / perceive; open heads include encoder DeBERTa, LoRA distill, openjev-lm, kev. Compaction job is backend-agnostic (Jev Score/Noul vs GLiNER2.5 encoder). Indexer cousin: GLiNER extract + escalate-S2 (10–50× unfilled). Computer-use observe→score-among-candidates→code-acts is backend-agnostic (Jev Ultrafast ↔ GLiNER2 Ultrafast ↔ Cua-S1 specialist ↔ Stagehand experimental Jev stack; Cua-S1 is not TypeSafe Jev; Stagehand pick is a fast path, not a replacement) | `references/judgment-class.md` | | Open weights vs constrained decoding vs proprietary API | Three open paths: encoder open-jev / AR constrained decode (TypeAR + pcdServer; decision-token LoRA; packed one-forward logprob on open LLMs; **CUDA/PyTorch local replica** jevify — uncalibrated likelihoods ≠ Noul) / trained decision-only (Laya + ONNX port, Nimble, kev, **blackwood-rlcd** multimodal now, Archer Watch still Watch). Local `/v1/systemone` surfaces: jev-local (stub until `hf`), kev (trained pointer), von (tiny SAN — not a replica). Softmax over allowed tokens ≠ Noul | `references/judgment-class.md` (when-to-use table) | | Entropy as allocator (low / medium / high) | Typed low+medium decisions → System One marginals; high-entropy synthesis → frontier decoder. Product rhetoric, not a meter. **Hypothesis** | `references/judgment-class.md` | | Formal / semi-formal (proof vs judgment) | Sensor vs constraint vs searchlight; Alloy vs Apalache; DST trio; TOCTOU-of-Noul, AI×FM | `references/formal-methods.md` (one-screen: `references/formal-semi-formal.md`) | -| Mixed architecture (judgment model + LLM) | Provider judges, LLM writes, code owns control; not a stack replacement. Advisory sidecar never changes host routing. Dual-process: S1 decides, S2 generates (routing accuracy unmeasured); S1 specialists + S2 coordinator is the same split (description-only greenfield this hour). Internals ≠ FSM; placement is a **component node**. Browser-use strength = DOM-as-text + speculative fan-out, not vision. Fail polarity is per act: skip-wake fail-open vs merge-gate BLOCK fail-closed. Hybrid local decide + remote fill; `DONE` ≠ verified success. Specialist computer-use: plan ≠ execute, dry-run default (Cua-S1; not TypeSafe Jev). Judgment as a language primitive (Ruby `almost_certain?`/`pick`/`rate`). Decision-native RAG: retrieve wide → decide → evidence set → LLM. Classify-first MCP (topology A): content to the judge without entering main agent context first. Generative UI: model decides, compiler emits. Draft-gate silence ≠ safer (heartbeat). Living class-pattern atlas (not a 342-title dump) | `references/mixed-architecture.md` | +| Mixed architecture (judgment model + LLM) | Provider judges, LLM writes, code owns control; not a stack replacement. Advisory sidecar never changes host routing. Dual-process: S1 decides, S2 generates (routing accuracy unmeasured); S1 specialists + S2 coordinator is the same split (description-only greenfield this hour). Internals ≠ FSM; placement is a **component node**. Browser-use strength = DOM-as-text + speculative fan-out, not vision. Fail polarity is per act: skip-wake fail-open vs merge-gate BLOCK fail-closed. Hybrid local decide + remote fill; `DONE` ≠ verified success. Specialist computer-use: plan ≠ execute, dry-run default (Cua-S1; not TypeSafe Jev). Judgment as a language primitive (Ruby `almost_certain?`/`pick`/`rate`). Decision-native RAG: retrieve wide → decide → evidence set → LLM. Classify-first MCP (topology A): content to the judge without entering main agent context first. Generative UI: model decides, compiler emits. Draft-gate silence ≠ safer (heartbeat). Living class-pattern atlas (not a 342-title dump). Stagehand experimental Jev: pick-and-copy extract + act tree + observe/cache-check; LLM fallback; draft stack | `references/mixed-architecture.md` | | Context sieve | Relevance Noul per block; always-keep set in code; stub + recall key. Encoder cousin: GLiNER2.5 retention Choice + char-offset spans (gliner25-compaction); fail-closed keep_full; shadowMode default. Stdout cousin: jev-pruner (Jev Noul after hard ≤10k/JSON-diff envelope; fail-safe original; archive). Session-ledger cousin: carryforward (verbatim facts; Jev scores recall; rules never judged; fail-open dump). Classify-first MCP cousin: jev-sift (batch path/url/text → Jev without entering main agent context; uncertain/errors/truncation ≠ irrelevant) | `references/applied-mappings.md#1-context-sieve` | -| Exact-text keep / drop | Choice include/exclude/mixed over candidates code already holds. Extractive quotes / pointer-not-generator (model never writes the excerpt; char-offset compaction same species). Observed a11y/DOM controls: score among them; code clicks (Jev or GLiNER2 or Cua-S1 option-attention) | `references/applied-mappings.md#2-exact-text-keep--drop` | +| Exact-text keep / drop | Choice include/exclude/mixed over candidates code already holds. Extractive quotes / pointer-not-generator (model never writes the excerpt; char-offset compaction same species). Observed a11y/DOM controls: score among them; code clicks (Jev or GLiNER2 or Cua-S1 option-attention). Harness productization: Stagehand extract pick-and-copy (Jev picks; code copies; schema/gate else LLM) | `references/applied-mappings.md#2-exact-text-keep--drop` | | Environment / harness triage | Scan every step for env failure; LLM autopsy only on flags. Merge-gate cousin: cluster in code, judge labels cause, policy owns PASS/BLOCK (latch; judge never says ignore alone) | `references/applied-mappings.md#3-environment--harness-triage` | | Moderation and ranking | Hold-before-publish vs graded rerank; fail policy per action | `references/applied-mappings.md#4-moderation-and-ranking` | | Skill / tool routing | Choice over a closed catalog + whether-anything-fits; code dispatches. Route ≠ memory: cheap intent gate skips memory tours on easy routes | `references/applied-mappings.md#5-skill--tool-routing` | @@ -133,7 +133,7 @@ classical method you already trust, substitute it, classify the win | Decision tables / circuits / state machines | Judgment predicates, code owns transitions. Language primitive: Ruby `chance`/`pick`/`rate` as control flow (hunch; English-as-config; fail polarity per action) | `references/mappings.md#3-semantic-predicates--decision-circuits` | | Retrieve + expensive relevance fn | Bounded rerank of a retrieved shortlist. Decision-native RAG: retrieve wide → decide explicitly → evidence set → conflict resolve → reason only over kept evidence (embeddings stay candidate generators; no universal benchmark). Classify-first MCP: same sandwich on agent I/O (path/url/text → judge; main LLM opens survivors) | `references/mappings.md#4-retrieval--bounded-semantic-reranking` (independent TREC DL2019 benchmark: Jev zero-shot best MAP 0.4748, nDCG@10 0.683 vs tuned monoBERT 0.718 — competitive, not dominant) | | Store as semantic index (SQL / SQLite / zoxide / dataframe) | Cheap exact predicates first; typed questions on the remainder. In-engine extension (sqlite-jev) vs CLI rewrite (jevql) vs path index (joxide) vs dataframe columns (jevpandas / jevframe) | `references/mappings.md#4-retrieval--bounded-semantic-reranking` | -| Soft judgment inside a hard envelope | Model may only match the deterministic policy or be more conservative (bitrate ABR; query-planner override-when-confident; compaction mutations/shell operators → keep_full; stdout prune: ≤10k/JSON-diff-whole-doc untouched, then Noul; Cua-S1: plan≠execute, dry-run, fail-closed checkbox/fill; pre-exec toolgate: allow/block/review — Jev is not authorization; guard error/timeout stops) | `references/mappings.md#12-runtime-assurance-sandwich-hypothesis`; `references/mappings.md#18-structural-prove--soft-remainder-hypothesis-as-domain-general-empirical-as-named-shapes` | +| Soft judgment inside a hard envelope | Model may only match the deterministic policy or be more conservative (bitrate ABR; query-planner override-when-confident; compaction mutations/shell operators → keep_full; stdout prune: ≤10k/JSON-diff-whole-doc untouched, then Noul; Cua-S1: plan≠execute, dry-run, fail-closed checkbox/fill; Stagehand extract: schema/completion-gate/screenshot-always-LLM then pick, else LLM; pre-exec toolgate: allow/block/review — Jev is not authorization; guard error/timeout stops) | `references/mappings.md#12-runtime-assurance-sandwich-hypothesis`; `references/mappings.md#18-structural-prove--soft-remainder-hypothesis-as-domain-general-empirical-as-named-shapes` | | Value of information / gather as an act | Pay for another observation only if EV(decision) improves more than cost; abstain from calling *any* model when a regex already answers (meta-VOI). Fail-open wake/resume: skip the LLM turn only if the judge answers and p(wake) is low (Horvitz). Selective memory: verbatim ledger + scored recall (carryforward; rules never judged; 9×3 is a hint). Classify-first read: pay for a full agent open iff relevance (or typed question) says it might change the act (jev-sift; errors/truncation ≠ irrelevant) | `references/mappings.md#6-value-of-information--gather-as-an-enumerated-act` (**Hypothesis** until a labeled act/outcome log; 149-row receipt is Empirical as a shape; wakegate 21/21 is smoke) | | Signal detection / ROC | Criterion and operating point from costs and base rate, not accuracy | `references/mappings.md#7-signal-detection--criterion-not-accuracy` (**Hypothesis** for non-SWE plots) | | Org / safety control structure | Sensor ≠ constraint (Leveson); STPA if the sensor lies | `references/mappings.md#8-control-structure--sensor--constraint-leveson` | @@ -150,8 +150,8 @@ classical method you already trust, substitute it, classify the win | Effect-oriented state-machine loops | Soft predicates on transitions; code owns the transition | `references/mappings.md#19-effect-oriented-state-machine-loops-hypothesis` (**Hypothesis**; ZIO client, not Effect.ts) | | Agent self-supervision / on-track detection | Pre-gate → output judge → done-check → supervisor nouls. S1 reflex keeps control; optional S2 is one-use advice. Claim/evidence Stop (anti-hallucinated-done); bounded Pi supervisor (shadow recovery, never generates commands) | `references/agent-self-assessment.md` | | Optimizer/program frameworks (Ax, DSPy) | Typed fields → one provider request; judge metrics; threshold discipline. Ax and DSPy climb LM-program knobs only | `references/optimizer-integration.md` | -| Perception → decision pipeline / measure / hill-climb | Stages with a versioned state contract; frozen taskset; DSPy/Ax only on the LM-program slice. Specialist composition stays **Hypothesis**; open multimodal decide (blackwood-rlcd) is a named receipt. Structured observe→decide→verified-act (no screenshots) is a computer-use speed-layer receipt (Jev or GLiNER2 or Cua-S1 specialist; Cua-S1 source-only, not TypeSafe Jev). Same section as the row below | `references/validation.md#eval--hill-climb` | -| Eval & hill-climb | Decision-stage jevals hygiene; Harbor taskset × harness × runtime; one score-composition table. Shared bake-off exemplar: open-jev-laya-bench (ECE/NLL/Brier; LLM-as-judge is not the score). Harbor-style frozen protocol vs constrained LLMs: DMB (accuracy/calibration/latency/cost; raw logs). Feedstock: jevals-data CC-BY-4.0 boards + JSONL (recompute-from-logs). Collab-arm curriculum: llm_autonomous vs scripted_plus_jev vs llm_plus_jev (Wilson / McNemar). Negative: combinatorial grid assembly ≠ extractive (ARC-AGI Direct Jev 4/400). Harbor on/off routing: chess-engine tasks, hidden perft verifier, one-run preliminary (jev-gateway-bench). Pair CI merge-gate with Harbor + rh-guard. Harbor needle/noise stdout prune: jev-pruner (manual sweep theirs; plugin eval cannot reach Jev → fail-safe original; Terminal-Bench pilot is integration not a full bench). Cua-S1 specialist form: source-only (metric names, no checkpoint scores; not TypeSafe Jev). Pre-registered independent eval: jev-baselines-eval (**both AMBIGUOUS**; cascade sign-flip at exact parity; confidence=1.0 theater; encoder-with-labels wins; serving-path ≠ model-speed). Healthcare Harbor-shaped: explore-typesafe-ai (synthetic FHIR; not clinically validated). Honest-negative PDF: databricks-jev-pdf-lab (no quality-equivalent Jev payoff) | `references/validation.md#eval--hill-climb` | +| Perception → decision pipeline / measure / hill-climb | Stages with a versioned state contract; frozen taskset; DSPy/Ax only on the LM-program slice. Specialist composition stays **Hypothesis**; open multimodal decide (blackwood-rlcd) is a named receipt. Structured observe→decide→verified-act (no screenshots) is a computer-use speed-layer receipt (Jev or GLiNER2 or Cua-S1 specialist; Cua-S1 source-only, not TypeSafe Jev). Stagehand experimental Jev is the same job inside a major harness (pick-and-copy extract; 37/75 no-LLM ~0.5s vs 4.37s is *their* card; pick ≠ replacement). Same section as the row below | `references/validation.md#eval--hill-climb` | +| Eval & hill-climb | Decision-stage jevals hygiene; Harbor taskset × harness × runtime; one score-composition table. Shared bake-off exemplar: open-jev-laya-bench (ECE/NLL/Brier; LLM-as-judge is not the score). Harbor-style frozen protocol vs constrained LLMs: DMB (accuracy/calibration/latency/cost; raw logs). Feedstock: jevals-data CC-BY-4.0 boards + JSONL (recompute-from-logs). Collab-arm curriculum: llm_autonomous vs scripted_plus_jev vs llm_plus_jev (Wilson / McNemar). Negative: combinatorial grid assembly ≠ extractive (ARC-AGI Direct Jev 4/400). Harbor on/off routing: chess-engine tasks, hidden perft verifier, one-run preliminary (jev-gateway-bench). Pair CI merge-gate with Harbor + rh-guard. Harbor needle/noise stdout prune: jev-pruner (manual sweep theirs; plugin eval cannot reach Jev → fail-safe original; Terminal-Bench pilot is integration not a full bench). Cua-S1 specialist form: source-only (metric names, no checkpoint scores; not TypeSafe Jev). Stagehand extract pick-and-copy: 37/75 no-LLM ~0.5s vs baseline 4.37s (*their* 25×3; pick is a fast path, not a replacement; draft stack #2951–#2955). Pre-registered independent eval: jev-baselines-eval (**both AMBIGUOUS**; cascade sign-flip at exact parity; confidence=1.0 theater; encoder-with-labels wins; serving-path ≠ model-speed). Healthcare Harbor-shaped: explore-typesafe-ai (synthetic FHIR; not clinically validated). Honest-negative PDF: databricks-jev-pdf-lab (no quality-equivalent Jev payoff) | `references/validation.md#eval--hill-climb` | | (meta) Finding new mappings & applications | Toolbox sweep: judgment-shaped component of a known method, substituted + falsified | `references/toolbox-mapping.md` | | Named methods / operators / theorems | Substitution tiers: operand-judgments, preconditioned theorems, non-substitutable | `references/methods-catalog.md` | | (meta) Where a judgment model sits relative to any construct | 11 positions + logical-operator rules + position×construct traversal as the application generator | `references/composition-algebra.md` | diff --git a/.agents/skills/augustus/references/agent-self-assessment.md b/.agents/skills/augustus/references/agent-self-assessment.md index d4d7720..498a597 100644 --- a/.agents/skills/augustus/references/agent-self-assessment.md +++ b/.agents/skills/augustus/references/agent-self-assessment.md @@ -66,6 +66,10 @@ relying: foreman, pi-jev, pi-warden, winnow, fast-jev-compaction, jev-judgment. [Cua-S1](https://github.com/trycua/cua/tree/main/libs/cua-s1) — option-attention among observed elements; plan ≠ execute; dry-run default; source-only (`notes.md` §54). + Harness cousin (draft stack): + [Stagehand #2955](https://github.com/browserbase/stagehand/pull/2955) + — pick-and-copy extract + act tree; LLM fallback; pick ≠ + replacement (`notes.md` §57). Productized Kahneman cascade for *any* cheap-decide / expensive-write loop (business/life, not only SWE): [dual-process-ai](https://github.com/taro1985/dual-process-ai) — diff --git a/.agents/skills/augustus/references/applied-mappings.md b/.agents/skills/augustus/references/applied-mappings.md index 35b4e7b..338f557 100644 --- a/.agents/skills/augustus/references/applied-mappings.md +++ b/.agents/skills/augustus/references/applied-mappings.md @@ -168,6 +168,22 @@ screenshot). Specialist-form cousin, **not TypeSafe Jev:** option-attention among observed elements (fill/check/click/skip); code owns execution order; dry-run default; source-only (`notes.md` §54). +**Harness pick-and-copy (Empirical as PR-body architecture + their +local eval, 2026-09-19 ~00:48; draft stack):** +[Stagehand #2955](https://github.com/browserbase/stagehand/pull/2955) +(5/5 of [#2951](https://github.com/browserbase/stagehand/pull/2951)–#2955, +all OPEN draft) — Jev **picks** observed a11y elements; **code copies** +text. `extract` `"off"` | `"judge"` | `"pick"`. Judge replaces the +metadata LLM `completed` check (throw → LLM). Pick: schema plan +(scalars / bools-enums / lists of flat objects; else LLM); must +validate + completion gate else LLM; screenshot extract always LLM. +Their card (gemini-3.8-flash, 25×3): **37/75** no-LLM in **~0.5 s** +vs baseline **4.37 s** / two LLM calls; 69/75 vs 23/25 (**92% both**); +LLM-off **36/75** — pick is a **fast path, not a replacement**. +Same observe→score-among-candidates→code-acts *job* as jev-ultrafast / +gliner2-ultrafast / solari-reflex / cua-s1, inside a major harness. +Do not merge clocks. Do not copy `experimentalJevAct` +(`notes.md` §57). **Score-among-observed atlas (Empirical as public showcase class pattern, 2026-09-19 ~00:38):** [jevable.com](https://jevable.com/) — candidates already on the diff --git a/.agents/skills/augustus/references/faq.md b/.agents/skills/augustus/references/faq.md index 26c5225..330ce18 100644 --- a/.agents/skills/augustus/references/faq.md +++ b/.agents/skills/augustus/references/faq.md @@ -631,14 +631,27 @@ uses local GLiNER2 (`fastino/gliner2-multi-v1`); uses a Laya head over DOM element indices; [Cua-S1](https://github.com/trycua/cua/tree/main/libs/cua-s1) uses a byte encoder + option-attention head (fill/check/click/skip) — **not -TypeSafe Jev**, source-only this pass. Same lesson as compaction +TypeSafe Jev**, source-only this pass. +[Stagehand #2951–#2955](https://github.com/browserbase/stagehand/pull/2955) +is the same hole **inside a major harness**: Jev picks; code copies +or acts; LLM fallback; extract `"off"`/`"judge"`/`"pick"`. Their +card: 37/75 no-LLM ~0.5 s vs baseline 4.37 s; LLM-off 36/75 — **pick +is a fast path, not a replacement.** Draft stack; do not copy the +opt-in flag. Same lesson as compaction (Jev Noul/Score vs GLiNER2.5). Screenshot multimodal (blackwood-rlcd: letters on an image) is a **different input**, not a better version of this hole. Hybrid local decide + remote fill is mixed-architecture economics, not dual-process-ai. `DONE` is loop termination, not verified success. Plan ≠ execute; dry-run default on Cua-S1. Not GLiNER2.5. Not a bake-off against the Flights demo clock. -`judgment-class.md`; `mixed-architecture.md`; `notes.md` §52, §54. +`judgment-class.md`; `mixed-architecture.md`; `notes.md` §52, §54, §57. + +## Does Stagehand extract replace the LLM? + +No. Pick-and-copy is a **fast path**. Schema leftovers, screenshot +extract, failed gates, and abstention still call the LLM. 36/75 with +the LLM disabled is the honesty number. `applied-mappings.md` §2; +`notes.md` §57. ## Is Cua-S1 TypeSafe Jev? diff --git a/.agents/skills/augustus/references/judgment-class.md b/.agents/skills/augustus/references/judgment-class.md index 4fac6c5..7ca56d5 100644 --- a/.agents/skills/augustus/references/judgment-class.md +++ b/.agents/skills/augustus/references/judgment-class.md @@ -150,7 +150,20 @@ below, next to the when-to-use table. or selectors. Plan ≠ execute; dry-run default; `execute` and `submit` independent opt-ins; fail-closed on unknown checkbox state. Profile `cua-s1-form-v0` is source-only — no weights, no - checkpoint scores (`notes.md` §54). Do not copy `uv` / MCP. + checkpoint scores (`notes.md` §54). Do not copy `uv` / MCP. + **Harness productization (Empirical as PR-body architecture + + their local eval, 2026-09-19 ~00:48; draft Watch):** + [Stagehand #2955](https://github.com/browserbase/stagehand/pull/2955) + (5/5 of #2951–#2955, all OPEN draft) puts the same + observe→score-among-candidates→code-acts hole inside + Browserbase Stagehand. Jev picks a11y elements; code copies text + or acts. Extract `"off"` | `"judge"` | `"pick"`. Schema / completion + gate / screenshot-always-LLM in **code**; LLM fallback. Their + extract card (gemini-3.8-flash, 25×3): 37/75 no-LLM ~0.5 s vs + 4.37 s; 69/75 vs 23/25 (92% both); LLM-off 36/75 — pick ≠ + replacement. Do not merge with jev-ultrafast / gliner2-ultrafast + / Cua-S1 clocks. Not multimodal pixels on the pick path + (`notes.md` §57). Do not copy `experimentalJevAct`. - **Decide.** Typed Choice/Score/Noul with a decision/proper-scoring objective. That is Jev's product claim. Open heads copy the *shape*; distillation copies the *teacher* (openjev-lm, jev-gate-student-b). diff --git a/.agents/skills/augustus/references/mappings.md b/.agents/skills/augustus/references/mappings.md index 1c32f67..aaba19e 100644 --- a/.agents/skills/augustus/references/mappings.md +++ b/.agents/skills/augustus/references/mappings.md @@ -508,6 +508,15 @@ classifier step (`notes.md` §52). same substituted-classifier *job* on a specialist form contract (option-attention; plan ≠ execute; not TypeSafe Jev; source-only, `notes.md` §54). +**Harness productization of the same job (Empirical as PR body, +2026-09-19 ~00:48; draft):** +[Stagehand #2951–#2955](https://github.com/browserbase/stagehand/pull/2955) +— the *algorithm* is Stagehand's act/observe/extract loop; the +substituted classifier step is Jev pick among a11y candidates, then +code copies or acts. LLM fallback when the pick/gate/schema fails. +Extract 37/75 no-LLM ~0.5 s vs 4.37 s is *their* card; pick ≠ +replacement. Cache-check errors never block replay. Do not merge +with demo-loop clocks (`notes.md` §57). **Robotics text-state, same job different body (Empirical as showcase class pattern, 2026-09-19 ~00:38):** @@ -656,6 +665,15 @@ without advertised token `set_value`). The option-attention head may only pick among observed elements and extracted `Label: value` entities. Not TypeSafe Jev. No checkpoint scores (`notes.md` §54). +**Named harness extract envelope (Empirical as PR body, 2026-09-19 +~00:48; draft Watch):** +[Stagehand #2955](https://github.com/browserbase/stagehand/pull/2955) +— the monitor is **code** (schema plan: scalars / bools-enums / +lists of flat objects else LLM; completion gate; screenshot extract +always LLM). Jev may only pick among a11y candidates; code copies +text. Invalid / abstain → LLM. Pick is a fast path, not a +replacement (`notes.md` §57). + ## 13. DST multiverse triage (Hypothesis) **Method**: Antithesis / Resonate DST artifacts → failure taxonomy → diff --git a/.agents/skills/augustus/references/mental-models.md b/.agents/skills/augustus/references/mental-models.md index 29dd488..0c4a2e2 100644 --- a/.agents/skills/augustus/references/mental-models.md +++ b/.agents/skills/augustus/references/mental-models.md @@ -470,13 +470,14 @@ Use these as *existence proofs of a position*. Write your own card. | Agent context | compact completed tool results without inventing prose | retention Choice + char-offset locate (**Empirical**: gliner25-compaction; same *job* as fast-jev-compaction / pi-jev-compaction) | mutation/shell envelope → keep_full; fail-closed keep_full; shadowMode before replace; copy exact bytes | | Agent context | prune Bash stdout before the LLM without inventing prose | Noul per chunk after a hard size/format envelope (**Empirical**: jev-pruner) | ≤10k / JSON-diff-whole-doc untouched; fail-safe original; archive dropped spans | | Dataframe labeling | classify / score rows | Noul/Choice/Score + full `p__` (**Empirical** as jevframe / jevpandas *shape*) | pandas/Polars, thresholds in code | -| Computer-use speed | one verified act per step | score / Choice among numbered a11y/DOM controls (**Empirical**: solari-reflex Jev; gliner2-ultrafast GLiNER2; laya-mind2web Laya DOM indices; cua-s1 option-attention, source-only, not TypeSafe Jev) | Guard check; deny-list absence; no screenshots; `DONE` ≠ success; plan ≠ execute | +| Computer-use speed | one verified act per step | score / Choice among numbered a11y/DOM controls (**Empirical**: solari-reflex Jev; gliner2-ultrafast GLiNER2; laya-mind2web Laya DOM indices; cua-s1 option-attention, source-only, not TypeSafe Jev; Stagehand experimental Jev harness, draft) | Guard check; deny-list absence; no screenshots on pick path; `DONE` ≠ success; plan ≠ execute; LLM fallback; pick ≠ replacement | | Agent turn | skip memory tour on easy intent | intent Choice (**Empirical**: jev-hermes) | Memory still writes; complex still searches | | Document / lab routing | which pages need the expensive observation | Noul on remainder after a text layer / recipe | local extract, merge order (**Empirical** as OCR-router *shape*) | | Shell / tool allowlist | unlisted remainder after a **proof** | five Nouls on unknown verbs | Proven/Refused in code; cannot block (**Empirical**: jevgate) | | SWE | residual AGENTS.md / CLAUDE.md rules | one Score per named instruction-file rule | linter owns hard rules; bands + fail-open (**Empirical**: Abide replay, `notes.md` §47) | | Screenshot candidates → act | lettered elements code already marked | Choice over those letters | Click in code (**Empirical** as blackwood-rlcd *shape*; CC BY-NC) | -| Browser / DOM candidates → act | numbered elements from a **text** snapshot | score among those ids (**Empirical**: atlas browser-use / jev-ultrafast / gliner2-ultrafast *shape*: DOM-as-text, not vision; cua-s1 specialist form, source-only) | Click in code; no screenshots; hybrid remote TYPE optional; plan ≠ execute | +| Browser / DOM candidates → act | numbered elements from a **text** snapshot | score among those ids (**Empirical**: atlas browser-use / jev-ultrafast / gliner2-ultrafast *shape*: DOM-as-text, not vision; cua-s1 specialist form, source-only; Stagehand a11y + editable-id side channel) | Click / copy in code; no screenshots; hybrid remote TYPE optional; plan ≠ execute; schema/gate else LLM | +| Extract from a page | values already in element text | pick elements; copy bytes (**Empirical** as Stagehand #2955: 37/75 no-LLM ~0.5s vs 4.37s *their* card) | Schema plan; completion gate; screenshot → LLM; 36/75 LLM-off honesty | | Knowledge / recall | fact that is not in the document | **Do not ask.** Retrieve the passage first; then a self-contained Choice (**Empirical**: history suite A wrong@0.90 → C right@0.97) | Index, citation, the passage in `state` | | Dual-process cascade | cheap classify / route vs write | S1 typed decision + τ; S2 generates only on low conf (**Empirical as a productized metaphor**; routing accuracy **unmeasured** — dual-process-ai) | Safety still fail-closed in code | | CI merge-gate | ignore infra noise without merging a real bug | cause Choice per cluster (**Empirical**: latch demo PASS vs BLOCK) | Cluster + fingerprint + `--gate` table; reporter never fails the runner | diff --git a/.agents/skills/augustus/references/methods-catalog.md b/.agents/skills/augustus/references/methods-catalog.md index ced5d86..caf48e9 100644 --- a/.agents/skills/augustus/references/methods-catalog.md +++ b/.agents/skills/augustus/references/methods-catalog.md @@ -48,7 +48,7 @@ judgment component is new). | Beam search over taxonomies | Which branches deserve expansion | Choice distributions as branch priority; keep K paths where ambiguity is early | Frontier, budget, final selection | **Empirical recipe** (beam K=3 cookbook) | | Structure induction over a bag | Pairwise "does i depend on j?" (or Choice over order) | One judgment per pair; DAG / scheduler in code | Topology, cycles, execution | **Empirical as a shape** (dag-jev experiment; empty README; no metrics, `notes.md` §48) | | Collab-arm product loop | Scripted legal set vs LLM-propose vs unconstrained | Choice over legal actions; stop on low p rather than guess | Legality, Wilson/McNemar, ceiling flags | **Empirical as a harness shape** (jev-testbench; bake into jevals/Harbor, `notes.md` §48) | -| Computer-use observe → score → act | Which observed control matches the current requirement | Score / Choice / option-attention among a11y/DOM candidates (Jev *or* GLiNER2 *or* Laya *or* Cua-S1); code clicks. Robotics cousin: Choice on geometry-as-text, not pixels | Observation, freshness, dates; never generate selectors or values; independent outcome check (`DONE` ≠ success); plan ≠ execute; kinematics / Hz in code | **Empirical recipe** as shipped loops (jev-ultrafast / solari-reflex Jev; gliner2-ultrafast GLiNER2; laya-mind2web DOM-index Laya); **Empirical as README/MODEL_CARD** for Cua-S1 (source-only, not TypeSafe Jev, `notes.md` §54); contrast blackwood-rlcd screenshot, `notes.md` §48, §52. **Empirical as showcase** MuJoCo / MOSS text-state (`notes.md` §56); drawing-pixel claim ≠ Archer | +| Computer-use observe → score → act | Which observed control matches the current requirement | Score / Choice / option-attention among a11y/DOM candidates (Jev *or* GLiNER2 *or* Laya *or* Cua-S1); code clicks. Harness cousin: Stagehand pick then copy/act. Robotics cousin: Choice on geometry-as-text, not pixels | Observation, freshness, dates; never generate selectors or values; independent outcome check (`DONE` ≠ success); plan ≠ execute; kinematics / Hz in code; schema/gate else LLM | **Empirical recipe** as shipped loops (jev-ultrafast / solari-reflex Jev; gliner2-ultrafast GLiNER2; laya-mind2web DOM-index Laya); **Empirical as README/MODEL_CARD** for Cua-S1 (source-only, not TypeSafe Jev, `notes.md` §54); **Empirical as PR body** for Stagehand #2951–#2955 (draft; 37/75 no-LLM ~0.5s vs 4.37s *theirs*; pick ≠ replacement, `notes.md` §57); contrast blackwood-rlcd screenshot, `notes.md` §48, §52. **Empirical as showcase** MuJoCo / MOSS text-state (`notes.md` §56); drawing-pixel claim ≠ Archer | | Screening / Wald sequential tests | Pass / fail / keep-looking per candidate | One Noul gate per candidate in one batched request; budget in code | Sequential rule, stop boundaries | **Hypothesis** | | STPA / STAMP control structure | Sensor reading vs enforced constraint | Judgment as sensor; constraints in policy/code/interlock; STPA table if the sensor lies | The constraint, the actuator, the probe | **Contract** as ownership; **Hypothesis** as domain product (`mappings.md` §8) | | PufferLib / Ocean env contracts | Does this episode look like a known trainer-bug mode? | Cluster failing episodes; never "the policy is correct" | Seeded serial env, Ocean sanity, observed rewards | **Hypothesis** as placement; **Contract** that Ocean is not a comparative baseline (`formal-methods.md` DST trio) | @@ -72,7 +72,7 @@ judgment component is new). | Method | Judgment-shaped component | Jev substitution | Stays in code | Status | |---|---|---|---|---| | Claim–evidence entailment (NLI) | supports / contradicts / not-established per claim–source pair | One Choice per pair + review flag; judge against the cited source text only. Stop-hook cousin: claims vs **session** evidence | Quote extraction, citation graph, audit log. When the answer *is* a span you already hold, **point** at line ids and copy verbatim — the model never writes the excerpt (jev-reviewer). Keyword retrieve is not semantic (clear-head) | **Empirical recipe** (citation_check cookbook; jev-reviewer sample study, `notes.md` §48; clear-head Stop, `notes.md` §51). Atlas: paraphrase_support and reversed_meaning_high_overlap correctly judged when both texts are in `state` (`notes.md` §49) | -| Extractive selection + offline re-threshold | Keep/drop over sentences, ids, character offsets, observed DOM controls, or stdout chunks code already holds | Per-item Noul/Choice + one broadcast; join in order; `redecide` on the log with no new calls. Compaction: retention Choice + span locate. Computer-use: score among a11y/DOM candidates. Stdout prune: Noul per chunk after hard envelope | Numbering, header skip, thresholds, publish permission; mutation envelope; copy exact bytes; click in code; never generate selectors; archive dropped stdout | **Empirical recipe** (testimonial-miner 8-request fixture; jev-reviewer; gliner25-compaction char-offset copies, `notes.md` §48, §50; gliner2-ultrafast observe→score→act, `notes.md` §52; jev-pruner, `notes.md` §53). Cousin of applied-mappings §2 | +| Extractive selection + offline re-threshold | Keep/drop over sentences, ids, character offsets, observed DOM controls, or stdout chunks code already holds | Per-item Noul/Choice + one broadcast; join in order; `redecide` on the log with no new calls. Compaction: retention Choice + span locate. Computer-use: score among a11y/DOM candidates. Stdout prune: Noul per chunk after hard envelope. Harness extract: pick elements, copy text | Numbering, header skip, thresholds, publish permission; mutation envelope; copy exact bytes; click in code; never generate selectors; archive dropped stdout; schema/gate else LLM | **Empirical recipe** (testimonial-miner 8-request fixture; jev-reviewer; gliner25-compaction char-offset copies, `notes.md` §48, §50; gliner2-ultrafast observe→score→act, `notes.md` §52; jev-pruner, `notes.md` §53). **Empirical as PR body** Stagehand #2955 pick-and-copy, `notes.md` §57. Cousin of applied-mappings §2 | | Combinatorial grid / program synthesis | Consistent whole-object from many cells | **Rejected as extractive.** Cell-wise Choice does not assemble ARC grids (4/400 Direct Jev) | Search, a program, a simulator | **Empirical as a negative** (`notes.md` §49) | | Spec vs artifact conformance (model checking *mindset*) | Property holds / violated / unverifiable for a named requirement | One Noul/Score per requirement, batched; violated → named rule back into context (pi-warden / Abide shape). This is **not** TLC/Apalache/GNATprove | Requirement enumeration, enforcement, logging; the **linter** if the rule is lintable; the real checker if you have one | **Empirical recipe** (pi-warden: 6→0 rule breaks, 150 paired runs; jev-pref: YOU define the rule; Abide: productized compile/calibrate/tune/replay, `notes.md` §47; if-ai: plain-English PR check, fail-closed on error, `notes.md` §51). Ownership split: `formal-methods.md` | | AST ∩ semantic lint | Semantic remainder after a parser already extracted units | Typed questions on Tree-sitter targets; do not execute scanned code | Parser, selection, fail-on; `tenbin` owns the lint *skill* | **Empirical as a shape** (jevscan 0.2.0rc4; not a calibration claim; `notes.md` §48) | diff --git a/.agents/skills/augustus/references/mixed-architecture.md b/.agents/skills/augustus/references/mixed-architecture.md index abd9b2d..dea495c 100644 --- a/.agents/skills/augustus/references/mixed-architecture.md +++ b/.agents/skills/augustus/references/mixed-architecture.md @@ -96,6 +96,13 @@ among a11y/DOM candidates → code acts (`notes.md` §52). Hybrid: local decide; remote fill only for TYPE. `DONE` is not verified success. +**Harness productization this hour (draft Watch):** +[Stagehand #2951–#2955](https://github.com/browserbase/stagehand/pull/2955) +puts the same DOM-as-text loop inside Browserbase Stagehand: Jev +picks; code copies or acts; LLM fallback. Extract pick-and-copy is +a fast path, not a replacement (`notes.md` §57). Do not copy the +opt-in flag. + **Dual-process cascade (Kahneman productized; routing accuracy unmeasured).** [`taro1985/dual-process-ai`](https://github.com/taro1985/dual-process-ai) @@ -193,7 +200,9 @@ not a global virtue: | Plain-English PR check | **Fail closed** on error / empty / low confidence (`if-ai`) | A skipped or timed-out check is not a pass. Threshold is policy, not measured correctness | | Route to a tool / start a side effect | **Fail closed** (don't call) | A wrong tool is an action | | Execute a proposed tool call | **Fail closed** on block / timeout / guard error (`toolgate`) | Execution is the irreversible act. `review` needs authenticated human approval, not self-approval. Jev is not authorization. Distinct from ndolinschi allow/ask_human/deny *vocab* | -| Actuate an observed browser control | **Fail closed** (code validates the node) | Freshness / visibility / disabled / occlusion in code; model never emits selectors (`gliner2-ultrafast`, jev-ultrafast, solari-reflex). `DONE` does not authorize "success". Cua-S1: dry-run default; `execute`/`submit` opt-in; fail-closed unknown checkbox; fill execution fails closed without token `set_value` | +| Actuate an observed browser control | **Fail closed** (code validates the node) | Freshness / visibility / disabled / occlusion in code; model never emits selectors (`gliner2-ultrafast`, jev-ultrafast, solari-reflex). `DONE` does not authorize "success". Cua-S1: dry-run default; `execute`/`submit` opt-in; fail-closed unknown checkbox; fill execution fails closed without token `set_value`. Stagehand: same; LLM fallback when Jev abstains | +| Replay a cached browser action | **Fail open** on the freshness check (`stagehand` cacheCheck) | Errors/timeouts never block replay; a stale verdict re-infers. Opt-in: the check costs a snapshot + a request | +| Skip the LLM on extract / act | **Fail open** to the generator (Stagehand pick/judge) | Schema/gate/screenshot envelope in code; pick is a fast path, not a replacement. Invalid extract → LLM | | Rerank a retrieved list | Fail open: keep retrieval order (`WiktorB2004/llama-index-jev`, **Empirical recipe** on BEIR nfcorpus: MiniLM 0.340 nDCG@5 → MiniLM+Jev 0.396; rerank fails open, *select* fails closed). Listwise/cross-encoder scores belong here, not on the row above. | Ranking errors are quality; selection errors are control-flow | Worked placements (2026-09-18 topic:jev hour + prior archive): @@ -482,7 +491,8 @@ decision-design card. Do not clone APIs from READMEs. | Decision-as-business-tool | Named judgment; gate is part of the result | Registry, arithmetic, hard guards | jev-decision-layer (unofficial) | | NL cases → checked e2e | Jev selects observed controls | Playwright expectations; PASS/FAIL/BLOCKED | jev-e2e (alpha) | | Extractive quotes / pointer evidence | Per-sentence, per-line-id, or char-offset Noul/Choice | Verbatim join; place; `redecide` / CSV; model never writes the excerpt | testimonial-miner; jev-reviewer; gliner25-compaction | -| Structured observe → decide → act | Score / Choice among numbered a11y/DOM controls | Guard check; deny-list absence; no screenshots; no generated selectors; TYPE is the only generation; `DONE` ≠ verified success | solari-reflex (Jev); jev-ultrafast (Jev); gliner2-ultrafast (GLiNER2); laya-mind2web (Laya, DOM indices); cua-s1 (option-attention fill/check/click/skip; not TypeSafe Jev; source-only) | +| Structured observe → decide → act | Score / Choice among numbered a11y/DOM controls | Guard check; deny-list absence; no screenshots; no generated selectors; TYPE is the only generation; `DONE` ≠ verified success | solari-reflex (Jev); jev-ultrafast (Jev); gliner2-ultrafast (GLiNER2); laya-mind2web (Laya, DOM indices); cua-s1 (option-attention fill/check/click/skip; not TypeSafe Jev; source-only); Stagehand experimental Jev (harness; draft #2951–#2955) | +| Harness pick-and-copy extract | Choice among a11y candidates; completion Noul | Schema plan + validation gate in code; screenshot always LLM; LLM fallback; pick ≠ replacement | Stagehand #2955 (`off`/`judge`/`pick`; 37/75 no-LLM ~0.5s vs 4.37s *their* card) | | Specialist form S1 (plan ≠ execute) | Option-attention among observed elements | Dry-run default; snapshot-bound tokens; reobserve; submit opt-in; fail-closed checkbox/fill | cua-s1 (`cua-s1-form-v0` profile; no weights this pass) | | Hybrid local decide + remote fill | Local encoder scores observed controls | Code owns actuators; remote OpenAI-compat helper writes field text only | gliner2-ultrafast (GLiNER2 local + Mercury 2.5 default) | | Dataframe semantic columns | Noul / Choice / Score per row; full `p__` | pandas/Polars, indexes, never silent renormalize | jevpandas; jevframe (PyPI + Polars) | diff --git a/.agents/skills/augustus/references/question-design.md b/.agents/skills/augustus/references/question-design.md index 2d3ea17..d67ca26 100644 --- a/.agents/skills/augustus/references/question-design.md +++ b/.agents/skills/augustus/references/question-design.md @@ -44,7 +44,7 @@ request, and treat a stale pin as a prior, never a setting. ## Criteria shape - Criteria are an extension of the instruction and must ask the same thing, in the same direction (a Noul whose `true` side describes "no" performs worse). -- Choice options: contrastive `what` / `not_for` / short concrete `examples` (instances, not descriptions of instances). Add an `other` / `none-of-the-above` when the list may not cover inputs. Skipping that hatch is not a style nit: the model will pick a listed option at confidence 1.00, and no downstream gate will see a problem (`notes.md` §46). Request-shape lint (wellposed / `tenbin`) puts the hatch on the offered set; **training must confront it as a wrong alternative too**, with varied wording, or the model learns "this wording ⇒ pick it" ([kev](https://github.com/jaredpalmer/kev) first-run shortcut; dedicated `none_of_the_above` eval; `notes.md` §45 delta). Corpus-scale cousin (maker claim, not re-run): SEO internal-link audit **584** placed / **139** refused because nothing honestly fit — Choice-with-`other` at catalog scale (`notes.md` §45–§46, §56). +- Choice options: contrastive `what` / `not_for` / short concrete `examples` (instances, not descriptions of instances). Add an `other` / `none-of-the-above` when the list may not cover inputs. Skipping that hatch is not a style nit: the model will pick a listed option at confidence 1.00, and no downstream gate will see a problem (`notes.md` §46). Request-shape lint (wellposed / `tenbin`) puts the hatch on the offered set; **training must confront it as a wrong alternative too**, with varied wording, or the model learns "this wording ⇒ pick it" ([kev](https://github.com/jaredpalmer/kev) first-run shortcut; dedicated `none_of_the_above` eval; `notes.md` §45 delta). Corpus-scale cousin (maker claim, not re-run): SEO internal-link audit **584** placed / **139** refused because nothing honestly fit — Choice-with-`other` at catalog scale (`notes.md` §45–§46, §56). Computer-use cousin: Stagehand pick asks `best` (no none) **and** `strict` (with none; vetoes above 0.9); ambiguity **stops rather than guesses** (`notes.md` §57). - Score levels (2–10): describe **situations**, one dimension each, each standing alone (Jev sees neither the level's number nor its neighbors — "worse than previous" means nothing). No numerals. Levels may be objects `{"summary", "signals"}`. Give a rare extreme its own level when code treats it differently. - Composite scoring: one Score per dimension, normalize by `len(criteria)-1`, weight and combine in code. Change policy by changing weights — never by rewriting questions. - Taxonomy walk: one Choice per tree level, walk in code; each option's value is its subtree (direct children + sample leaves); follow several branches when probabilities are close. diff --git a/.agents/skills/augustus/references/toolbox-mapping.md b/.agents/skills/augustus/references/toolbox-mapping.md index 9d07313..ac87859 100644 --- a/.agents/skills/augustus/references/toolbox-mapping.md +++ b/.agents/skills/augustus/references/toolbox-mapping.md @@ -79,7 +79,7 @@ component; keep the rest of the method in code. | Discrete math: width vs depth | Fan out in width (parallel ≈ free), pay depth linearly; two-stage only when next options depend on an earlier answer | **Empirical recipe** (fan-out: 12.2× cheaper, 10× faster) | | Psychology: Kahneman | System 2 generates/proposes (LLM), System 1 discriminates (Jev); never the reverse. S1 keeps control; optional S2 is one-use advice and does not fly. Productized cascade: `conf ≥ τ` → S1 decides else S2 writes; routing fails open / safety fails closed; **routing accuracy unmeasured**; keyword fallback ≠ S1. S1 specialists + S2 coordinator is the same split (reification-labs/foreman is description-only Phoenix scaffold this pass — do not invent an Elixir API). Indexer: S1 GLiNER extract on the bulk, escalate LLM on the tail (10–50× unfilled). Healthcare: S1 remainder after NEWS2/code, S2 blinded review (explore-typesafe-ai; synthetic; not clinically validated) | **Empirical recipe** (mcts-agent role split; 2026-09-18 mixed-architecture discourse; jev-reflex-autonomy-lab as a control-loop shape, `notes.md` §46; [dual-process-ai](https://github.com/taro1985/dual-process-ai) as a business/life cascade, `notes.md` §49; s1-graphify-indexer, `notes.md` §51; explore-typesafe-ai, `notes.md` §55). **Route ≠ memory:** a cheap intent gate skips memory/tool *tours* on easy routes; memory still writes; complex still searches (jev-hermes, `notes.md` §48) | | IR / cascades | Cheap relevance / irrelevance before an expensive ranker or generator | **Empirical recipe** (RAG cookbook; jevprune; git-jev-stage; LlamaIndex Jev rerank; jev-pruner stdout after size/format envelope, `notes.md` §53). **Empirical as architecture** (decision-native-rag-skills retrieve-wide→decide→evidence-set; Hypothesis as a measured win, `notes.md` §55). **Empirical as README** (jev-sift classify-first MCP; mocks ≠ accuracy; errors/truncation ≠ irrelevant, `notes.md` §56) | -| IR / extractive keep-drop | Number candidates in code; model selects; copy verbatim; `redecide` thresholds on the log with no new calls. Compaction: same pointer job on tool results (Jev Noul/Score *or* GLiNER encoder). Computer-use: same pointer job on observed a11y/DOM controls (Jev *or* GLiNER2 *or* Cua-S1 option-attention). Stdout: same pointer job on Bash chunks after a hard envelope (Jev Noul). Showcase: score-among-observed (ads, on-screen posts) | **Empirical recipe** (testimonial-miner; jev-reviewer pointer-not-generator, `notes.md` §48; gliner25-compaction char-offset + fail-closed keep_full, `notes.md` §50; gliner2-ultrafast observe→score→act, `notes.md` §52; jev-pruner, `notes.md` §53; cua-s1 specialist form, source-only, `notes.md` §54). Atlas class pattern: Your Signal / Near Here (`notes.md` §56) | +| IR / extractive keep-drop | Number candidates in code; model selects; copy verbatim; `redecide` thresholds on the log with no new calls. Compaction: same pointer job on tool results (Jev Noul/Score *or* GLiNER encoder). Computer-use: same pointer job on observed a11y/DOM controls (Jev *or* GLiNER2 *or* Cua-S1 option-attention *or* Stagehand harness pick). Stdout: same pointer job on Bash chunks after a hard envelope (Jev Noul). Showcase: score-among-observed (ads, on-screen posts) | **Empirical recipe** (testimonial-miner; jev-reviewer pointer-not-generator, `notes.md` §48; gliner25-compaction char-offset + fail-closed keep_full, `notes.md` §50; gliner2-ultrafast observe→score→act, `notes.md` §52; jev-pruner, `notes.md` §53; cua-s1 specialist form, source-only, `notes.md` §54). **Empirical as PR body** Stagehand #2955 pick-and-copy (37/75 no-LLM ~0.5s vs 4.37s *theirs*; pick ≠ replacement, `notes.md` §57). Atlas class pattern: Your Signal / Near Here (`notes.md` §56) | | Spec / lint | Project-defined semantic rules as predicates over a diff; linter owns hard rules. AST remainder: Tree-sitter units, then typed questions; do not execute scanned code. Plain-English PR check: one condition + min-confidence; fail-closed on error | **Empirical recipe** (jev-pref contract; Abide productized path — replay 93 sessions, edit precision ~26% / turn ~73% before tune, `notes.md` §47; JevLint file-level Noul; pi-warden; snifftest unsure-band; jevscan AST∩semantic, `tenbin` owns the lint skill, `notes.md` §48; if-ai, `notes.md` §51). jev-marshal is Watch / empty this pass | | Formal methods / DST / safety | Judgment triages counterexamples, failing seeds, and named-rule conformance; proof/MC/DST stay with their tools. Alloy finder ≠ Apalache BMC ≠ Quint run. DST trio: Antithesis hypervisor / Resonate HQ Lean+oracle+SDK (durable async) / PufferLib env+seed. Noul is a sensor, not a discharged PO. Semi-formal diagrams are vocabularies, not enforcers | **Hypothesis as product**, **Contract** as ownership (matching `mappings.md` §8 and `methods-catalog.md`; worked shape pi-warden — `formal-methods.md`, `formal-semi-formal.md`) | | Decision analysis: VOI | Gather as an enumerated act; pay iff expected decision-loss drop > cost. Wake/resume: skip the LLM turn only if the judge answers and p is low (Horvitz); user-message / skip-limit already answer without a model. Selective memory: score a verbatim ledger; dump on failure; never judge the rules. Classify-first: pay for a full agent open iff relevance might change the act | **Hypothesis** as calculator (`mappings.md` §6; `mental-models.md`). wakegate 21/21 is smoke (`notes.md` §51). carryforward 9×3 is a hint (`notes.md` §55). jev-sift mocks ≠ accuracy (`notes.md` §56) | diff --git a/.agents/skills/augustus/references/validation.md b/.agents/skills/augustus/references/validation.md index d0e609d..5224b64 100644 --- a/.agents/skills/augustus/references/validation.md +++ b/.agents/skills/augustus/references/validation.md @@ -324,7 +324,7 @@ Rules: | LM-program knobs only | DSPy/Ax (narrow) | never primary System One calibration score | | Reward-hack / eval gaming | [rh-guard](https://github.com/24601/rh-guard) | structural deny + System One sidecar | | Project soft-rule lint | [Abide](https://github.com/coldteadotai/abide) | Score per rule on the diff; bands; fail-open; replay + independent review | -| Collab / computer-use product loop | [jev-testbench](https://github.com/ufx7/jev-testbench); [solari-reflex](https://github.com/hitakshiA/solari-reflex); [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast); [cua-s1](https://github.com/trycua/cua/tree/main/libs/cua-s1) | Wilson/McNemar arms; independently checked task time; `DONE` ≠ success; Cua-S1 source-only (metric names, no checkpoint scores) | +| Collab / computer-use product loop | [jev-testbench](https://github.com/ufx7/jev-testbench); [solari-reflex](https://github.com/hitakshiA/solari-reflex); [gliner2-ultrafast](https://github.com/sahibzada-allahyar/gliner2-ultrafast); [cua-s1](https://github.com/trycua/cua/tree/main/libs/cua-s1); [Stagehand #2955](https://github.com/browserbase/stagehand/pull/2955) | Wilson/McNemar arms; independently checked task time; `DONE` ≠ success; Cua-S1 source-only (metric names, no checkpoint scores); Stagehand 37/75 no-LLM ~0.5s vs 4.37s *their* card; pick ≠ replacement; draft | | Agent routing on vs off | [jev-gateway-bench](https://github.com/vinilana/jev-gateway-bench) | Hidden perft; cost/quality; one-run signal this pass | | Command-output prune (needle/noise) | [jev-pruner](https://github.com/tamaratran/jev-pruner) | Manual `trimOutput` sweep (theirs); plugin eval cannot reach Jev (fail-safe original); Terminal-Bench paired pilot is integration, not a full bench | | Pre-registered cascade vs nano/frontier/encoder | [jev-baselines-eval](https://github.com/ickma2311/jev-baselines-eval) | Both experiments **AMBIGUOUS**; cascade sign-flip at exact parity; confidence=1.0 theater; encoder 0.933/9ms with labels; serving-path ≠ model-speed; same-day errata ×3 | @@ -359,7 +359,16 @@ default; source-only this pass. Offline utilities *name* accuracy, abstention, coverage, wrong actions/targets, and unsafe-when-should- abstain; **no checkpoint scores**. Tests exercise implementation, not quality. Do not invent a vs-Jev table. Watch for `cua-s1-form-v0` -(`notes.md` §54). **Collab-arm curriculum:** +(`notes.md` §54). **Harness extract card (their PR body, not +re-run; 2026-09-19 ~00:48):** +[Stagehand #2955](https://github.com/browserbase/stagehand/pull/2955) +— gemini-3.8-flash, Browserbase, local, 25 tasks × 3: 69/75 vs +23/25 baseline (92% both). **37/75** no-LLM in **~0.5 s** vs +baseline **4.37 s** and two LLM calls; LLM-off **36/75**. Pick is a +**fast path, not a replacement.** Draft stack #2951–#2955. In-sample +thresholds on the act suite (#2953). Not Harbor. Do not merge with +solari / Flights / Cua-S1 clocks (`notes.md` §57). +**Collab-arm curriculum:** [jev-testbench](https://github.com/ufx7/jev-testbench) — `llm_autonomous` vs `scripted_plus_jev` vs `llm_plus_jev`; Wilson + McNemar; Jev is not a peer arm. Bake into jevals/Harbor hygiene, do diff --git a/CHANGELOG.md b/CHANGELOG.md index f561f46..a854e57 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -309,6 +309,18 @@ Pegged against [`typesafe-ai/skills` v0.5.7](https://github.com/typesafe-ai/skil silence ≠ safer — not a 342-title hit list). Maker clocks stay claims unless already a named receipt. No wrapper. No invented metrics. +- Stagehand experimental Jev stack (`research/notes.md` §57, + [#2955](https://github.com/browserbase/stagehand/pull/2955) 5/5 of + #2951–#2955, all OPEN draft): architecture notes, not an SDK + how-to. Major harness productization of + observe→score-among-candidates→code-acts (cousins jev-ultrafast / + gliner2-ultrafast / cua-s1 / solari). Jev picks a11y elements; + code copies text. extract `"off"` | `"judge"` | `"pick"`. Their + card (gemini-3.8-flash, 25×3): **37/75** no-LLM ~0.5 s vs baseline + **4.37 s**; 69/75 vs 23/25 (92% both); LLM-off **36/75** — pick is + a fast path, not a replacement. Screenshot extract always LLM. + Cache-check errors never block replay. Do not merge clocks. No + invented metrics. - Effect-oriented loops (`notes.md` §28, `mappings.md` §19): Ward's ZIO client keeps Jev as the outer Choice and the handler as the effect. Not Effect.ts. GLiNER author: GLiNER2 "like jev" is GLiGuard diff --git a/README.md b/README.md index 72399c1..403c280 100644 --- a/README.md +++ b/README.md @@ -63,12 +63,12 @@ never launder a Noul as a proof. (retrieve wide → decide → evidence set); classify-first MCP (jev-sift); draft-gate heartbeat; living class-pattern atlas - `.agents/skills/augustus/references/applied-mappings.md` — context sieve, - exact-text keep/drop (extractive / pointer-not-generator; char-offset compaction; observed a11y/DOM controls; Bash stdout prune; verbatim session ledger / carryforward; classify-first MCP / jev-sift), env triage (OpenSmoke + latch merge-gate), moderation/ranking (decision-native RAG evidence set; living class-pattern atlas), skill routing (route ≠ memory) + exact-text keep/drop (extractive / pointer-not-generator; char-offset compaction; observed a11y/DOM controls; Bash stdout prune; verbatim session ledger / carryforward; classify-first MCP / jev-sift; Stagehand extract pick-and-copy), env triage (OpenSmoke + latch merge-gate), moderation/ranking (decision-native RAG evidence set; living class-pattern atlas), skill routing (route ≠ memory) - `.agents/skills/augustus/references/faq.md` — "just classification", stack replacement, Jev vs open head vs encoder vs LoRA vs constrained AR vs kev vs blackwood, wait-for-Archer, missing-other confident-wrong, soft project rules vs linter (Abide), extractive/pointer-not-generator, compaction summarize vs pointer, encoder vs Jev compaction, fail-closed keep_full, shadow-mode rollout, fail-open vs fail-closed wake vs CI gate, observe→score→act backend-agnostic, hybrid local decide + remote fill, DONE ≠ verified success, stdout prune vs session compaction, Cua-S1 vs TypeSafe Jev, plan ≠ execute / dry-run, local drop-in vs stub scorer, route ≠ memory, when-it-holds / extractable-from-state, decision-model vs constrained LLM, dual-process S1/S2, combinatorial grid ≠ extractive, GLiNER vs GLiClass vs CLIP, LLM-as-judge, in-engine vs CLI store, hard envelope (bitrate / planner), not-another-how-to, - uncalibrated local likelihoods ≠ Noul, decision-native RAG, classify-first MCP, living applied-mappings atlas / class patterns, draft-gate silence ≠ safer, robotics text-state vs pixels, cascade + uncalibrated local likelihoods ≠ Noul, decision-native RAG, classify-first MCP, living applied-mappings atlas / class patterns, draft-gate silence ≠ safer, robotics text-state vs pixels, Stagehand extract pick-and-copy / fast-path not replacement, cascade sign-flip / calibration theater, Precision PDF honest negative - `.agents/skills/augustus/references/mappings.md` — classical-method mappings with boundaries, counterexamples, acceptance tests (including @@ -82,7 +82,9 @@ never launder a Noul as a proof. Harbor-adjacent soft-rule measurement; solari-reflex Harbor-style computer-use; gliner2-ultrafast encoder-backend cousin (`DONE` ≠ success; demo is not a bake-off); Cua-S1 specialist form source-only - (metric names, no checkpoint scores; not TypeSafe Jev); jev-testbench collab arms; ARC-AGI Direct Jev as + (metric names, no checkpoint scores; not TypeSafe Jev); Stagehand + extract pick-and-copy 37/75 no-LLM ~0.5s vs 4.37s (*their* card; + pick ≠ replacement; draft #2951–#2955); jev-testbench collab arms; ARC-AGI Direct Jev as combinatorial-≠-extractive negative; jev-gateway-bench Harbor on/off routing one-run signal; jev-pruner Harbor needle/noise + Terminal-Bench integration pilot, not a full bench; jev-baselines-eval pre-registered diff --git a/docs/ecosystem.md b/docs/ecosystem.md index 29f2362..b244580 100644 --- a/docs/ecosystem.md +++ b/docs/ecosystem.md @@ -131,7 +131,7 @@ the READMEs, not a monopoly. - **AppitStudio/testimonial-miner** — extractive selection + multi-question broadcast + offline `redecide`. Model never writes the quote. - **choxos/jev-reviewer** — pointer-not-generator: line ids; verbatim copy with place; *not found* is an answer. - **us/jev-local** — contract-compatible `POST /v1/systemone`. Default scorer is a **stub** until `JEVLOCAL_SCORER=hf`. -- **hitakshiA/solari-reflex** — observe → decide → verified act; no screenshots. Author table vs Codex on Solari ~3–7× wall. Encoder-backend cousin: gliner2-ultrafast (`notes.md` §52). Specialist-form cousin: cua-s1 (`notes.md` §54). +- **hitakshiA/solari-reflex** — observe → decide → verified act; no screenshots. Author table vs Codex on Solari ~3–7× wall. Encoder-backend cousin: gliner2-ultrafast (`notes.md` §52). Specialist-form cousin: cua-s1 (`notes.md` §54). Harness cousin: Stagehand experimental Jev stack (`notes.md` §57). - **ktaletsk/jevframe** — pandas/Polars `.jev` accessor; full `p__`; sibling of jevpandas. - **de-niji/jev-hermes** — route ≠ memory: cheap intent gate skips memory tours. - **ngallodev-software/agent-workflow-typesafe-ai** — advisory sidecar receipts; never changes host routing (Apache-2.0). @@ -215,6 +215,12 @@ Architecture notes, not a plugin / showcase catalog. `notes.md` §56. TypeSafe J - **kbhuw/jev-sift** — classify first, read selectively. Batch path / public URL / inline text → Jev relevance or 1–8 typed questions. Content to Jev without entering main agent context first. Envelope (theirs): 50 items, 60k char, 2 MB / 20 s, public-IP only, no JS/cookies/login, PDFs unsupported. Uncertain/errors/truncation ≠ irrelevant. Transport tests (mocks) ≠ accuracy. No LICENSE this pass. Same retrieve-wide → decide → evidence-set family as decision-native-rag-skills. Topology A MCP; **not** nekowasabi/jev-routing (host adapter). - **jevable.com** — living applied-mappings atlas. Claimed **342** curated projects; JSON-LD first page **36**. Categories: Agents, Browser extensions, Creative tools, Data & research, Developer tools, Experiments, Finance, Games, Marketing, Productivity, Robotics. No public API this pass. Class patterns: intent columns, score-among-observed, VOI gates, generative UI decide, robotics text-state, draft-gate silence ≠ safer. Maker clocks stay claims unless already a named receipt. +### Hourly ~18:48 Boise 2026-09-18 / 00:48 UTC 2026-09-19 (Stagehand experimental Jev pick-and-copy) + +Architecture notes, not an SDK catalog. `notes.md` §57. TypeSafe Jev is the exemplar, not the monopoly. Archer still Watch. Draft stack. No invented metrics. + +- **browserbase/stagehand #2951–#2955** (MIT parent; all OPEN draft; author miguelg719). 5/5 user link: [#2955](https://github.com/browserbase/stagehand/pull/2955) extract completion **judge** + **pick-and-copy**. Jev picks a11y elements; code copies text. `extract` `"off"` | `"judge"` | `"pick"`. Both modes send page/extracted content to TypeSafe. Schema/gate/screenshot-always-LLM in code; LLM fallback. Their card (gemini-3.8-flash, Browserbase, local, 25×3): 69/75 vs 23/25 (92% both); **37/75** no-LLM ~0.5 s vs baseline **4.37 s** / two LLM calls; LLM-off **36/75** — pick is a fast path, not a replacement. Stack: #2951 editable ids (outline byte-for-byte unchanged); #2952 client + pick library (`best`+`strict`); #2953 act tree; #2954 observe + cache-check (errors never block replay). Same observe→score-among-candidates→code-acts *job* as jev-ultrafast / gliner2-ultrafast / cua-s1 / solari, inside a major harness. Do not merge clocks. Do not copy `experimentalJevAct`. + See `references/mixed-architecture.md` in the skill. Class-level family choice: `references/judgment-class.md`. Proof vs judgment (Alloy vs Apalache; DST trio Antithesis / Resonate HQ / PufferLib): diff --git a/research/archive/findings.md b/research/archive/findings.md index d67dea5..d2bffab 100644 --- a/research/archive/findings.md +++ b/research/archive/findings.md @@ -1180,6 +1180,35 @@ living atlas extracts class patterns, not a 342-row dump; (de) draft-gate silence ≠ safer (heartbeat); (df) robotics text-state ≠ pixels; (dg) 342 is their count / JSON-LD 36 is page 1. +## Batch #41 (2026-09-19 ~00:48 UTC / ~18:48 Boise 2026-09-18) — Stagehand experimental Jev pick-and-copy + +Note: `research/notes.md` §57. Docs-only. Folded into PR #2. Not a +second PR. Archer still Watch. No invented metrics. Do not re-fold +§50–§56. Draft stack — Watch merge. + +- **browserbase/stagehand #2951–#2955 (Empirical as PR-body + architecture + their local eval).** Parent MIT. All OPEN draft. + Author miguelg719. Created 2026-09-17T06:40Z. User link **#2955 + (5/5)**: extract completion **judge** + **pick-and-copy**. Jev + picks a11y elements; code copies text. `extract` `"off"` | + `"judge"` | `"pick"`. Both modes send page/extracted content to + TypeSafe. Schema leftovers / screenshot extract / failed gate → + LLM. Their card (gemini-3.8-flash, Browserbase, local, 25×3): + 69/75 vs 23/25 (92% both); **37/75** no-LLM ~0.5 s vs baseline + **4.37 s** / two LLM calls; LLM-off **36/75** — pick is a fast + path, not a replacement. Stack: #2951 editable ids (outline + unchanged); #2952 client + pick (`best`+`strict`; ambiguity + stops); #2953 act tree (LLM fallback); #2954 observe + cache-check + (errors never block replay). Same observe→score-among-candidates→ + code-acts *job* as jev-ultrafast / gliner2-ultrafast / cua-s1 / + solari, inside a major harness. Do not merge clocks. + +Cross-repo addition: (dh) harness pick-and-copy is pointer-not- +generator at product scale; (di) pick is a fast path, not a +replacement (36/75 LLM-off honesty); (dj) cache-check errors fail +open (never block replay); (dk) best+strict is NOTA at pick time; +(dl) draft stack #2951–#2955 Watch merge. + diff --git a/research/notes.md b/research/notes.md index c297f66..ceaed41 100644 --- a/research/notes.md +++ b/research/notes.md @@ -4849,3 +4849,199 @@ UI decide; draft-gate heartbeat); `faq.md`; `mental-models.md`; `question-design.md` (SEO 139-refused as `other`); `agent-self-assessment.md` (silence ≠ safer); `methods-catalog.md`; `toolbox-mapping.md`. No wrapper. + +## 57. Stagehand experimental Jev stack — harness pick-and-copy (2026-09-19 ~00:48 UTC / ~18:48 Boise 2026-09-18) + +America/Boise ~18:48 = 00:48 UTC 2026-09-19. Docs-only fold into +open PR #2 (`cursor/augustus-store-envelope-00b4`). Not a competing +PR. Archer 27B drop still **WATCH**. Identity lock vs `typesafe-ai` +/ `tenbin` / `decision-first` holds. No wrapper, no SDK/init how-to, +no copied `experimentalJevAct` as a class constant. No invented +metrics — clocks below are **their PR bodies**, not re-run. Do not +re-fold §50–§56 as this product, jev-ultrafast / gliner2-ultrafast / +cua-s1 / solari as a new species, or blackwood-rlcd as the extract +path. + +**Placement.** Major harness **productization** of +observe→score-among-candidates→code-acts (and pointer-not-generator +extract) inside [browserbase/stagehand](https://github.com/browserbase/stagehand) +(MIT; Browserbase Inc.). Same *job* as jev-ultrafast / solari-reflex +(Jev backends), gliner2-ultrafast (GLiNER2), cua-s1 (option-attention, +not TypeSafe Jev), laya-mind2web (Laya DOM indices). This is the +harness, not a demo loop. TypeSafe Jev is the exemplar, not the +monopoly. **Experimental; all five PRs OPEN draft this pass.** +Watch merge of extract + act paths. + +### Stack (all OPEN draft; each PR targets its predecessor) + +Author [@miguelg719](https://github.com/miguelg719). Created +2026-09-17T06:40Z. Opt-in via `experimentalJevAct` / +`STAGEHAND_EXPERIMENTAL_JEV_ACT` (**not** public create config — +cross-language create contract unchanged). Do not copy the flag. + +1. **[#2951](https://github.com/browserbase/stagehand/pull/2951)** + (1/5) — report **editable** element ids alongside the a11y + snapshot (`plaintext` / `richtext` AX `editable`). Side channel + on the private snapshot type: outline text, xpath map, and url + map are **byte-for-byte unchanged**, so nothing the LLM sees (or + any cache key) moves. No consumer in this PR. +2. **[#2952](https://github.com/browserbase/stagehand/pull/2952)** + (2/5) — TypeSafe Jev client + candidate-picking library. No + wiring yet (tree-shaken). `/v1/systemone`; 8 s timeout; per + endpoint+key circuit breaker (auth 60 s, 3 consecutive failures + 30 s); errors carry the code only, never the body or key. Outline + → per-intent views (pointer / input / select / scroll / option / + broad / text / link). Pick: role view first, then named elements; + lists over 40 cut to the 30 sharing words with the instruction; + two questions per request — `best` (no "none") and `strict` (with + "none", vetoes above 0.9); unease holds while the next tier + tries; **ambiguity stops rather than guesses**; twins share a + vote; huge lists sharded by character budget. Args parsed in + code; `%variable%` values redacted before anything leaves the + process. +3. **[#2953](https://github.com/browserbase/stagehand/pull/2953)** + (3/5) — experimental Jev **decision tree for `act()`**. Intent + (no snapshot) → args in code → candidates + pick → existing + `performUnderstudyMethod` → deterministic checks (fill read-back, + native `