Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
43 commits
Select commit Hold shift + click to select a range
ed9f037
Fold 12:58 Boise watch: store-index fork, hard envelopes, harness pra…
cursoragent Sep 18, 2026
efe4694
Fold kev as the runnable Archer reconstruction on the trained path
cursoragent Sep 18, 2026
7ce5043
Fold 14:03 Boise watch: open multimodal RLCD, bake-off, decision-toke…
cursoragent Sep 18, 2026
8d61cd6
Fold Abide as productized soft-rule preference lint
cursoragent Sep 18, 2026
d19541f
Fold kev delta: Hub weights and NOTA training confrontation
cursoragent Sep 18, 2026
b8afd59
Fold 14:52 Boise watch: extractive selection, local drop-in, observe-act
cursoragent Sep 18, 2026
6a6d020
Fold 15:52 Boise watch: boundary map, Harbor bake-off, dual-process
cursoragent Sep 18, 2026
900ba5d
Fold GLiNER2.5 extractive compaction as encoder keep/drop
cursoragent Sep 18, 2026
d220485
Fold 16:48 Boise watch: CI merge-gate, fail-open wake, Harbor on/off
cursoragent Sep 18, 2026
5225ab8
Fold GLiNER2 Ultrafast as encoder observe-score-act backend
cursoragent Sep 18, 2026
d1c3006
Fold jev-pruner as evidence-preserving Bash stdout prune
cursoragent Sep 18, 2026
d2d9ea3
Fold Cua-S1 as specialist form System One computer-use
cursoragent Sep 18, 2026
4aeca62
Fold 17:48 Boise watch: CUDA replica, decision-native RAG, AMBIGUOUS …
cursoragent Sep 19, 2026
59aee41
Fold classify-first MCP and living applied-mappings atlas
cursoragent Sep 19, 2026
1cd1f10
Fold Stagehand experimental Jev pick-and-copy harness
cursoragent Sep 19, 2026
b531f05
Fold 18:46 Boise watch: public wall, meaning-search, attention≠correc…
cursoragent Sep 19, 2026
465b100
Fold 19:48 Boise watch: capability kernel, DSPy control plane, calibr…
cursoragent Sep 19, 2026
6c30287
Fold 20:43 Boise watch: domain specialist vs few-shot, cascade leftov…
cursoragent Sep 19, 2026
b7760f7
Fold 21:39 Boise watch: triage anti-distill, evidence packets, meanin…
cursoragent Sep 19, 2026
34eb180
Fold 22:38 Boise watch: permission vs probability, judgment ≠ permiss…
cursoragent Sep 19, 2026
7673cf9
Fold 23:40 Boise watch: constrained optimizer + S1 features, privileg…
cursoragent Sep 19, 2026
4f6e374
Fold 00:39 Boise watch: evidence-gated packs, Jev supplies evidence, …
cursoragent Sep 19, 2026
223967d
Fold 00:39 Boise remainder: hot-click CU, verbatim compact+gate, anti…
cursoragent Sep 19, 2026
ee06e18
Fold 01:47 Boise watch: control-plane combinators, receipts-not-leade…
cursoragent Sep 19, 2026
f9c1b80
Fold 02:38 Boise watch: TLA+ never-confidently-wrong, SEAL coverage l…
cursoragent Sep 19, 2026
f58b8b3
Fold 03:38 Boise watch: Stop-hook attention redirect, pre-send views,…
cursoragent Sep 19, 2026
0265a9d
Fold 04:39 Boise watch: digital-design combinators, VOI cache, skill-…
cursoragent Sep 19, 2026
b0f7a95
Fold 05:46 Boise watch: jevassert record/replay, BBQ, decider≠executo…
cursoragent Sep 19, 2026
18dfcce
Fold 06:43 Boise watch: SGR-judge contract, jev-use/pi control planes…
cursoragent Sep 19, 2026
063a917
Fold 07:49 Boise watch: 1-token logprob endpoint, jevinf replica, Eli…
cursoragent Sep 19, 2026
5f44bf4
Fold classifier-dev: productized System One HTTP, escalate-under-thre…
cursoragent Sep 19, 2026
03fddc6
Fold choxos/jev-reviewer: systematic-review pointer, two-pass Choice+…
cursoragent Sep 19, 2026
daa70b7
Fold githubnext/localjev: wire-compat ≠ logit-equiv, prompted JSON vs…
cursoragent Sep 19, 2026
bcf66f1
Fold NandhaKishorM/laya: packaging ≠ new species, Router script-befor…
cursoragent Sep 19, 2026
db654b5
Fold @airesearch12 census: external openjev list ≠ scored bake-off
cursoragent Sep 19, 2026
40a5f12
Fold JevBench v1.2 board: geometric-mean I/C/S/K, Harbor honesty
cursoragent Sep 19, 2026
74c0c18
Fold hourly 0842: already-folded apply-the-five recipe, skip thin, so…
cursoragent Sep 19, 2026
4b8298c
Fold khordoo/jev-reflex-autonomy-lab delta: S1 never stalls, consumpt…
cursoragent Sep 19, 2026
a5c9d33
Fold awlevin/typesafe-computer-use: OCR+AX observe→score→act, exclusi…
cursoragent Sep 19, 2026
38ae518
Fold moritzkremb/jev-voice-browser: ASR observe→score→act, partial-sp…
cursoragent Sep 19, 2026
a18d6b4
Fold AgentGhost wrap-as-execution + studio_yebisu JP genre atlas
cursoragent Sep 19, 2026
e6e9e42
Fold Akshay Pachaar “Jev Clearly Explained” as pedagogy
cursoragent Sep 19, 2026
f1e2847
Fold uehaj/jev-semgrep: meaning-grep AND/OR/NOT, proposition≠embeddin…
cursoragent Sep 19, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
77 changes: 44 additions & 33 deletions .agents/skills/augustus/SKILL.md

Large diffs are not rendered by default.

339 changes: 328 additions & 11 deletions .agents/skills/augustus/references/agent-self-assessment.md

Large diffs are not rendered by default.

1,085 changes: 1,077 additions & 8 deletions .agents/skills/augustus/references/applied-mappings.md

Large diffs are not rendered by default.

168 changes: 161 additions & 7 deletions .agents/skills/augustus/references/composition-algebra.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,8 +16,8 @@ mappings.md conventions.
|---|---|---|---|---|---|
| 1 | **Operand** | F(Jev(...)) — judgment's number feeds the function | Nouls as CatBoost features; Score as PUCT leaf value | It's a calibrated belief in your rubric's units, not a natural quantity; version feature/question defs with the consumer | **Empirical recipe** |
| 2 | **Post-judge** | F(x) → Jev judges the result | Output judge (leaks_secret, failure_class); citation check on generated text | Only the post-judge sees what the call printed; the pre-gate cannot | **Empirical recipe** |
| 3 | **Gate** | if Jev(x): apply F — Jev decides *whether* F runs, or whether F's result is admitted | Pre-action gates (destructive .90/exfil .70); winnow context sieve | A gate is a filter, not authorization — validate operation+target in code. Error paths fail **per action**, not always open: an advisory guard fails open *because* a hard interlock or sandbox sits underneath; a gate that selects or authorizes a side effect fails closed (`mixed-architecture.md` prefilter table; `mappings.md` §18) | **Empirical recipe** |
| 4 | **Selector (of F or its parameters)** | Jev picks which F runs: Choice over functions/models/effort levels | jev-router (cheapest model), jev-codex-router (model+effort), DiffJury review_depth | Dispatch stays in code; per-option consequences are your cost model; confidence-gate the selection | **Empirical recipe** |
| 3 | **Gate** | if Jev(x): apply F — Jev decides *whether* F runs, or whether F's result is admitted | Pre-action gates (destructive .90/exfil .70); winnow context sieve; pi-heed side-effect check | A gate is a filter, not authorization — validate operation+target in code. Error paths fail **per action**, not always open: an advisory guard fails open *because* a hard interlock or sandbox sits underneath; a gate that selects or authorizes a side effect fails closed (`mixed-architecture.md` prefilter table; `mappings.md` §18) | **Empirical recipe** |
| 4 | **Selector (of F or its parameters)** | Jev picks which F runs: Choice over functions/models/effort levels | jev-router (cheapest model), jev-codex-router (model+effort), DiffJury review_depth, jeffrey next-tool | Dispatch stays in code; per-option consequences are your cost model; confidence-gate the selection. **Pick ≠ fill:** the LLM may write args; Jev does not | **Empirical recipe** |
| 5 | **Comparator** | Replace a semantic comparator inside sort/rank: "more relevant / more severe" as a key | Rerank; skillranker; order statistics over semantic keys | Comparability needs a shared rubric; measure recall separately from rerank quality | **Empirical recipe** |
| 6 | **Prior / initializer** | Jev distribution seeds a deterministic method that refines it | MCTS PUCT priors; beam-search branch priority | It's a heuristic prior, not a posterior; refine with real observations | **Empirical recipe** |
| 7 | **State estimator, F = controller** | Jev estimates named probabilities; deterministic policy with hysteresis acts | foreman (progress/stuck/complete → continue/stop/retry/verify) | The model never commands; interventions enumerated in code | **Empirical recipe** |
Expand All @@ -40,6 +40,34 @@ mappings.md conventions.
never the model.
- **→ (implication) / chains**: decompose into gate → act → post-judge;
never encode multi-hop logic in one question (indirection costs accuracy).
- **Named combinators** ([jev-combinators](https://github.com/voidning/jev-combinators);
renamed from decision-combinators, same repo):
Then / Gate / Vote / Cascade / Weighted plus **extended**
Router / Loop / Retry / Fallback / Memory are
**control-plane** wiring, not a license to treat parallel
Nouls as independent. Digital-design slogan (transistors /
logic gates / chip) is a *metaphor* for soft classifiers;
∧/∨ aggregation still follows the rule above. Fallback is
the fail-closed node; Memory gates what to remember. No
measurements. `notes.md` §66, §69.
- **Decider ≠ executor** ([jeffrey](https://github.com/thomasbrueggemann/jeffrey)):
position 4 (Selector of next F) stays Jev; arg fill is
generation, not a Jev position. The loop is Jev→tool→Jev.
Pick ≠ fill. Mapping §9 still rejects the fused
planner-writer. `notes.md` §70.
- **Selector of the next word** ([jev-gpt](https://github.com/florian-hoenicke/jev-gpt)):
position 4 applied to generation itself. Each token is a
Choice over a closed lexicon; the model never
free-generates. Architecture demo (~400 calls / 75 s /
2¢ *theirs*). Distinct from jeffrey (selector of next
*tool*). `notes.md` §71.
- **Hand no-text steps** ([jev-use](https://github.com/shitianfang/jev-use)):
selector + gate; writing stays generation. Vercel drops
confidence so margin is a different statistic.
`notes.md` §71.
- **TLA+ kernel around votes** ([jev-labs](https://github.com/copyleftdev/jev-labs)):
aggregation (quorum, stability, escalate) is the spec,
not a multiplied joint of five Nouls. `notes.md` §67.

## Rules that hold across every position

Expand Down Expand Up @@ -143,10 +171,22 @@ Reusable shapes when generating applications:
meeting action items ~150 ms after each utterance.
7. **Formula embedding**: JUDGE/SCORE/CHOOSE as first-class spreadsheet formulas.
8. **Pixel-free computer use**: accessibility tree → compact actionable-JSON → one
batched question set per step → execute via AX actions.
batched question set per step → execute via AX actions. Encoder-backend
cousin: gliner2-ultrafast scores observed a11y/DOM controls with
local GLiNER2; code clicks; `DONE` ≠ success (`notes.md` §52).
Specialist-form cousin: Cua-S1 option-attention (fill/check/click/skip);
not TypeSafe Jev; plan ≠ execute; source-only (`notes.md` §54).
9. **Shadow-mode harness** (jev-harness): policy + confidence gate + shadow mode +
offline eval CLI replaying fixtures, asserting on actions; 24-row filter 48.9 s
(Claude CLI) vs 1.3 s Jev at concurrency 8.
(Claude CLI) vs 1.3 s Jev at concurrency 8. Compaction rollout:
gliner25-compaction public default `shadowMode: true` (log proposed
reduction; do not replace history) (`notes.md` §50).
Stdout-prune cousin: [jev-pruner](https://github.com/tamaratran/jev-pruner)
archives full stdout before scoring; fail-safe keep original
(`notes.md` §53). Marketplace id still `fast-jev-output`.
Recovery cousin: [jevons](https://github.com/LilDojd/jevons) default
recovery **shadow** (record, do not interrupt); steering never
generates commands (`notes.md` §51).

Calibration warning (calibre): routing thresholds and ROI do **not** transfer across
datasets — every gate is a per-dataset measurement (see validation.md).
Expand All @@ -155,10 +195,124 @@ datasets — every gate is a per-dataset measurement (see validation.md).
Jev as gate/selector/verifier *around* a generator, never instead of one.
Cost-sensitive prefilter (drop chunks/lines/hunks before the LLM);
tool/skill routing (Choice + fits-Noul, code dispatches); preference lint
(project-defined rules as criteria). Fail-open vs fail-closed is per
(project-defined *soft* rules as criteria; linter owns hard rules;
Abide is the productized path, `notes.md` §47). Fail-open vs fail-closed is per
action — LlamaIndex Jev rerank fails open (keep retrieval order), select
fails closed. Full card: `references/mixed-architecture.md`.
11. **Structural prove ∩ remainder judge** (jevgate, doc-router): code
(allowlist, text layer) decides the easy cases; typed questions only
11. **Structural prove ∩ remainder judge** (jevgate, doc-router, Abide): code
(allowlist, text layer, linter) decides the easy cases; typed questions only
on leftovers; fail-open unless a real sandbox sits under. Full card:
`mappings.md` §18.
12. **1-token selector / tree of Choices** (chakuho, jev-gpt): the
generator is reduced to a next-label or next-word Choice. Softmax
over declared labels is not a Noul. Numeric rules and writing stay
exact. Full cards: `judgment-class.md`, `mixed-architecture.md`.
13. **Productized System One HTTP** (classifier-dev): the public
contract is label + calibrated confidence, not a paragraph. Batch
state, escalate-under-threshold, and a `FALLBACK` marker are
*code*. Full cards: `mixed-architecture.md`, `validation.md`.
14. **Evidence-synthesis pointer** (choxos/jev-reviewer, ≠ egma-ai):
Jev picks line ids; code copies verbatim; a second absolute Noul
checks "does this line itself answer?"; *Not found* is an answer;
the human tick is the product. Full cards: `applied-mappings.md`
§2, `mixed-architecture.md`.
15. **Prompted-JSON wire** (githubnext/localjev, ≠ kunchenguid/local-jev):
the TypeSafe SDK talks to a local `/v1/systemone`; the model
*writes* probabilities rather than exposing logits. Entropy
confidence is computed in code from that vector. Full cards:
`judgment-class.md`, `mixed-architecture.md`, `validation.md`.
16. **Script-before-p router** (NandhaKishorM/laya packaging of Hub
Laya): pick the checkpoint from script/lang/task *before* the
forward pass, because confidence will not drop on OOD (Khmer
0.000@0.952). Post-T ECE is not raw ECE; 0.85 is still soft.
Full cards: `judgment-class.md`, `mixed-architecture.md`,
`faq.md`, `validation.md`.
17. **External class census** (@airesearch12 / Benchmark
Heaven): a named list is not a rank; GLiNER2 and
routers on the list are a class-boundary, not
identity; incompleteness is lag. Watch the board
URL; do not paste live scores into the census card.
Full cards: `mixed-architecture.md`, `faq.md`,
`validation.md`, `toolbox-mapping.md`.
18. **Geometric-mean product** (JevBench v1.2): four
axes at 25% each; a weak axis cannot be bought
back; weighting is a product design, not a law;
calibration on the rank is a choice (v1.1 kept it
off); instruction models in the same table as NAR
rebuilds; ×2 latency and est. costs are assumptions
to name. Full cards: `mixed-architecture.md`,
`validation.md`, `faq.md`, `mental-models.md`.
19. **Already-folded class as a recipe** (hourly 0842):
when the named HIGHs are already on the branch,
extract how-to-apply instead of re-carding —
wire-compat ≠ logit-equiv; productize label+p and
mark `FALLBACK`; packaging ≠ new species / script-
before-p; pointer-not-generator (two-pass; *Not
found*; human tick); external census ≠ scored
bake-off / geo-mean weights are a design. Skip
thin noise. Hard-gating a Noul as a PR/quality
gate is soundness theater. Full cards:
`mixed-architecture.md`, `faq.md`,
`mental-models.md`, `validation.md`.
20. **S1 keeps flying / S2 one-use** (khordoo/jev-reflex-autonomy-lab
delta of §46): position 4 (Selector of next
*action*) stays on the reflex every tick;
position 7 (state estimator) is the optional
planner — advice, not a command. Escalate-
under-threshold **without stalling**. Log
consumption, not arrival. Local rule-based vs
Live API is an A/B of backends, not a scored
bake-off; Local controller **≠** githubnext/localjev.
Seed = geometry ≠ async replay. No pixels.
Confidence ≠ selected probability. 20% still
soft. S2 never grants. Full cards:
`mixed-architecture.md`, `faq.md`,
`mental-models.md`, `agent-self-assessment.md`,
`validation.md`.
21. **OCR+AX observe→score→act** (awlevin/typesafe-computer-use):
position 10 (Discretizer: screen → numbered items)
then position 4 (Selector of next action). Writer
is generation, not a Jev position. Split kind/item/site
is width-is-cheap. Overlap is concentration theater.
Perception in code rebuilds pixel-free reasoning.
Decision never ships screenshots; the answer reader
may. 155× is one screenshot *theirs*. Full cards:
`mixed-architecture.md`, `faq.md`,
`applied-mappings.md` §9, `validation.md`.
22. **ASR observe→score→act** (moritzkremb/jev-voice-browser):
position 10 (Discretizer: waveform → transcript +
numbered elements) then position 4 (Selector).
Width-is-cheap: 9–11 questions on one request.
Partial-speech wait is VOI (closed-set vs free-text).
Spoken confirm is not a grant. Overlay numbers are
exact, not a second model. Compose with item 21
(OCR). Full cards: `mixed-architecture.md`, `faq.md`,
`applied-mappings.md` §9, `validation.md`.
23. **Wrap-as-execution ALLOW/ASK/DENY** (reddpy/AgentGhost):
position 3 (constraint on the actuator path) then
position 4 (Selector on leftovers). Rules prove;
Jev remainder; ASK is an error not a log line.
Fail-closed on judge error. Distinct from actiongate
(evidence ≠ authority) and jev-use (fail-open).
Full cards: `mixed-architecture.md`, `faq.md`,
`applied-mappings.md` §7, `mappings.md` §8/§18.
24. **Application genre atlas** (@studio_yebisu):
position 11 (catalog of holes, not a score).
Stars are research-time. Same discipline as item 17
(class census ≠ bake-off). Full cards:
`mixed-architecture.md`, `faq.md`, `validation.md`.
25. **External pedagogy / how-to-apply** (@akshay_pachaar):
position 11 (explainer of already-owned placements,
not a new construct). LLM hammer; code owns
branches; schema-safe ≠ correct; shadow +
questions-as-code. 200×/400× are TypeSafe ceiling.
Full cards: `mixed-architecture.md`, `faq.md`,
`mental-models.md`.
26. **Boolean composition of soft Nouls** (uehaj/jev-semgrep):
position 5 (Comparator over lines) then code ∧/∨/¬
on *thresholded* bits — never multiply parallel p
(the logical-operator caveat above). Proposition ≠
embedding (contrast-set refund). Not a Gate (position
3). **≠** jev-combinators digital-design metaphor
**≠** semgrep.dev. Full cards: `mixed-architecture.md`,
`faq.md`, `applied-mappings.md` §4, `mappings.md` §4.
Loading