Conversation
…ctice Archer drop still Watch — no architecture rewrite. Novel vs the 11:59 digest: sqlite-jev in-engine vs jevql CLI; bitrate-advisor and the JOB planner as soft judgment inside a hard envelope; jev-routing host adapter; jev-claw classify-then-policy. Already-folded HIGH get frames only (memory gate, receipts economics, encoder vs decoder replica). Docs-only; no Jev wrapper. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com>
Place jaredpalmer/kev beside Laya / Nimble / Archer Watch: Qwen2.5-0.5B LoRA + pointer, Apache-2.0, POST /v1/systemone drop-in. Public gold, not a Jev teacher. Contrast vs TypeAR, encoder DeBERTa, proprietary Jev. Cite README isolation/ECE/acc/permute/IIA/forgery; laptop-local development/eval, not a knowledge substitute. Docs-only; no serve how-to. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com>
…n LoRA Cite Hub/GitHub receipts only. Archer 27B remains Watch. Skill cards carry mental models (omni decide without waiting, ECE/NLL/Brier bake-off, decision-token training, allowlist-as-proof, forced Choice, S1 keeps control), not a hit list or serve how-to. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com>
Cite README + replay receipt only. Mental models: soft rules vs linter, edit vs turn observation window, banded fail-open, rubric as artifact. jev-pref stays the contract; rh-guard stays eval-integrity. Not a hit list, not multimodal, no hook how-to. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com>
Do not rewrite §45. Hub jaredpalmer/kev-0.5b; PEFT FEATURE_EXTRACTION. Choice other must appear as a wrong alternative too — wellposed request lint is not enough. No invented NOTA rates. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com>
Docs-only fold of novel categorization/scoring placements into PR #2: extractive quotes + offline redecide, pointer-not-generator, contract- compatible /v1/systemone (stub until hf), structured observe→decide→act, dataframe accessors (jevframe beside jevpandas), route≠memory, advisory sidecar, structure induction, collab-arm measurement, AST∩semantic lint. Archer remains Watch. No wrappers or copied how-tos. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com>
Independent when-it-holds atlas (extractable-from-state), DMB frozen protocol vs constrained LLMs, jevals-data recompute-from-logs, and Kahneman S1/S2 cascade. Brief: ARC combinatorial negative, packed open-LLM System One, tiny SAN local drop-in, kev star delta. Archer still Watch. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com>
Place m-newhauser/gliner25-compaction as architecture notes: pointer compaction vs summarizers; fail-closed keep_full under a mutation envelope; same job as pi/fast-jev-compaction with a GLiNER backend; shadowMode default true. Not Jev, not multimodal. Docs-only. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com>
Docs-only architecture notes for latch (cluster-then-policy PASS/BLOCK), wakegate (fail-open VOI resume), GLiNER S1 indexer (10-50x unfilled), clear-head claim/evidence Stop, Harbor on/off routing, and cousins. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com>
Docs-only: sahibzada-allahyar/gliner2-ultrafast proves observe→score-among-candidates→code-acts is backend-agnostic (Jev Ultrafast ↔ local GLiNER2). Hybrid local decide + remote fill; DONE ≠ verified success. Numbered §52 after the 16:48 Boise watch. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com>
Docs-only: tamaratran/jev-pruner Noul-prunes command output after a hard ≤10k/JSON-diff envelope, before the main LLM sees it. Same extractive family as fast-jev-compaction / gliner25-compaction; different job. Fail-safe keep original; archive for recovery. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com>
Docs-only: trycua/cua libs/cua-s1 option-attention among observed elements (fill/check/click/skip). Plan ≠ execute; dry-run default; source-only. Not TypeSafe Jev — parallel S1 naming in CUA. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com>
…eval Fold jevify (uncalibrated local likelihoods), decision-native RAG, carryforward, hunch, explore-typesafe-ai, and jev-baselines-eval (honest negative + cascade sign-flip) into PR #2 skill cards. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com>
Docs-only: jev-sift classify-first (retrieve-wide → decide → evidence-set on agent I/O) and jevable.com class patterns (not a 342-title dump). No invented metrics; Archer still Watch. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com>
Docs-only: draft stack #2951–#2955 as harness productization of observe→score-among-candidates→code-acts. Extract judge + pick-and-copy uses their 37/75 ~0.5s vs 4.37s card; pick is a fast path, not a replacement. No invented metrics; Archer still Watch. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com>
…tness Docs-only skill/architecture notes for ask-jev-ai, jevgrep, egma-ai/jev-reviewer, jev-oxlint, jev-adaptive-thinking, and Jev-RAG (plus MED pointers). Harbor/jevals practice, not a hit list. TypeSafe Jev remains exemplar, not monopoly. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com>
…ation arena Interlock as ring-0 sensor≠policy (secrets never in agent; type-safe ≠ correct). Typed control plane around DSPy (drafts after route+action). Native-probability jev-arena plus sonar/vickrey/bracket fan-out. Stockfish owns truth / Jev owns judgment. Human-confirmed port-cleanup gate UX. Brief toolbelt; skip parody/farming. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why this PR exists
#1 merged as v0.3.0. This branch is the next hourly System One watch fold. Docs-only: skill, mental-model, and architecture notes — not a hit list.
Archer Hume 27B drop is still WATCH. Omni perception→decision does not wait for it.
Non-negotiable: exact work stays in code or policy; the model owns narrow judgment; never launder a soft Noul as a proof. No wrapper, no SDK/plugin how-to, no copied ports/key-file paths. No invented metrics. TypeSafe Jev is the exemplar, not the monopoly.
This update — 01:48 UTC 2026-09-19 / ~19:48 Boise 2026-09-18 (
research/notes.md§59)Docs-only fold. Do not rehash §50–§58. Keep Augustus as how-to-apply / mental-model / architecture / toolbelt + jevals/Harbor practice — not a thin Jev skill. Backend-agnostic categorization/scoring/decision-only class.
HIGH
somoore/interlock (Python; MIT; created 2026-09-19T01:41:58Z; 0★). Capability kernel: LLM ring 3 / Interlock ring 0. Secrets never enter the agent (canaries/placeholders). Closed action space. Jev is SENSOR;
policy.pydecides BLOCK/ASK/ALLOW. Anti-pattern: launch-week firewalls that ask “dangerous?” after the LLM already decided with real secrets in scope. Type-safe ≠ correct; irreversible stays behind a threshold and a human. Distinct from toolgate. 38-case local-judge set, not a blind paper. rh-guard crossover (different hole).manikanda-kumar/jev-dspy-control-plane (Python; MIT; created 2026-09-19T01:36:58Z; 0★). Benchmark-first: can a constrained classifier be a typed control plane around DSPy? Classifier → ontology validation → security override → confidence → state machine → tool allow-list; DSPy drafts AFTER route+action are fixed. OpenJEV / DSPy / JSON Schema share ontology. Offline heuristic + contract stubs = plumbing, not quality. Accuracy alone is not enough. Ax/DSPy stay LM-program knobs.
meetr1912/jev-arena (+ jev-sonar / jev-vickrey / jev-bracket) (Python; MIT; 0★). Native Jev probabilities, not verbalized confidence. Live theirs (
jev-1.13.0, 145 noul, 2 requests / 710 ms): Brier 0.0059, ECE 0.0620; overconfident in low bins. Fan-out as measurement economics. Sonar = heatmap-as-policy. Vickrey = threshold CDF; Jev never bids. Bracket = Brier vs Elo; live trailed Elo (honest).JoelLewis/game-coach (TypeScript; GPL-3.0; created 2026-09-19T01:08:53Z; 0★). Wave 0 PRD. Stockfish owns truth; Jev owns judgment; templates + capped writing model own words. Anti-soundness-theater alongside egma attention≠correctness. Treat PRD as Empirical-as-spec; shipped product is Hypothesis.
epiphany-dynamics/port-cleanup (Swift; MIT; created 2026-09-19T00:58:50Z; 0★). Jev recommends; human is the only kill trigger; identity re-check before SIGTERM; shields override; displayed explanations are app-owned mapped text, not raw model prose. Kill recs need conf ≥ 0.8. Gate UX + rh-guard cousin.
MED (toolbelt / patterns)
1jehuang/jev-pr-labeler (conceptual scope, not line counts). RubyBrewsday/jevcumber (
.featureonly; pointer among observed controls). shkumbinhasani/typedecide (class SDK; not on npm). douglance/jevon (CLI+MCP; key not in agent config). buberlo/dsh-jev (can only gate, never widen; not on npm). zaycruz/fast-jev-compaction-pi (pi port). planstack-ai/jev-tetris-benchmark (legal set in code; not a rigorous eval). fabricioctelles/modelsystem (modelsystem.one; 1★; not affiliated). emirbartu/opencode-system-one (fail-open plugin; license null; 1★). phanngoc/browser-ai (Go CDP; design done). acorn181/semantic-bookmark (user-authored semantic rules).Skip: SPFreedom/jef-mcp (parody), rayelzz/jevregist (account farming). Star spike this day: SemIf 1491→1606 (this pass 1607★); jevlike 851→896 (this pass 897★).
Reviewer entry points (this fold)
research/notes.md§59 — receiptsreferences/applied-mappings.md§7 — capability kernel vs toolgate; human-confirmed kill; §2 jevcumberreferences/mappings.md§8 / §18 — sensor≠constraint; closed action space / human killreferences/mixed-architecture.md— fail table + gallery (kernel, control plane, engine-owns-truth)references/validation.md— jev-arena Brier/ECE; dspy-control-plane metrics; tetris demoreferences/faq.md— type-safe ≠ correct; Jev is sensor not policy; Ax/DSPy knobs vs control plane; native vs verbalized confidencereferences/optimizer-integration.md— typed control plane around DSPyreferences/formal-methods.md— Stockfish truth / Jev judgment; anti-soundness-theaterSections changed (this fold)
SKILL.md;applied-mappings.md;mappings.md;mixed-architecture.md;faq.md;validation.md;optimizer-integration.md;formal-methods.md;mental-models.md;methods-catalog.md;toolbox-mapping.md;agent-self-assessment.mdnotes.md§59;sources.json(318 sources, 315 unique URLs, retrieved 2026-09-19T01:54Z);refresh-log.md;archive/findings.mdbatch #43CHANGELOG.md;README.md;docs/ecosystem.mdevaluate_decisions.pyPrior on this branch — 00:46 UTC (
research/notes.md§58)HIGH: ask-jev-ai public judgment wall (six parallel questions; policy-in-code; cost-to-1M from tokens; license null). jevgrep meaning-search without embeddings; 79% top-5 vs BM25 40% / grep 20% on stripped repos; keyword still wins exact strings. egma-ai/jev-reviewer attention P0/P1/P2 ≠ correctness; not choxos. jev-oxlint skills→oxlint AST∩remainder; Phoenix fixtures; not a hard gate. jev-adaptive-thinking session-sticky first-prompt; fail-closed fallback. Jev-RAG one-run ≥70% cost / 72% latency vs Spark rerank; full-context Spark still faster.
Prior on this branch — 00:48 UTC (
research/notes.md§57)HIGH: browserbase/stagehand #2955 (5/5 of #2951–#2955, all OPEN draft). Extract completion judge + pick-and-copy. Jev picks; code copies. Their card (gemini-3.8-flash, 25×3): 37/75 no-LLM ~0.5 s vs baseline 4.37 s; 69/75 vs 23/25 (92% both); LLM-off 36/75 — pick is a fast path, not a replacement.
Prior on this branch — 00:38 UTC (
research/notes.md§56)HIGH: jev-sift classify-first MCP (retrieve-wide → decide → evidence-set on agent I/O; mocks ≠ accuracy; no LICENSE). jevable.com living applied-mappings atlas (claimed 342 vs JSON-LD first page 36; class patterns, not a hit list). Draft-gate silence ≠ safer.
Prior on this branch — 17:48 Boise (
research/notes.md§55)HIGH: CUDA replica jevify (uncalibrated likelihoods ≠ Noul); decision-native RAG; carryforward; hunch; explore-typesafe-ai (not clinically validated); jev-baselines-eval both AMBIGUOUS.
Prior on this branch — 17:21 / 17:15 / 16:56 Boise
Cua-S1 specialist form (
notes.md§54). jev-pruner stdout prune (notes.md§53). gliner2-ultrafast observe→score-act (notes.md§52).Prior on this branch — 16:48 / 16:22 Boise; extractive; kev; Abide
latch / wakegate / s1-indexer / clear-head / jevons (
notes.md§51). GLiNER2.5 compaction (notes.md§50). Boundary map / DMB / dual-process (notes.md§49). Extractive quotes, jev-local stub, solari-reflex (notes.md§48). kev Hub/NOTA (notes.md§45). Abide soft-rule lint (notes.md§47). blackwood-rlcd / open-jev-laya-bench (notes.md§46).Verification
evaluate_decisions.py --self-testoksources.jsonparses (318 sources, retrieved 2026-09-19T01:54Z; 315 unique URLs — three pre-existing duplicate URLs left untouched)cursor/augustus-store-envelope-00b4; no duplicate PR