From 6488bc2501c48ee12383ff1e713b715e5d95e7d4 Mon Sep 17 00:00:00 2001 From: Tauan BF <11513929+tauanbinato@users.noreply.github.com> Date: Mon, 28 Sep 2026 19:31:48 -0300 Subject: [PATCH] Release 0.30.0 --- CHANGELOG.md | 22 ++++++++++++++++++++-- Cargo.lock | 2 +- Cargo.toml | 2 +- README.md | 2 +- npm/package.json | 2 +- plugin/.claude-plugin/plugin.json | 2 +- site/src/ci.md | 4 ++-- 7 files changed, 27 insertions(+), 9 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 1000e75..7d0fa7c 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -4,6 +4,12 @@ Notable changes to JevGate. Versions follow [Semantic Versioning](https://semver ## [Unreleased] +## [0.30.0] - 2026-09-28 + +0.30.0 carries the five versions of the roadmap, 0.26 to 0.30, in one release. A pull request check judges only what the change touches and fails only on rules measured right at least 80% of the time on projects JevGate was never tuned on; `jevgate hook` puts that gate in a coding agent's loop; each question's answer is cached apart and each finding says how often findings like it were right; a team's own conventions become questions that gate its code; and nine more languages are read, in preview. Each part below says what changed and how it was measured. + +### More codebases + More codebases: nine more languages are read, in preview, where JevGate's own rules never fail the default gate, and a syntax error leaves out the unit it sits in instead of its whole file. On the corpus, no file is skipped for want of a parser any more; 872 were. On 37 projects never used for tuning, the new languages' reviews were right from 45% (Bash) to 90% (Lua) of the time. - C, C++, Kotlin, Swift, Bash, Dart, Scala, Elixir and Lua are judged, in preview. They were discovered and skipped for want of a parser: 872 files of the corpus, among them whole projects (vapor, ktor-samples, phoenix_live_dashboard). A generic tier reads them through one tree-sitter tag query per language, written in the captures GitHub's code navigation uses (`@definition.function`, `@definition.class`, `@reference.call`), and a table per language names the nodes that hold statements, nest control flow and hold literals, and where its tests live. The grammars' own `tags.scm` tag what names a definition, not the definition (a C prototype's declarator, a Swift method's whole class), so the queries are JevGate's; a test checks that every definition those `tags.scm` find is a unit or owns one. These files get function simplification, file organization, shared logic and comments, and every request names the language; hardcoded values and security need a language's own sites and sources and are not asked. @@ -29,6 +35,9 @@ More codebases: nine more languages are read, in preview, where JevGate's own ru - The generic tier's files follow the same rule. Of the 34 that the generic tier alone skipped whole over a syntax error, 32 are judged for their intact units (C 24, Swift 6, Kotlin 1, C++ 1) and 2 stay skipped; 56 more name what they leave out; 2 are skipped now, a Swift benchmark whose one function holds the error (its comments had been judged as top-level code) and a Bash script with no unit. - The release binary grows from 26.4 MB to 46.1 MB (6.6 MB to 8.7 MB compressed with gzip -9) for the nine grammars; a clean release build took 28 seconds against 27 for 0.29 on a 12-core machine, the grammars compiling in parallel. + +### Your rules + A team's own conventions now gate its code. A convention is written as a yes/no question, or drafted from a line of the project's `AGENTS.md` for a person to accept, and JevGate asks it of every function, test, comment, documentation section, file or changed hunk it names, and gates, baselines and allows its findings like a built-in rule's. The question's examples, asked by `jevgate rules test`, show when a new model or a reworded question stops telling code that breaks the rule from code that keeps it, and five questions measured on corpus projects ship in a gallery. Proposed from ky's `AGENTS.md`, accepted and given one line of guidance, a question failed a pull request check on a change that broke ky's rule, passed once the change was fixed, and got all six of its examples right; quoted alone, without the guidance, the rule answered 0.78 against its threshold of 0.80, and the change passed. The same cycle runs through the binary against a scripted provider, where the agent hook also keeps a turn that breaks the rule working until the agent fixes it. - **Custom questions.** A convention is a yes/no question whose yes is a violation: `[[question]]` in `jevgate.toml`, or one per file in `.jevgate/questions/.toml`, with `question`, `background` and `guidance`, `unit` (`function`, `test`, `comment`, `section`, `file` or `hunk`), `paths`, `threshold` (default 0.80), `level` (`review`, `consider` or `note`; default `review`) and `next_step`. Each is the rule `custom/`, in the `custom`, `default` and `all` groups, named that way wherever a rule is, and listed by `jevgate rules` and the MCP rules tool; a check names each question a `rules` list in `jevgate.toml` leaves out. Every mistake is an error naming the question and its file, and `jevgate.schema.json` checks them in editors. At its threshold an answer is a finding at the question's level; at or below one minus it the unit is clear, and between them it is undecided. [See the guide](https://tech-byte-frontier.github.io/jevgate/custom-questions.html). @@ -47,6 +56,9 @@ A team's own conventions now gate its code. A convention is written as a yes/no - Proposed from ky's `AGENTS.md` ("Do not add special handling for `null`") and accepted as a review as proposed, a question answered 0.78 on a function that turns a `null` timeout off: undecided against its threshold of 0.80, so `check --base` passed, and `rules test` found it missing one of six examples from ky's code (0.75). With one line of guidance saying what gives `null` a meaning of its own, the same change failed the gate at 0.89, the fixed function cleared at 0.12, and all six examples were right. A proposal quotes a rule as written, so the proposal file, `rules accept` and the docs say to add guidance and a failing and a passing example, and to run `rules test`, before raising its level; `rules accept` says what a question that already fails the gate lacks. - **Question gallery.** Custom questions measured on real projects, for conventions linters cannot check. `jevgate rules add NAME` writes one into `.jevgate/questions/NAME.toml` offline, with the wording this version measured; it refuses an id `jevgate.toml` defines and a file that differs from the gallery's unless `--force` replaces it, leaves an identical one as it is, and warns when Git ignores the directory. Five ship, each asked of 6 to 12 projects without the built-in questions, with every finding labeled from the code (a debatable one counting as not right): at review, `todo-without-owner` (31 of 31 right at its threshold of 0.95), `swallowed-errors` (14 of 17, 9 of the right ones in one of the maintainer's projects), `resource-leak` (12 of 14, 9 in javavulnlab, an intentionally vulnerable application) and `thin-handlers`, asked of request handlers under common controller, handler, route and view paths (12 of 13, 11 in lobsters); as a note, which never fails the gate, `n-plus-one` (5 of 7). Any level but note fails the gate once added, and these counts are below the 20 findings on unseen projects that JevGate's own rules need before they fail by default, so `rules add` prints each question's numbers and whether it fails the gate, and the page says to try one with `--fail-on custom=report`. Ten more were measured and left out, under 60% right or with too few findings: state shared between callers (3 of 6), an error logged and also returned (4 of 7), flaky tests (6 of 11), test names that promise what the test does not check (7 of 17), leftover debug prints (5 of 13, and linters catch most), untranslated interface text (4 of 18), money in floats (6 of 50), stale comments (1 of 4), undocumented special return values (1 finding in 1,565 functions) and names that hide writes (none in 916). The page gives each question's projects, its right and wrong findings, and its file; about $0.47 on the corpus. [See the gallery](https://tech-byte-frontier.github.io/jevgate/question-gallery.html). + +### Cheaper, steadier, and how often each finding is right + Each finding now says how often findings of its rule and level were right on projects JevGate was never tuned on, in place of the probability of one answer; each question's answer is cached apart, so a reworded question is the only one asked again; and a function's source is sent once, with every rule's questions about it. With every rule and tests, the first pass of the corpus's 117 projects outside Bend sends 20% fewer requests and bills about 11% less input, and JevGate's own `--rule all` sends 37% fewer requests for about 16% fewer tokens; the default rules, of which only function simplification packs functions, plan exactly the requests 0.27 plans. The default gate is unchanged: replayed from the answer cache with the default rules and tests on the 94 labeled corpus projects, this release turns 131 shared-logic considers into notes, 116 of them outside test code, and changes nothing else, and the check still fails only on function-simplification reviews (20 of 23 right on unseen projects) and, once the documentation rules run, agent-context considers (22 of 24). - **Every finding says how often findings like it were right.** Every review and consider ends with how often findings of its rule and level were right on the 25 projects JevGate was never tuned on, in place of the probability of the answer that set its level: "Right 87% of the time (23 labels)." or, below 20 labels, "Not yet measured.", in the agent text, GitHub annotations and the job summary, GitLab issues and SARIF results; a law finding adds that it was labeled only on Bend 2 projects, which the table leaves out. The probability says how sure one answer was, not how often such findings are right: on those projects, reviews with a probability below 0.90, below 0.95, below 0.98 and above were right 55%, 46%, 56% and 61% of the time, none near the 80% a level needs to fail the check. The numbers are the labels `jevgate rules` shows. The JSON report gives each review and consider `precision` (`{"right": 20, "labeled": 23}`; none for notes, which are never labeled) and keeps `concern_probability`; SARIF results carry `precision` beside `probability`, and the HTML report shows it under each finding. The agent hook's finding lines put the sentence after the why, which they cut when it is long, never the sentence; the MCP tools' findings end their `message` with it and carry `precision`, and the server's instructions tell the agent to weigh a finding by it rather than by `probability`. Messages no longer end with the probability, in the JSON report too, and a warning whose rule is still being measured says why it does not fail without repeating the numbers. @@ -64,6 +76,9 @@ Each finding now says how often findings of its rule and level were right on pro - Security: a server template's code that writes client data unescaped is judged by injection only (every scriptlet of a JSP page by every security rule), so with injection off, such an ERB, EJS or Handlebars template was sent in a request that asked no question. Nothing is sent for it now. - Privacy and cost: what TypeSafe, OpenRouter and Vercel AI Gateway say about keeping what they receive and training on it, checked on 2026-09-28 against their own documents. TypeSafe does not train on inputs, keeps personal data "for as long as necessary", may use customer data in perpetuity to derive telemetry it can process without restriction, and offers zero data retention to enterprise customers. OpenRouter stores no prompts unless logging is turned on, and lists TypeSafe's Jev endpoint among its zero-data-retention endpoints. Vercel AI Gateway says it retains no prompts; its zero data retention with TypeSafe applies only when a team turns it on, and its model list marks Jev without it. + +### In the agent's loop + JevGate now works in a coding agent's loop. `jevgate hook` checks each edit and the end of each turn and keeps the agent working while findings fail the gate; `jevgate init --agent` sets it up in one command for Claude Code, Codex, Cursor, Gemini CLI and OpenCode, a Claude Code plugin bundles it with the MCP server, and the MCP tools return structured results, with the units Jev left undecided for the agent to verify. A turn is judged as `check --base` judges a change, from a snapshot taken when the turn began, and blocked only on what the gate fails. Replayed on 30 edits of 10 corpus projects in 9 languages, each inserting one comment line inside a function, the hook answered an edit in 1.45 s at the median and 2.2 s at the 95th percentile when the edit's requests were new, and in 0.31 s and 0.52 s from the cache; 3 of the edits gave the agent findings and no turn was blocked, where judging the edited files whole under 0.25.0's gate had given findings after 13 edits and blocked 4 turns. Writing a new file, which is judged whole, took 1.5 s at the median over 10 files of 26 to 1,013 lines, and 5.0 s for the largest (60 requests, every rule with tests), 0.28 s from the cache. The replays cost $0.025. A scripted session is blocked, fixed and let through, in process and through the binary against a scripted provider, and an outage, an HTTP 402 or a missing key never blocks the agent and is always said. - **`jevgate hook`** puts JevGate in a coding agent's loop, as a hook of Claude Code (and Devin CLI), Codex, Gemini CLI, Cursor, Copilot CLI or VS Code, and of OpenCode through a plugin; VS Code names its edit tools its own way, not verified yet, so there only the end of a turn is sure to be checked. The agent is detected from the event, or named with `--agent`. @@ -93,6 +108,9 @@ JevGate now works in a coding agent's loop. `jevgate hook` checks each edit and - The JSON report says what each undecided unit left open. Each entry of `dimensions.*.undecided` gains the unit's `fingerprint`, made as a finding's is, its `locations`, and `open`: each question it left undecided as it was asked (`text`), the state paths the question names (`evidence`, such as `functions[0].source`), what each answer means (`options`) and the answer itself. A question answered again by a recheck or a trace is quoted as that follow-up asked it. On the 117 corpus projects replayed from the answer cache, all 1,701 undecided units quote every question they left open, and the reports grew by 1.3%. Nothing is asked again, and no finding changes. - Reruns: the [versions and stability](https://tech-byte-frontier.github.io/jevgate/stability.html#reruns-of-an-unchanged-commit) page says why a rerun of an unchanged commit sends no request and reports the same findings, and what makes one ask again. `rerun.sh`, published beside it, shows it on any repository: it checks twice and compares the two reports, each file's status, findings and raw answers. On 14 corpus projects in 9 languages (2,839 files, 3,870 findings, 129,403 answers), every rerun sent no request and matched. + +### A gate on the change, and more ways to run it + A pull request check now judges only what the change touches and fails only on what has been measured right, and OpenRouter and Vercel AI Gateway keys work as TypeSafe keys do. A check with the defaults (`jevgate check --base`, as the action runs it), replayed from the answer cache on the last commit of 118 corpus projects (90 open-source, 28 of the maintainer's own private repositories), failed 18 of them with 0.25.0 and left one more incomplete (its replay lacked a cached answer), on 61 reviews, 32 of the 46 labeled right (9 wrong, 5 debatable). With 0.26 it fails 9, on 11 function-simplification reviews, 9 of them right (none wrong, 2 debatable). 8 of those 9 projects, and 10 of the 11 findings, are the maintainer's own; the other is ky's `Ky` constructor, labeled right. It reports 29 reviews and 70 considers where 0.25.0 reported 61 and 140. On the 25 of those projects never used for tuning, it fails 2 instead of 4. - **The default gate fails only on what is measured right.** By default, only the rules and levels measured right on projects JevGate was never tuned on fail the check. The default gate level is now `mature`: a rule's reviews or considers fail the check when at least 80% of them were right on the 25 projects never used for tuning (11 held out, 14 fresh), over at least 20 findings labeled from the code. Every other finding is still reported, marked as still being measured, and the check passes. Two levels are mature: function-simplification reviews (20 of 23 right, 87%) and agent-context considers (22 of 24, 92%). Agent context is a documentation rule, which runs only when selected: with `--rule documentation` or `--rule all`, its considers fail the check, the first consider level to do so. Its 22 right findings on unseen projects come from 4 of the maintainer's own repositories. On the unseen projects' full runs, the default gate failed on 122 findings, 57% of the labeled ones right (64% leaving the 13 debatable ones out), and on 17 of 22 projects; it now fails on 23, 87% right, and on 9 projects. On the 72 projects used for tuning: 468 findings at 73% to 95 at 83%, and 61 projects to 28. The 49 right reviews on unseen projects that no longer fail the check are still reported. The concern probability could not do this: unseen reviews were right 55%, 46%, 56% and 61% of the time with a probability below 0.90, below 0.95, below 0.98 and above. The labels were made on 0.24.1's findings and joined to 0.25.0's, replayed from the answer cache; 0.25.0 had turned 17 labeled security findings into notes, one on an unseen project, none in a mature level. Bend 2's labels are kept apart, so `tests/laws`, which judges only Bend 2 code, has no row: its findings are reported without failing the check, and say they were labeled only on Bend 2 projects. Right / labeled on unseen and tuned projects, a debatable label counting as not right: @@ -145,7 +163,6 @@ A pull request check now judges only what the change touches and fails only on w - Retries: an answer worth retrying (a rate limit, overload, or a server or gateway error) is sent up to 6 times instead of 4, pausing 1, 2, 4, 8 and 8 seconds, each up to a quarter longer (the first three as before), or longer when the provider asks, up to 30 seconds as before. On 2026-09-28 TypeSafe answered 503 to about two attempts in three for at least ten minutes, directly (22 of 34 attempts) and through OpenRouter (141 of 224), each failed attempt taking about 10 seconds: with 4 attempts, a self-check with a TypeSafe key and all four canaries through OpenRouter ended incomplete, 1 of 13 and 16 of 99 requests having given up. Simulating the queue, 6 attempts leave 6% of requests unanswered in that brownout instead of 16%; when 1 attempt in 5 fails, a 1,000-request run completes 95% of the time instead of 21%; and a provider failing every attempt is outlasted for 26 seconds instead of 8. They cost time only while attempts fail: a hard outage takes up to three times as long to end a run incomplete, 8 minutes instead of 3 for 100 requests. A timed-out request is still sent twice at most, and one whose connection failed before sending 4 times. - The test suite leaves alone the repository it runs in. Git exports `GIT_DIR` to a hook, `git rebase --exec` or `git bisect run` in a linked worktree, and `GIT_INDEX_FILE` to a pre-commit hook; `cargo test` run there made the tests' `git init`, `add` and `commit` act on that repository instead of their temporary projects, which set `core.bare = true` in a clone's shared configuration and committed a test's files into a worktree. The tests' Git, and a check's in the unit tests, now runs without the variables that point Git at a repository, and `tests/lint_policy.rs` rejects starting Git anywhere else. A check still honors them: `jevgate check --base HEAD` in a pre-commit hook reads the index being committed, and a repository kept apart from its work tree is found through `GIT_DIR`. - ## [0.25.0] - 2026-09-27 Fixes from running JevGate on widely used projects under daily development (rtk, headroom, paperclip, hermes-agent, cc-switch, freellmapi, herdr, multica, OmniRoute, dify, openclaw, n8n), each checked against the code. @@ -407,7 +424,8 @@ These changes come from running 0.11.0 on six open-source repositories it had ne - First release: the maintainability CLI. -[Unreleased]: https://github.com/Tech-Byte-Frontier/jevgate/compare/v0.25.0...HEAD +[Unreleased]: https://github.com/Tech-Byte-Frontier/jevgate/compare/v0.30.0...HEAD +[0.30.0]: https://github.com/Tech-Byte-Frontier/jevgate/compare/v0.25.0...v0.30.0 [0.25.0]: https://github.com/Tech-Byte-Frontier/jevgate/compare/v0.24.1...v0.25.0 [0.24.1]: https://github.com/Tech-Byte-Frontier/jevgate/compare/v0.24.0...v0.24.1 [0.24.0]: https://github.com/Tech-Byte-Frontier/jevgate/compare/v0.23.1...v0.24.0 diff --git a/Cargo.lock b/Cargo.lock index 3111210..22c4433 100644 --- a/Cargo.lock +++ b/Cargo.lock @@ -966,7 +966,7 @@ checksum = "8f42a60cbdf9a97f5d2305f08a87dc4e09308d1276d28c869c684d7777685682" [[package]] name = "jevgate" -version = "0.25.0" +version = "0.30.0" dependencies = [ "anyhow", "clap", diff --git a/Cargo.toml b/Cargo.toml index 243bc80..32757fa 100644 --- a/Cargo.toml +++ b/Cargo.toml @@ -1,6 +1,6 @@ [package] name = "jevgate" -version = "0.25.0" +version = "0.30.0" edition = "2024" rust-version = "1.90" description = "Code-review gate for CI and coding agents: asks TypeSafe Jev small questions about functions, files, tests and docs, and reports maintainability, test, security and documentation findings with locations and how often findings like them were right" diff --git a/README.md b/README.md index 2fef51c..1b26f8f 100644 --- a/README.md +++ b/README.md @@ -74,7 +74,7 @@ jobs: - uses: Tech-Byte-Frontier/jevgate-action@v1 with: api-key: ${{ secrets.TYPESAFE_API_KEY }} - version: 0.25.0 + version: 0.30.0 ``` It reviews only what the pull request changed (the functions, tests and comments on changed lines, and copies where either copy changed), annotates each finding on its line and writes a job summary; unchanged code is answered from the cache for free. [Continuous integration](https://tech-byte-frontier.github.io/jevgate/ci.html) covers pre-commit, other CI systems, pull requests from forks, budgets and a gate policy the change cannot edit. diff --git a/npm/package.json b/npm/package.json index ca19d01..45b4b24 100644 --- a/npm/package.json +++ b/npm/package.json @@ -1,6 +1,6 @@ { "name": "@tech-byte-frontier/jevgate", - "version": "0.25.0", + "version": "0.30.0", "description": "JevGate, the code-review gate for CI and coding agents: this package downloads the matching release binary, checks its SHA-256, caches it and runs it", "keywords": [ "code-review", diff --git a/plugin/.claude-plugin/plugin.json b/plugin/.claude-plugin/plugin.json index 66ae6f7..bf4d0b4 100644 --- a/plugin/.claude-plugin/plugin.json +++ b/plugin/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "name": "jevgate", "displayName": "JevGate", - "version": "0.25.0", + "version": "0.30.0", "description": "JevGate in the agent's loop: checks each edit and the end of each turn, keeps Claude working while findings fail the gate, and adds JevGate's MCP tools and a skill for acting on findings. Runs the jevgate command, 0.27 or later, which you install separately.", "author": { "name": "Tech Byte Frontier", diff --git a/site/src/ci.md b/site/src/ci.md index 2f0d43d..63ea17a 100644 --- a/site/src/ci.md +++ b/site/src/ci.md @@ -17,7 +17,7 @@ jobs: - uses: Tech-Byte-Frontier/jevgate-action@v1 with: api-key: ${{ secrets.TYPESAFE_API_KEY }} - version: 0.25.0 + version: 0.30.0 ``` The action installs a checked release binary, keeps `.jevgate/cache` in the Actions cache and runs `jevgate check --base --format github`; `args` passes more flags, such as `--rule default --rule security` to add the security rules to the default ones. Naming a rule replaces the selection, and a selection without a mature rule level never fails the default gate: `--rule security` alone reports security findings without ever failing the gate. It runs on Linux, macOS and Windows runners. For an OpenRouter or Vercel AI Gateway key, leave `api-key` out and set the key's variable in the step's `env`, such as `OPENROUTER_API_KEY: ${{ secrets.OPENROUTER_API_KEY }}`, with `version` 0.26.0 or later; jevgate-action 1.2 adds `api-key-kind: openrouter` or `api-key-kind: vercel` for the same. @@ -66,7 +66,7 @@ Before each commit, with [pre-commit](https://pre-commit.com), review what is st ```yaml repos: - repo: https://github.com/Tech-Byte-Frontier/jevgate - rev: v0.25.0 + rev: v0.30.0 hooks: - id: jevgate-system # the jevgate on PATH; `jevgate` builds it with Rust instead ```