Conversation
Refine shared-logic classifications and bound optional evidence excerpts. Add file pagination and estimated batch cost to HTML reports.
Omit trailing occurrence regions when the combined request would be rejected, and keep the general maintainability questions in that same request.
When the cascade is enabled, an uncertain result gets one follow-up containing only the undecided operation, the repeated lines, or the file without the other evidence. The same question is asked again and replaces the status only at the existing 0.80 threshold.
Separate operational scripts, declaration files, and tests from the source the gates judge. Mixed files keep tests out of the application request, and --include-tests judges test code on its own.
Remove --classification-cascade. Every check routes repeated occurrences and rechecks an uncertain dimension. The specialist comparison still keeps the general verdict when the route is mixed or unresolved.
System One questions now carry labeled boundaries, so review and clear are not one paragraph. A file that cannot fit a request is needs-context with its size and operation names, and a repeat check reuses the last report instead of sending the tree again.
Local analysis (units, member groups, Type-2 clone candidates, test map) builds small units; each request asks short literal questions about one function pack, file outline, candidate pair or test pack. Composition is pure and every review carries a finding. Adds test value and redundancy rules, --fail-on with a baseline, schema v2, transport retries with a shared cooldown, and cache entries that do not expire for pinned models. Removes the role cascade, --roles-only and the roles evaluation script.
Function simplification asks two Scores over the same state: a task Score with structured levels whose top level raises a review, and a one-job Score that defines a task and can clear while the task Score does not lean toward several tasks. Weak test-value signals raise a consider without blocking a clear. On the frozen set, uncertain function units fell from 55% to 18% and uncertain test units from 85% to 16%.
Questions: replace task and purpose counting with benefit Scores for functions and files, and clear a unit when its actionable top level is ruled out at 0.80. Ask flattening only for deep nesting or long branch chains, word the own-logic test question literally, and send tests one per request. Evidence: import-aware callers, class and constructor links, callbacks registered through calls, Python unittest and pytest structural tests, generated-code headers, clone groups instead of pairwise findings, a three-statement clone minimum, a 100-line floor for file organization and an outline recheck with the file's source. Copies inside one test raise at most a consider. Transport: identify as jevgate/<version>, name provider edge blocks, and stop only after three consecutive ones. Version 0.2.0.
Self-review: fix all 13 review findings and most considers. Split long functions into named steps (planning, clone search, purpose decisions, request dispatch, report output, the local server), give the parser cache, test location, nesting, token budget, decision policy, provider errors, upload boundary and finding wording their own modules, share test helpers and make repeated test cases table-driven. Output: a reader that closes stdout early (`| head`) ends the output without a panic, and the exit code still reflects the gate; messages go through write helpers that ignore closed streams. Keep src/html_report.rs out of uploads: its escaping test holds an XSS string that the provider's edge firewall blocks. Version 0.2.1.
Composition: where a Score's middle level says the code reads well as it is (splitting, flattening, moving members), a consider now also needs the top level at 0.50; middle mass alone is an optional note. Notes are listed only with --verbose and never fail the gate. Copies whose every site is inside test cases are one level lower, and a file split is suggested only when the proposed group has callers of its own in other files (unknown callers keep the answer). Location: a split review or consider gets one follow-up Choice among the function body's top-level blocks; a statement that wraps most of the body is read through. The chosen block becomes the finding's first location. Report: one row design for files and their findings, strongest first; no separate ranked list; the notice appears only when a run failed. Self-review: split units/mod.rs into plan types, request evidence and answers with follow-ups; split mixed tests and share repeated setup.
Cargo.toml denies the unused lint group, dead_code and clippy::too_many_arguments, and denies #[allow] attributes: an exception must be an #[expect] with a reason. tests/lint_policy.rs rejects any allow or expect of dead_code, unused, too_many_arguments or complexity. The lints are deny rather than forbid because derive macros emit their own allows. Removed the four existing allows: the token budget travels in FileContext, a Questions builder owns each request's question bodies and their mapping to units, and the parser passes a definition's syntax nodes as one Definition.
…l cases New rule maintainability/hardcoded-values. The parser lists each function's literal values (skipping 0, 1, 2, one-character strings and escapes, documentation, attributes and imports) and each file's module-level constants. Per function, one pack of up to eight asks whether a value must change in another environment, whether naming a value would help a reader (both Scores with the note tier), and whether the function special-cases a particular user, account or record (Noul). Constants are asked only about the environment. There is no recheck: a per-value recheck added more false findings than it resolved. The environment criteria name what is not environment-specific (the program's own routes, project-relative paths, public addresses): with only "URL, path" examples, routes and repository files were flagged. Development results on a labeled set of 96 units from four codebases, a real hardcoded-values refactor and ten controls are kept in the ignored evaluation directory. Self-review: name the values the new rule found (backoff jitter, link weights, outline clip limits, response tolerances, token bounds, clone placeholders) and move a definition's signature and doc line into analysis/summary.rs.
Each rule's result now lists its undecided units (name, line and the questions whose answers stayed split), so an uncertain file explains itself instead of showing only a badge. Only the questions that decide a unit are listed: weak test signals and the required-repetition check never leave a unit undecided. The report counts them in the file row and lists them in its detail; --verbose prints them under each rule. Also name the separator of hashed identities once (fingerprints are unchanged) and give the hardcoded-values rule its name in the report.
Copied libraries are skipped as vendored whatever their size, test calls inside library code no longer make a file a test file, and security reads quoted SQL identifiers and browser requests as handled and asks where error text goes before calling it a leak. The README lists vendored files.
…ide the read arm `if let` guards in match arms are not stable in Rust 1.90, the declared rust-version, so the msrv CI job failed from b12028c onward. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… root On macOS the temp dir is under the /var symlink, and sources behind symlinks are refused, so 76 unit tests failed there while Linux CI passed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
0.12.0 did not build on Rust 1.90, its declared minimum; 0.12.1 does. Test projects also resolve their temp directory, so the suite passes on macOS. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Pushing an annotated tag vX.Y.Z on main runs the CI checks, including the Rust 1.90 build, then publishes to crates.io through Trusted Publishing and creates the GitHub release from the tag's message body. 0.12.0 was published by hand while its msrv job was failing; a release can no longer skip the checks. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
C# (tree-sitter-c-sharp 0.23.5):
- Methods, constructors, destructors, properties and indexers with bodies
are units owned by their class, struct or record, inside block or
file-scoped namespaces; records, interfaces, enums and delegates are
types; local functions are functions.
- Top-level statements are the program's setup; minimal API routes and
inline middleware (`app.MapGet("/orders", …)`, `app.Use(…)`) are units
named by their registration, while configuration lambdas stay in setup.
- `using` directives are imports, and a file reaches a class it names (or
its `I…` interface), since C# imports whole namespaces.
- Member access, generic and conditional calls are callees; foreach,
switch expressions, using and lock nest.
- `const` and `static readonly` fields are the constants of hardcoded
values; interpolated strings are built text, and one that only joins
values (`$"{baseUrl}{path}/{id}"`) is no value of its own.
- `new …Exception(…)` and any object a `throw` creates are created errors;
XML documentation tags are stripped from summaries and attribute lists
from signatures.
Tests: classes of xUnit, NUnit or MSTest tests ([Fact], [Theory], [Test],
[TestCase], [TestMethod], or a [TestFixture]/[TestClass] class) are
structural tests wherever they are, and C# files of test projects named
like `Shop.Tests` are tests. `.Designer.cs` and `.g.cs` files are
generated, and a C# file of only usings and namespaces has no code.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Security questions about C# files carry ASP.NET Core's names for what they ask, as examples of each answer: FromSqlRaw and SqlCommand against FromSqlInterpolated, Process.Start of a shell, certificate callbacks that return true (and token options such as ValidateIssuer or RequireHttpsMetadata, which are not certificate checks), AllowAnyOrigin with credentials, cookie options, exceptions a function lets propagate, `exception.Message` in an exception middleware, and seeding methods. Parameters of controller actions come from another party. Other languages keep their wording: their requests are unchanged. C# traces ask four more checks: - a type named by input or chosen by deserialized data (CWE-502); - a developer exception page outside development (CWE-489); - token signature or lifetime checks turned off, or a token that is only decoded (CWE-347); only the issuer or audience check is not one; - a signing or encryption key written in the code (CWE-321). The trace shows the `const` and `static readonly` fields the code names, often in another file (`AuthorizationConstants.JWT_SECRET_KEY`). A reset key derived from an email address counts as a guessable secret. Error handlers: UseExceptionHandler with a handler, IExceptionFilter, ExceptionFilterAttribute, IExceptionHandler and middleware classes whose Invoke catches what the pipeline throws are judged once each, with the program's exception classes (primary-constructor classes included). A broad weak-setting answer that no specific check leans toward names no setting to change and is at most a note. On eShopOnWeb and dvcsharp-api the uncertain rate is 0.7% and 1.9%; the regression corpus (jevgate, fastapi-template, chi) reports the same as main, with every request answered from main's cache. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… public secrets Next.js was not a listed framework. On a labeled fixture of Next.js route handlers, Server Actions, pages/api, middleware and next.config (12 real issues), JevGate found 8: a Server Action splicing its argument into sql.raw was a consider on "its parameters", open redirects were notes about a URL "it requests", a NEXT_PUBLIC_ secret was not asked about, and next.config CORS headers were never sent, since they call nothing. After this change all 12 are reviews. - Files carry their Next.js role as `file.framework` beside path and language: route handlers, pages/api routes, middleware and proxy, error boundaries, pages and layouts, next.config (paths within a package that depends on `next`), and Server Actions or client components from 'use server' / 'use client' directives in any package. Every question's note points to it. - Injection asks one more literal check, a redirect target from a variable (CWE-601); parameters alone stay a note, like paths and URLs. - The SQL check names binding tagged templates (Drizzle and postgres.js sql``, Prisma $queryRaw``) as handled and $queryRawUnsafe and sql.raw as not; the markup check names dangerouslySetInnerHTML, which is now a site; the code check no longer reads a query as evaluated code. - Unsafe settings ask whether a secret comes from an environment variable the build puts into browser code (CWE-200). - A next.config file's setup is every top-level statement that holds an object, with its innermost objects as sites. On next-saas-starter, nextjs-subscription-payments and a slice of umami the error-detail reviews on Supabase auth messages became considers, and a Server Action that redirects to any path its caller names is now named as a redirect note. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
On next-saas-starter, nextjs-subscription-payments and umami (src/app, src/lib, events queries), 56 files stayed uncertain on main and still 56 with the Next.js roles (4/5/47, then 7/5/44); most were client components that navigate to fixed paths or render values as attributes, fetch helpers of client components, and route handlers whose CORS or logging split. They are now 34, and every rule stays under 3% undecided units. - Each undecided security check gets its own settle Choice, asked only while it is undecided and able only to clear it: where redirect targets come from (written in the code, returned by the program's own server, passed by callers, checked, or no redirect), how markup is rendered, which sites may send credentialed requests, what the logs write, and where code that requests a URL runs. The last is also asked for a URL note: a note that a client component's fetch "places a parameter into a URL it requests" only puzzled readers. Offered beside "a whole URL handed to it", the browser lost for such helpers, so it is a Choice of its own. - The redirect check counts only targets a request carries: umami's short link redirect to the destination its owner saved was a review. - The Next.js role goes only to security and hardcoded-value questions. On split and outline questions it moved answers without informing them (three functions turned uncertain). - `route.test.ts` and `page.stories.tsx` no longer take the role of a route or page. A Choice about what a query builder joins into SQL was tried for considers on parameters and dropped: it cleared a sort column taken from the request as readily as clauses with placeholders. On the regression corpora the only changes are notes: a split note on jevgate's `status` and a Jinja rendering read as evaluated code in fastapi-template are gone, and chi's Profiler and RedirectSlashes redirects are named; uncertain files went 34/11/6 to 34/9/4. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Django code (Python that imports Django or Django REST framework, and settings modules) is asked its own wording and checks; other code keeps the common ones. - Views are sent with the URL routes that reach them, the templates they render with `|safe` or autoescaping off, and the module constants they name; management commands are marked as run by hand. - Injection checks name raw SQL (`raw`, `extra`, `RawSQL`), `mark_safe`, the storage API, open redirects and deserializers (`pickle`, `yaml.load`). - A settings module is one unit whose statements are its settings, secret literals redacted, sent with the lines that select it (`DJANGO_SETTINGS_MODULE`) and the settings modules that import and override it. Debug mode, `csrf_exempt` and literal secrets are checked; a weak setting no check names is at most a note. - `handler500`-style error views, middleware `process_exception` and Django REST framework's `EXCEPTION_HANDLER` are error handlers. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- An injection note in Django code that no check found gets the settle Choices, as an uncertain unit does: redirects to a view's own paths with ids in them no longer leave "values from another party" notes. - A `setdefault` of DJANGO_SETTINGS_MODULE is shown as only a default, so shared settings that a production module sets again are not read as deployed. - In a settings module, assignments into a setting such as `OPTIONS["ssl_cert_reqs"] = None` are setup and rank right after the security settings as sites. - The CSRF check asks about views acting for the session's user; the Django redirect check counts targets the request carries, not stored records; the Django error-handler question names Django REST framework's exception_handler. - The Django checks join main's language-scoped checks as variants of the common ones with the same id, and the Django markup Choice uses the common options. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Ruby (`.rb`, parsed with tree-sitter-ruby):
- Methods (`def`, `def self.`, `class << self`, `define_method`, `private
def`) are units owned by their class or module; a constant bound to a
lambda is a function; a class or module without methods is a type.
- A block passed at file or class level is a unit named by its call
(`get('/invoices')`); a block that holds definitions (`helpers do`) is read
for them; `RSpec.configure` and test declarations are left to the test
rules, keeping only the helper methods defined inside them.
- `require`, `require_relative`, `load` and `autoload` are imports and not
values; constants in the file, module and class bodies feed hardcoded
values; `raise`/`fail` are created errors; `if`/`elsif` chains, `unless`,
`case`, loops, modifiers and blocks count as nesting; bodies, blocks and
clones read Ruby's statement lists.
- `Invoice.new(...)` is recorded as a call to `Invoice`, like `new Invoice()`.
Tests:
- RSpec groups and examples written as statements (`describe`, `context`,
`it`, `specify`, `its`, titled by a string or not at all) and classes whose
superclass ends in `Test`, `TestCase` or `Spec` with `test_*` methods or
Rails `test "..." do` blocks are structural tests; `*_spec.rb` files and
Ruby files under `spec/` or `step_definitions/` are test files.
- Each case gets its groups as its suite, examples titled alike in different
groups are named with their innermost groups, and a case calls each bare
name it does not bind, plus the calls of the `let` and `subject`
definitions it reads. Its hooks are its groups' `before`, `around`,
`setup` and `let!`, and the lazy definitions it reads.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… adds On sinatra (Minitest and RSpec) and factory_bot (RSpec), a third of the Ruby tests that were rechecked stayed undecided: the recheck showed every hook of the file, with the file's setup described as mocks, so factory definitions and `mock_app` routes read as mocks. Copied RSpec examples for an alias and its original (`each` and `each_pair`) or for two predicates of one record were reviews. - A Ruby test is sent with the groups it is declared in (`suite`), since an RSpec example reads as a sentence continuing them and the outer group often names the class under test. - Its recheck shows what runs for it: its groups' `before`, `around` and `setup` hooks, `let!`, the `let` and `subject` definitions it reads, then the test helpers it and those hooks call, from its own file or the nearest support file sharing a directory with it (an RSpec group's methods stay in its file). The note says a value built there is input to the code under test unless a mock returns it. - The mock-only question counts a value the code builds from definitions, routes, records or settings the test gives it as the code's behavior, even when that input spells out the expected value. On the regression corpora this also cut chi's undecided Go tests from 12% to 5-7% and jevgate's from 3% to 1%. - Ruby test pairs carry their groups and hooks when these differ, and are also asked whether each test checks something the other does not; "one adds nothing" is a review only when that is ruled out at 0.80, otherwise a consider. - README lists RSpec, Minitest and Ruby DSLs; the cascade describes the Ruby test evidence. Question wording version 8. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Overlapping test pairs were grouped by subject alone, so two pairs of one subject that share no test became one "4 tests of `get` overlap" finding. On sinatra that joined rack-protection's redirect and deny tests, the `respond_to` and `respond_with` copies of five different tests, and 17 etag tests across three request kinds; on factory_bot, lint's raise tests with its strategy tests. On fastapi-template it joined the superuser and current-user reads of a user with two email-exists updates. - A group is now the tests a chain of overlapping pairs connects, per subject; the pairs themselves are unchanged. On sinatra 19 groups became 14 (the etag tests are five groups of one test across safe, idempotent and post requests), on factory_bot 15 became 12, and fastapi-template's two mixed groups are gone. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
PHP files were skipped. On DVWA (intentionally vulnerable), Slim-Skeleton,
the Laravel skeleton and BookStack (app/Uploads, app/Http, app/Access) they
are now judged: 991 security units on DVWA with 1.7% undecided, 1407 on
BookStack with 1.2%. Every vulnerabilities/*/source/low.php of DVWA with
a vulnerability in its code is found (or the index.php that runs it, for
file inclusion and DOM XSS), and the review findings are 93% right by hand
labels on DVWA; BookStack has one review left.
Parsing (tree-sitter-php, the PHP-with-HTML grammar):
- Functions, class methods, closures and arrow functions are units;
route closures (`$app->get('/users', function …)`, `Route::post(…)`)
and configuration closures (`return function (App $app) {…}`) are
named by their registration.
- A file's top-level statements outside functions and classes, with its
`<?= … ?>` echoes, are one more unit, `top-level code`: a page script
reads the request and writes the response, so every security rule
judges it like a function.
- `use`, `require` and `include` are imports; member, static and nullsafe
calls and `new` are callees; `.` joins text and double-quoted strings
and heredocs that interpolate are built text; `echo`, `print`,
`include` and backticks are sites; `new …Exception(…)` is a created
error.
- Tests: `test…`, `@test` and `#[Test]` methods of a class extending a
`…TestCase`, and Pest `test`/`it` calls; `…Test.php` files are test
paths.
- Error handlers: `set_exception_handler`, and the `respond`, `render`
and `register` methods of subclasses of Slim's `ErrorHandler` and
Laravel's `ExceptionHandler`.
Questions (src/units/questions/php.rs) read in PHP's own terms, through
the same rewording as C#; every other language keeps its wording:
- The SQL check counts driver escaping inside quotes and numbers as
handled (DVWA's escaped and quoted guestbook inserts were reviews);
shell, code, markup, path, URL, redirect, TLS, hashing, random and
cookie checks name PHP's functions.
- `echo` is the response, not a log; a page that only calls
`generateSessionToken()` or a session helper is not judged for it; the
message of the program's own exception class, caught by name, is its
own text (four BookStack upload controllers were error-detail reviews);
the origin counts only variables placed into the text or path, not an
uploaded file's contents (a BookStack helper was answered for the
upload); Laravel's `Str::random` is cryptographic.
- `unserialize` and uploaded file names are PHP checks, asked only of
source that names them. The first shares the `deserialize` kind of
Django's pickle check: one CWE-502 row whose remedy names both.
- Three settle Choices of their own, asked whenever their check is not
clear and able to clear a found concern: what a unit joins into HTML
unescaped, what its command lines hold, and where its paths come from.
A markup check whose Choice names a request value or stored record is
a review whatever the origin question said. The settle table now names
the files each Choice is planned for, so how markup is rendered is not
asked of PHP.
- A page script's consider or note names values whose origin it does
not show, not parameters; an injection finding with several checks
names each.
On the regression corpora (jevgate, fastapi-template, chi) the review and
consider findings are the same as main's; DVWA and Slim-Skeleton keep
every finding and hand label, with main's wording for weak settings.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Java files were skipped as unsupported: on spring-petclinic and jsoup, 258 of 261 files were skipped and nothing was judged. - tree-sitter-java 0.23.5 parses `.java`. Methods, constructors and record constructors belong to their class, interface, enum, record or enum constant with a body (`OPEN::enter`); annotation types are types. Imports name their class or static member; a wildcard names none, and a class of the same package counts as imported. - Units carry Java calls (`findById`, `new ArrayList<>()`), nesting (enhanced for, switch expressions, try-with-resources, synchronized and lambdas with a block body), numeric literals, `String.format`, `formatted` and `printf` as built text, and `new …Exception(…)` or any `new` or call a `throw` makes as created errors. - `static` fields and interface constants are constants. `equals` and `hashCode` overrides are boilerplate: they offer no copies and no values to name. - A class is test code, whole, when it holds a JUnit 4 or 5, TestNG or jqwik test method (a composed annotation ending in `Test` included), a lifecycle method such as `@BeforeEach`, a `@Nested` test class, or extends JUnit 3's `TestCase`. `…Test`, `…Tests`, `…TestCase` and `…IT` files are test paths. A test's setup is the class head with its mocks and its `@Before…` methods. - A test outline names a subject by the one type that owns it (`StringUtil::isBlank`); the top-level test class is not a suite, `@Nested` classes are. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A file whose source is too long to send whole got no recheck, so its undecided first answer on splitting was never followed up and the file stayed uncertain. On jsoup that was Element.java (2,211 lines) and its largest test classes, such as ElementTest.java (3,621 lines). - The kind question falls back to the outline alone when the file's source does not fit, and is asked after an undecided first answer when there is no recheck. The kind is read beside that first split. - On jsoup, Element.java is one type and ElementTest.java tests one subject: file-organization units left uncertain went from 5 of 89 to 3. No file of the jevgate, fastapi-template or chi corpora is affected: every outline there fits a recheck. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A copy of three short statements is the smallest a clone can be, and in Java it was as often an idiom as a missing helper. On jsoup three of its fifteen shared-logic reviews were such copies, and two of them were not worth sharing: `Attribute::html` and `Attributes::html` borrow a pooled builder, render and release it, as 31 methods of jsoup do; the Cleaner's TextNode and DataNode branches copy one node type each. - A pair whose every site spans three lines or fewer is at most a consider, worded "repeat related steps; a person should decide". - Longer copies keep their level. No shared-logic review of the jevgate, fastapi-template or chi corpora is that short, so none changes. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…nd values On jsoup, both shared-logic units left uncertain paired constructors that only store their fields (`this.lineMap = lineMap; …` in Range and QueryParser, and the copy constructor of HttpConnection.Request with itself), and four of eight hardcoded-value considers were numbers with nothing to name: `new ArrayList<>(4)` and `cost()` methods returning 7. - A Java constructor statement that stores a parameter, another object's field or a literal in a field breaks a clone window, like an excluded line. Constructors that also do work are still compared. - The one number argument of `new …List/Map/Set/Builder/Buffer/…(n)` is an initial capacity, not a value. - A method whose whole body returns one number is named by the method. A returned string stays a candidate: it may be an address. On jsoup, shared-logic units left uncertain went from 2 of 19 to none, and hardcoded-value considers from 8 to 6 (the `cost()` 7 and the capacity 4 are gone). Nothing changes for other languages. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A MockMvc or RestTemplate test calls its controller through a path
(`mockMvc.perform(get("/owners/{ownerId}", 1))`), so it named no code
under test: its recheck had no controller source, and tests of one
endpoint shared no subject, so their overlap was never asked.
- Controller methods carry the routes of their `@GetMapping`,
`@PostMapping`, `@PutMapping`, `@DeleteMapping`, `@PatchMapping` or
`@RequestMapping` annotations (paths, arrays of paths, `method =`),
under the class's `@RequestMapping` prefix.
- A test's `get`/`post`/…, `getForEntity`-style calls with a literal
path are its requests. Each request calls the most literal route that
serves it, by the method's full name (`PetController::processCreationForm`),
since controllers share method names. A route variable matches any
segment; a variable the test fills in matches only a route variable.
- The recheck shows that method's source and its route
(`GET /owners/{ownerId}`).
- Test outlines name a subject by its owning type only in Java, with
owners from Java files: a Go test's `resp.Body.Close()` had been named
after the one type of its package with a `Close` method. The jevgate,
fastapi-template and chi reports are again identical to main's.
On spring-petclinic, PetControllerTests' form-error tests are now
compared and six redundancy considers name parameterizable tests of
`processCreationForm` and `processUpdateForm`. Test-value units left
uncertain went from four to five of 76: with the controller source, Jev
still splits on whether a MockMvc test that checks a view name or JSON
shape and mocked data checks only its mocks, and VetController's JSON
test, cleared before, is now one of them.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
On spring-petclinic, the only shared-logic review among tests paired the
`george()` fixture of OwnerControllerTests with an owner built inside
ClinicServiceTests::shouldInsertOwner: five `owner.setX("…")` lines whose
only differences were the values. That is test data, not a missing helper.
- A Java statement that calls a `set…` method on an object with one
literal breaks a clone window, like a constructor storing its fields.
Setters given computed values (`dto.setCity(owner.getCity().trim())`)
are still compared.
- `check --help` lists Java among the languages discovered by default.
On spring-petclinic the review is gone; no other finding changes. Nothing
changes for other languages.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… body On spring-petclinic, 5 of 76 test-value units stayed uncertain on "checks only its mocks", and on jsoup (helper and safety tests) 4 of 53 redundancy pairs stayed spread over the three overlap levels. Pairs of validator tests were reported as "4 tests of `setBirthDate` overlap". - The mock question gives examples: a test with no stub at all (a setter read back, a round trip, a benchmark), which stub or handler the code chose, the view or status a handler chose for stubbed data, and a result picked from stubbed input check the code. Without them such tests stayed near a third, as if every assertion were about the mocks. - A pair of tests is about a shared function that is neither a camelCase getter or setter nor called by most of the file's tests, the one the fewest tests call, instead of the first shared name in order. A fixture every test uses (`create_user`) is no longer the subject of a pair. - A pair left undecided is asked the overlap again with the body of that function: `con.response()` throws before `.parse()` or `.body()` runs, which only its body shows. A Ruby pair is asked again whether each test checks something the other does not, since the recheck replaces both. Question wording version 9; test value goes to rule version 5 and test redundancy to 2. On spring-petclinic, test-value units left uncertain went from 5 of 76 to 2; on jsoup's helper and safety tests, redundancy pairs left uncertain went from 4 of 53 to 1. Against main on the regression corpora, chi's test-value units left uncertain went from 10 of 131 to 1 and jevgate's from 4 of 237 to 3, with no review or consider finding changed. One jevgate test that scripts Jev's answers and checks the status composed from them rose from 0.19 to 0.3 on the mock question and is now uncertain, and two "several unrelated behaviors" notes moved across 0.80 by 0.01 to 0.03 when their requests were asked again. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…umentation checks
The documentation rules read Markdown only. On pallets/flask (Sphinx
.rst docs) they judged nothing; on vercel/ai (MDX site, AGENTS.md) MDX
was read as raw Markdown: 998 duplication considers, mostly pages that
shared a `streamText` call, 16 of 17 staleness considers were the
reader's own files or a dependency's binary (`tsx`), and 15 files stayed
uncertain.
Formats and harnesses:
- MDX, reStructuredText and AsciiDoc are read as Markdown with the
file's own lines (src/docs/format.rs): MDX drops imports, exports,
comments and component markup but keeps the prose components carry;
reST titles become headings by the order of their adornment styles;
AsciiDoc titles by their `=` level; comments and attributes are
dropped and code blocks fenced with their language. `_build/` is
skipped. Frontmatter `title` names a page.
- Kiro steering files, Junie guidelines and rules, and Roo Code rules
are instruction files, loaded by each harness's own rules.
Evidence:
- Manifests carry the runtime versions they require (`engines`,
`packageManager`, `requires-python`, `rust-version`).
- Staleness drops names written with a code role (`:attr:`), paths the
ignore files cover, and files a section writes out for a code block;
dependencies count as scripts, scripts are read only where a command
starts, and a missing `.ts` names its one `.tsx` namesake.
- Duplication pairs sections on prose, commands and settings, never
program code; package READMEs are not paired with each other, and a
section repeated in many documents is asked against one head and
reported as one finding. Each section is sent with its document title.
Questions:
- "States everything" and "disagrees" are Scores whose middle ("mostly",
"only in detail") is acceptable; a pair about different subjects
settles what stays undecided.
- A pair, an instruction section or a large document still undecided is
asked what it is: how two sections relate, which kind of section, which
kind of document. The kind raises or clears only what is undecided; on
vercel/ai undecided pairs went 18 to 1, instruction sections 6 to 3,
and large documents 3 to 0.
- The linter question asks whether a section is only style the listed
tools check with their usual settings.
On vercel/ai, uncertain units went 43/1175 to 7/789 and uncertain files
15 to 5; of 47 considers, 45 were right by hand. On flask, 54 units were
judged, none undecided; both considers were right.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
On vercel/ai three staleness checks stayed between 0.24 and 0.49: a protocol method (`tools/call`), skill-relative example paths and a migration guide's `pnpm drizzle-kit` for the reader's own app. On fastapi-template all four candidates stayed undecided. - A check still undecided is asked, apart, what the section treats its missing names as: a current part of the repository, the reader's own project, an example, not a file at all, or something removed. The repository ruled out at 0.20 clears it; nothing else moves. All three vercel/ai checks cleared (0.91 "not a file", 0.90 "example", 0.52 "reader" with the repository at 0.14). - A relative link that climbs above the repository, such as a README badge's `../../actions/workflows/ci.yml/badge.svg`, is a route of its host, not a path. - A bare name is also checked as a directory against the ignore files, so `backend/app/frontend/` covers the build output a guide names. - Question wording version 9. vercel/ai: undecided units 7 to 4 of 789 (3 instruction sections torn between instructions and a description, and 1 pair), plus 1 pair over the request limit; uncertain files 5 to 2. fastapi-template: staleness candidates 5 to 0 (gitignored outputs and badge links), uncertain files 9 to 5 against main. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Java, C#, PHP and Ruby are judged, Django and Next.js are understood for security, and MDX, reStructuredText and AsciiDoc docs are read. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Contributor
|
Feel free to reopen this, the close was not intended please check the #2 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
TypeSafe remains the default. This adds opt-in OpenRouter selection through
jevgate.tomlor--provider openrouter, readsOPENROUTER_API_KEY, and uses Jev’s System One endpoint. It defaults totypesafe/jev-latestand keeps Jev’s typed answers.The README and CI setup are updated. Tests cover provider and credential selection, request mapping, and versioned responses.
Validated with
cargo fmt,cargo test --all-targets, andcargo clippy --all-targets -- -D warnings.Prepared with AI assistance and reviewed before submission.