Repository navigation
Keep finding fingerprints stable across merges of main or a parent branch #56
Description
Activity
Spike result (2026-10-03, jevgate 0.33.0 at 3f3ac98)
How a fingerprint is made
sha256(rule \0 owner path \0 identity)(units/compose/mod.rs::fingerprint). Baseline matching is exact set membership on that string (baseline.rs::apply). Line numbers are not part of the hash. What goes intoidentitydepends on the unit:Unit identity What changes it function, hardcoded value, test, security subject, doc section name + whitespace-free source any token edit inside the unit file organization (outline) names of all members any member added, removed or renamed shared-logic pair a.function,b.path,b.function, hash of the normalized window; the owner path is the first selected sitewhich files are selected, which pair represents the group, how far the window extends, renames custom hunk Git's hunk-header function line + the changed lines ( --unified=3)code above the hunk; neighbouring changes that join the same hunk grouped comments, test clusters, grouped plans the set of member identities membership Fixture
A small TypeScript repo (three files, two shared-logic findings, one function-simplification finding, all baselined on main), with a stacked parent→child and main moving underneath. Recipe, outputs and a git bundle are in the ignored
.jevgate/evaluation/issue-56/.Case Result Cause Lines inserted above the units, merged into the branch stable, all answers cached the hash has no positions A third copy added in an unrelated file pair dd2a…→cdb2…, re-raisedthe group's representative pair changed (rank = smallest copy size × number of copies) Only the third copy edited (shorter names) the old dd2a…comes back as NEWthe representative changed again One statement added inside the other copy dd2a…→0114…the matching window moved, so its normalized hash changed Stacked child, --base mainvs--base parentthe parent base doesn't find the pair at all (unchanged files are not loaded); with three copies, --basefinds a different pair than the whole run (NEW vs baselined)the candidate set depends on which files are selected The same pair with --contextholding the other filef6b1…, now reported on the other filethe owner is the selected site, not a canonical one The other copy's file renamed / the owner file renamed new fingerprints for the pair / for every finding in the file the path is hashed Custom hunk: main adds a function above the child's hunk fe0d…→46e4…Git's hunk header now names the new function Custom hunk: the parent edits the next line --base parentgivesfe0d…,--base maingivesf7ca…the two changes join into one -U3hunkCustom hunk: main adds a constant just above the fingerprint stays the same, but the request changed, so the cache missed and the question was asked again the context lines are sent as evidence nodehaven-online's baseline history agrees. 46 findings have more than one fingerprint for the same rule, path and message: 23 function-simplification (every one an edit to the function, since that fingerprint is a pure function of its content), 10 file-organization (member lists changed) and 13 shared-logic or hardcoded-values entries. In the shared-logic ones the second copy and the count of "N more copies" change between fingerprints, as the fixture shows.
A separate effect: a merge also changes the evidence sent with each request: callee signatures from the selected files, packs, the outline's line count, hunk context. The cache then misses, a fresh answer can raise a unit that was clear before, and it comes back as a "new" finding with a fingerprint that was never marked. No fingerprint change fixes this.
Proposal
- Shared logic: a canonical, selection-independent identity. Use the two endpoints
(path, function)sorted, and no window hash and no owner path.one_per_function_pairalready reports one finding per function pair, so nothing collapses that is reported apart today. Copies outside any function keep the normalized hash. A group finding also carries the identities of its member pairs as aliases. - Custom hunk: drop Git's header. Use the enclosing unit that JevGate's own parser finds, plus the changed lines taken from a
-U0diff, so neighbouring changes stay apart. The-U3diff is still sent as evidence. - File organization: use the rule and path only (there is one outline unit per file), or the path with a member-similarity match.
- Content units stay content-identified. An edited function is judged again; that is the intended behaviour.
- Renames: when Git reports a rename, also compute each unit's fingerprint with the old path and use it as an alias.
- Matching by aliases. A finding carries its new fingerprint and aliases (its v1 fingerprint, the pair members, the pre-rename path).
applyaccepts a finding when any of these is in the baseline.baseline writeandmarkwrite the new fingerprint and keep the reason found through an alias.
Risks
- False matches. A function pair that was accepted stays accepted while the copies diverge or grow. With "any member pair accepted", a new copy that joins an accepted group would be hidden (the safer rule is to raise the group again when a copy that has not been accepted joins it). An outline accepted by path stays accepted as the file grows. Two identical changed lines in one function need the existing occurrence suffix.
- Migration. Old entries keep matching through their v1 alias, so no rewrite is needed to keep gates green. An optional offline
jevgate baseline migratecan plan the units, with no inference, and rewrite each entry whose v1 matches a current unit. The one exception is shared-logic v1 entries written under a different selection: those match nothing in a whole-repo plan, so they stay until they go stale. SARIF should emit bothjevgateFingerprint/v1and/v2for one release so code-scanning alerts don't all reopen. GitLab's fingerprint, MCP ids, hook turn memory andchanges.rslineage need the same alias treatment, or they show one round of resolved and introduced entries.
Alternatives
- Fuzzy matching on rule, path and symbol, with a similarity threshold. This is less churn but needs the symbol in the baseline schema and has more false matches.
- In
--base, load every file for clone detection, which is local and costs no tokens, so groups match a whole run. This improves recall (copies in unchanged files) but changes behaviour and what is uploaded. - Process only: always check stacked PRs against the parent. This doesn't fix the representative choice, renames or hunk headers.
Decisions needed before implementation
- The shared-logic identity: a canonical function pair (recommended), or keep the window hash and only canonicalize the owner. For a group: match when any pair is accepted, or raise the group again when an unaccepted copy joins it.
- Whether
--baseshould find copies in unchanged files (more recall, more upload). - The outline identity: path only, or path with member similarity.
- Content units: keep strict content identity (recommended), or match a symbol fuzzily.
- Migration: v1 aliases plus an optional offline
baseline migrate(recommended), or a one-time forced rewrite; dual SARIF fingerprints for one release. - Follow renames through aliases: yes or no.
- Shared logic: a canonical, selection-independent identity. Use the two endpoints
Decisions (2026-10-03), after the spike
- Shared logic: identify a finding by its function pair,
(path, function)sorted, with no statement-span hash and no owner path. A group stays accepted, but it is raised again when an unaccepted copy joins it, so new duplication is never hidden. Copies outside a function keep the hash. - Outline (file organization): identify by path plus a similar member list. A mostly unchanged file matches; a file that grew a lot is asked again.
- Migration: a forced one-time rewrite. The next baseline write rewrites every entry to the new fingerprints; to do that, entries are matched through their v1 fingerprints during the rewrite. Expect one large baseline diff in each repository.
- Renames: findings follow Git-reported renames and match through the old path.
- Unchanged as proposed: function and other code units keep strict content identity. Custom hunks drop Git's header and use JevGate's own enclosing function, plus the changed lines from a diff with no context lines.
--basedoes not search unchanged files (consistent with Flag cross-file shared-logic findings so a change fixes only its own copy #55).
Implementation follows #55, because both change how shared-logic findings are owned.
- Shared logic: identify a finding by its function pair,
Decision (2026-10-03): the transition
After the upgrade, an entry still on a v1 fingerprint keeps matching through the finding's v1 fingerprint until a baseline write (
mark,baseline,--merge) rewrites it to v2, keeping its reason and note. Gates therefore stay green on upgrade, and one whole check plus a write produces the one-time diff. The user rejected the strict alternative, in which v1 entries stop matching at once.Builder defaults the coordinator accepts: an outline matches when the Jaccard similarity of member names is at least 0.8; shared-logic and outline entries store
members; the baseline file keepsversion: 1.Decisions (user, 2026-10-04)
- SARIF v1 fingerprint: drop
jevgateFingerprint/v1from SARIFpartialFingerprintsin the release after the one that ships v2 fingerprints (Keep fingerprints stable across merges, selections and renames #59). That release keeps both and notes v1 as deprecated.
- SARIF v1 fingerprint: drop
Goal
After main or a parent branch is merged into a branch, findings whose code did not change should keep their fingerprints, so their marks still match and are not raised again. Stacked PRs in nodehaven-online (#181 → #180 → #178) re-raised and re-churned marks after each merge.
Scope
Acceptance (spike)
Decisions
2026-10-03
Requested by the user. It starts as a spike.