Skip to content

Keep finding fingerprints stable across merges of main or a parent branch #56

Description

@tauanbinato

Goal

After main or a parent branch is merged into a branch, findings whose code did not change should keep their fingerprints, so their marks still match and are not raised again. Stacked PRs in nodehaven-online (#181 → #180 → #178) re-raised and re-churned marks after each merge.

Scope

  1. A research spike: reproduce on a small stacked fixture, then find which evidence or unit inputs change on a merge (line positions, hunk context, candidate sets, the base used) and why.
  2. Propose a fix with its trade-offs (for example, fingerprints from unit content rather than position) before implementing it.

Acceptance (spike)

  • A written finding: the cause, the fixture, the proposed fix, and the risk to existing baselines (migration).
  • Implementation waits for the user's decision on the proposal.

Decisions

2026-10-03

Requested by the user. It starts as a spike.

Activity

  1. tauanbinato commented on Oct 4, 2026

    @tauanbinato
    ContributorAuthor

    Spike result (2026-10-03, jevgate 0.33.0 at 3f3ac98)

    How a fingerprint is made

    sha256(rule \0 owner path \0 identity) (units/compose/mod.rs::fingerprint). Baseline matching is exact set membership on that string (baseline.rs::apply). Line numbers are not part of the hash. What goes into identity depends on the unit:

    Unit identity What changes it
    function, hardcoded value, test, security subject, doc section name + whitespace-free source any token edit inside the unit
    file organization (outline) names of all members any member added, removed or renamed
    shared-logic pair a.function, b.path, b.function, hash of the normalized window; the owner path is the first selected site which files are selected, which pair represents the group, how far the window extends, renames
    custom hunk Git's hunk-header function line + the changed lines (--unified=3) code above the hunk; neighbouring changes that join the same hunk
    grouped comments, test clusters, grouped plans the set of member identities membership

    Fixture

    A small TypeScript repo (three files, two shared-logic findings, one function-simplification finding, all baselined on main), with a stacked parent→child and main moving underneath. Recipe, outputs and a git bundle are in the ignored .jevgate/evaluation/issue-56/.

    Case Result Cause
    Lines inserted above the units, merged into the branch stable, all answers cached the hash has no positions
    A third copy added in an unrelated file pair dd2a… → cdb2…, re-raised the group's representative pair changed (rank = smallest copy size × number of copies)
    Only the third copy edited (shorter names) the old dd2a… comes back as NEW the representative changed again
    One statement added inside the other copy dd2a… → 0114… the matching window moved, so its normalized hash changed
    Stacked child, --base main vs --base parent the parent base doesn't find the pair at all (unchanged files are not loaded); with three copies, --base finds a different pair than the whole run (NEW vs baselined) the candidate set depends on which files are selected
    The same pair with --context holding the other file f6b1…, now reported on the other file the owner is the selected site, not a canonical one
    The other copy's file renamed / the owner file renamed new fingerprints for the pair / for every finding in the file the path is hashed
    Custom hunk: main adds a function above the child's hunk fe0d… → 46e4… Git's hunk header now names the new function
    Custom hunk: the parent edits the next line --base parent gives fe0d…, --base main gives f7ca… the two changes join into one -U3 hunk
    Custom hunk: main adds a constant just above the fingerprint stays the same, but the request changed, so the cache missed and the question was asked again the context lines are sent as evidence

    nodehaven-online's baseline history agrees. 46 findings have more than one fingerprint for the same rule, path and message: 23 function-simplification (every one an edit to the function, since that fingerprint is a pure function of its content), 10 file-organization (member lists changed) and 13 shared-logic or hardcoded-values entries. In the shared-logic ones the second copy and the count of "N more copies" change between fingerprints, as the fixture shows.

    A separate effect: a merge also changes the evidence sent with each request: callee signatures from the selected files, packs, the outline's line count, hunk context. The cache then misses, a fresh answer can raise a unit that was clear before, and it comes back as a "new" finding with a fingerprint that was never marked. No fingerprint change fixes this.

    Proposal

    1. Shared logic: a canonical, selection-independent identity. Use the two endpoints (path, function) sorted, and no window hash and no owner path. one_per_function_pair already reports one finding per function pair, so nothing collapses that is reported apart today. Copies outside any function keep the normalized hash. A group finding also carries the identities of its member pairs as aliases.
    2. Custom hunk: drop Git's header. Use the enclosing unit that JevGate's own parser finds, plus the changed lines taken from a -U0 diff, so neighbouring changes stay apart. The -U3 diff is still sent as evidence.
    3. File organization: use the rule and path only (there is one outline unit per file), or the path with a member-similarity match.
    4. Content units stay content-identified. An edited function is judged again; that is the intended behaviour.
    5. Renames: when Git reports a rename, also compute each unit's fingerprint with the old path and use it as an alias.
    6. Matching by aliases. A finding carries its new fingerprint and aliases (its v1 fingerprint, the pair members, the pre-rename path). apply accepts a finding when any of these is in the baseline. baseline write and mark write the new fingerprint and keep the reason found through an alias.

    Risks

    • False matches. A function pair that was accepted stays accepted while the copies diverge or grow. With "any member pair accepted", a new copy that joins an accepted group would be hidden (the safer rule is to raise the group again when a copy that has not been accepted joins it). An outline accepted by path stays accepted as the file grows. Two identical changed lines in one function need the existing occurrence suffix.
    • Migration. Old entries keep matching through their v1 alias, so no rewrite is needed to keep gates green. An optional offline jevgate baseline migrate can plan the units, with no inference, and rewrite each entry whose v1 matches a current unit. The one exception is shared-logic v1 entries written under a different selection: those match nothing in a whole-repo plan, so they stay until they go stale. SARIF should emit both jevgateFingerprint/v1 and /v2 for one release so code-scanning alerts don't all reopen. GitLab's fingerprint, MCP ids, hook turn memory and changes.rs lineage need the same alias treatment, or they show one round of resolved and introduced entries.

    Alternatives

    • Fuzzy matching on rule, path and symbol, with a similarity threshold. This is less churn but needs the symbol in the baseline schema and has more false matches.
    • In --base, load every file for clone detection, which is local and costs no tokens, so groups match a whole run. This improves recall (copies in unchanged files) but changes behaviour and what is uploaded.
    • Process only: always check stacked PRs against the parent. This doesn't fix the representative choice, renames or hunk headers.

    Decisions needed before implementation

    1. The shared-logic identity: a canonical function pair (recommended), or keep the window hash and only canonicalize the owner. For a group: match when any pair is accepted, or raise the group again when an unaccepted copy joins it.
    2. Whether --base should find copies in unchanged files (more recall, more upload).
    3. The outline identity: path only, or path with member similarity.
    4. Content units: keep strict content identity (recommended), or match a symbol fuzzily.
    5. Migration: v1 aliases plus an optional offline baseline migrate (recommended), or a one-time forced rewrite; dual SARIF fingerprints for one release.
    6. Follow renames through aliases: yes or no.
  2. tauanbinato commented on Oct 4, 2026

    @tauanbinato
    ContributorAuthor

    Decisions (2026-10-03), after the spike

    • Shared logic: identify a finding by its function pair, (path, function) sorted, with no statement-span hash and no owner path. A group stays accepted, but it is raised again when an unaccepted copy joins it, so new duplication is never hidden. Copies outside a function keep the hash.
    • Outline (file organization): identify by path plus a similar member list. A mostly unchanged file matches; a file that grew a lot is asked again.
    • Migration: a forced one-time rewrite. The next baseline write rewrites every entry to the new fingerprints; to do that, entries are matched through their v1 fingerprints during the rewrite. Expect one large baseline diff in each repository.
    • Renames: findings follow Git-reported renames and match through the old path.
    • Unchanged as proposed: function and other code units keep strict content identity. Custom hunks drop Git's header and use JevGate's own enclosing function, plus the changed lines from a diff with no context lines. --base does not search unchanged files (consistent with Flag cross-file shared-logic findings so a change fixes only its own copy #55).

    Implementation follows #55, because both change how shared-logic findings are owned.

  3. tauanbinato commented on Oct 4, 2026

    @tauanbinato
    ContributorAuthor

    Decision (2026-10-03): the transition

    After the upgrade, an entry still on a v1 fingerprint keeps matching through the finding's v1 fingerprint until a baseline write (mark, baseline, --merge) rewrites it to v2, keeping its reason and note. Gates therefore stay green on upgrade, and one whole check plus a write produces the one-time diff. The user rejected the strict alternative, in which v1 entries stop matching at once.

    Builder defaults the coordinator accepts: an outline matches when the Jaccard similarity of member names is at least 0.8; shared-logic and outline entries store members; the baseline file keeps version: 1.

  4. tauanbinato commented on Oct 4, 2026

    @tauanbinato
    ContributorAuthor

    Decisions (user, 2026-10-04)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions