Skip to content

EPIC: SBOM Merge strategies, including de-duplication #451

Description

@jimklimov

For some historic context: some in the community may remember that I've started with this some 3 years ago, ended up with code that "worked for us" internally but for various reasons its PRs were not merged by the upstream, and just used the resulting build in-house as it was slowly getting obsolete with new CycloneDX spec revisions coming out while we were stuck with support for v1.4 max which was less and less useful as new maven, dotnet, container-baseline and other generators were issuing newer-spec documents by default and we had to fiddle with their settings or write scripts to shoehorn those incoming documents into something that this aging build of the toolkit could parse to merge.

Recently (with help from Claude AI) I revisited that old solution to rewrite it in a (more) correct C# and rebase on top of current trunks of the cyclonedx-dotnet-library for most of the code, and cyclonedx-cli to easily call the new optional abilities. I hope this new iteration is less questionable regarding merge-ability so everyone can benefit :)

Since the core of the changes are in the library, I'm posting this "epic" issue here to tie the numerous old and new tickets together.

Problem overview

For various reasons, including anticipated customer demand and actual legal compliance, as well as for internal benefit for our developers to plug security holes in a timely manner (especially in old but supported releases of the large product my dayjob is making), we wanted to produce a "Release Bundle SBOM" -- single document which lists all components that end up in the customer's environment when some or all components of our product get installed there. One aspect to this is honestly listing the FOSS components and their license/author attributions, another is knowing their versions deployed (eventually the potential vulnerabilities to investigate), yet another is having a dependency graph so we can do this for any subset of our product.

To clarify, an installation of our product delivers roughly 20 to 50 separate programs (each with a dependency tree into Java, DotNet, NodeJS and other ecosystems). Some editions are used as plain programs, others can be delivered as containers for easier runs in arbitrary environments - but so add dependency trees from the base OS images involved. While each component's dependency tree is strictly-speaking unique, they largely overlap per ecosystem, especially starting a couple of levels below our import statements (e.g. "using Spring" may mean certain variability in our Java pom.xml files to pick the components our program uses, but then it is the Spring sources' curated choice of what gets pulled in - with little control from our side). Each incoming document is a complete SBOM of the dependency or product it describes - not just that one component entry, but also its dependency tree (and so the other components which this tree mentions).

At the scale of our project, a completely honest merge - which is this toolkit's default - deals with about 30000 components and their unique 1:1 dependency links (with a LOT of repeating names). When I began with this 3-4 years ago, it was not what our team expected, to the point of considering the retention of all duplicates a bug (what sort of a "merge" is that?), but during the review discussions I've learned it is done for quite a valid reason - the same potentially-vulnerable component can be a security issue when used in one codepath, and not a problem in another (e.g. exposed to network or not, or the component has many features that are enabled in some configurations and not called in others). So the vulnerability investigations, such as verdicts maintained in DependencyTrack server, should take each such link as a separate use-case that may be or not be vulnerable independently.

This is nice in theory, but we are not gonna get anyone looking at, and clicking through in DT, the 30k use-cases and ultimately verdicts - neither several times a week as the team iterates the codebase and issues patches to older releases, nor even once in a quarter of a year for new initial public releases. Just the sheer amount of this work would mean that we have to pull our most experienced and expensive team members, who know the project and various implications for different customers' unique deployments like a chess gross-meister playing a dozen games at once from memory, into man-weeks worth of this review instead of actual work on, you know, developing the product.

  • This stance might change now that the chore can be fed to AI, but even that can turn out prohibitively expensive

Given that the majority of our dependency links are outside our control anyway (the dependency webs of frameworks and ecosystems we build upon), the reasonable trade-off is to have a de-duplicated merged document, where each instance of a component is named once, and has multiple pointers (graph links) from whoever touches it. In our case, the merged document with 30k incoming components ends up with 700-1000 components, and a much smaller amount of vulnerability instances to investigate when new relevant CVEs come into our DT server.

Looking at the issue tracker here, using a "lossy" merge strategy is a reasonable trade-off for many other askers too :)

While my effort started with a single goal way back when, I found that actually "reasonable expectations" do differ from project to project, so I tried to introduce a concept of merge strategies which impact decisions in otherwise similar code structure, so the mostly-same iterations do not have to be written again and again once somebody has a bright idea of doing something differently.

Below I will list preceding tickets from the currently proposed solution, from the older iteration, and mention other issues that asked about a similar feature or voiced concerns about the original full or flat merge implementations (if only to have GitHub references from them, so original askers are notified that their wishes may have been heard).

Problems solved

As I followed this windy road, several problems were identified and addressed:

  • "One of the hardest problems in IT is naming": the opaque bom-ref identifier for some component (and inverse ref from dependency declarations) in its original document is usually the same string as its purl. When you add and keep dozens of copies of essentially the same component during such merges of related dependency trees, each copy (and corresponding back-references) should use an unique name.
  • Data completeness: for many same-named components we have different information from different incoming documents. For example, analysis of the source recipe could give us its name, dependencies and licensing information, maybe even a reference SCM URL, but a separate analysis of an actual binary/library we use gives us the expected file hashes, or the SCM commit identifier (hash, tag) that build was started from. We do have different incoming SBOMs telling us different things about the same entities, and in many cases they can be safely accumulated into one larger entity.
  • There are cases where it is not safe to fully merge different mentions of a component into one entity. The most frequent for us is a conflict around the scope values in different copies of dependency branches (e.g. used in tests=>excluded from the deliverable vs. used in production=>required). This led to introduction of different strategy options: a simpler merge conflates a few meanings to keep exactly one mention of a component (e.g. "optional+required" = "required") but aborts production of a document in case of conflicts that can not be reconciled ("required+excluded = ?"), and a recently completed optional strategy allows to rename such components so their bom-ref adds a scope=... suffix for each different copy and merges the other information into both, grouping the back-references reasonably (all excluded go here, all required go there).
  • "Dangling links": as a result of such rewrites and merges, with some strategies it is possible (potentially or practically) that components are added into the final document but are not mentioned in any dependency declarations. This can also happen due to incoming data being incomplete like this. A special option was added to identify such components and synthesize dependency sink(s) for them, so that all listed components are in the graph (in part, this was required for DependencyTrack server to handle such documents properly) and those not in a "real" dependency tree are easy to identify (e.g. during merge-strategy development and debugging).
  • Command-line size limits: the mundane issue of listing the hundreds of incoming documents to be merged - if posted as --input-file NAME they exceed the OS-dictated CLI limits. An option was added to pass an --input-file-list for this (additionally/instead).
  • Various troubles with the "Metadata" component - I have actually to refresh my memory on this one to write about it responsibly :)

Ready to go for a spin?

For smaller review scopes, the project changes were split into multiple pull requests listed further below. Here are my recent development branches which include all of these changes, and maybe something on top (like internal README notes about getting the custom-versioned builds to play together) so brave souls can just try it out like I do:

New proposal:

Old proposal

Most PRs are closed as superseded by those above:

Related issues in the tracker

Listing to show that not only our team saw this as a problem, and for GitHub to add backlinks in those tickets to this epic as a possible solution :)

Honourable mentions

Kudos to everyone who participated in discussions about the initial development, whether in GitHub reviews or in the Slack and other channels. After 3 years I probably would not tag everyone who helped move this forward, but at least gotta tip the hat to those I can: @nscuro @andreas-hilti @jkowalleck @coderpatros @fnxpt @stevespringett @pombredanne ...

Special thanks to @claude for making sense from my original ideas and reflection-heavy implementation into (more) proper code :)

Also thanks to my dayjob for putting up with the man-weeks or even months sunk into this work, back in the day and recently, and yet letting it be out in the open.

Activity

  1. jimklimov commented on Sep 17, 2026

    @jimklimov
    ContributorAuthor

    Follow-up questions, posted to AI, as these PRs focused on producing a functional equivalent to what I had done 3 years ago:

    • Do the merge strategy and general code base changes introduced by these PRs cover any new entity types introduced by the upstream in the past 3 years?
    • Will any subsequently introduced classes in future versions need special (boilerplate?) code to support merge strategies for object properties and lists of objects, or will they "just work" by default simply inheriting from a base class/interface?
    • When we merge lists of items, do we now just add non-equal items into the list, or for some types analyze if some existing list entries can be merged from A to B?

    Good questions — I went and read the actual interface family, the list-merge algorithm, and cross-checked against the current Bom model to answer precisely rather than from memory.

    1. Coverage of new entity types (last ~3 years)

    Partially covered, and there's a real gap. The IMergeable<T>/IEquivalent<T> retarget touched 18 classes, and that set does include the two newest additions that got wired into merge at all: CycloneDX 1.6's Declarations (Assessor, Attestation, Claim) and Definitions (Standard) — both interface-enabled and dispatched in the new FlatMerge.

    But two other additions are not covered at all, silently:

    • Bom.Formulation (List<Formula>, CycloneDX 1.5 workflows) — Formula implements no equality/merge interface.
    • Bom.Citations (List<Citation>, CycloneDX 1.7 — this year's addition) — same, no interfaces.

    Neither list is even copied into the merged output — not by my new code, and not by the original pre-existing plain FlatMerge(Bom, Bom) either (I checked both side by side). So this isn't a regression I introduced; it's a pre-existing blind spot in the library that neither the old nor new merge code closes. Worth a follow-up issue, but out of scope of what I touched. Also worth noting: Bom.Properties (the top-level one, distinct from Component.Properties) is likewise never populated by FlatMerge even though the Property class itself is fully merge-capable — it's just never referenced at that particular call site.

    2. Will future classes "just work"?

    Not automatically — two things are still manual, by design (the interfaces deliberately avoid reflection/a shared base class):

    • Opting a class in is genuinely trivial if it already implements IEquatable<T> (which most model classes here do): add , IMergeable<T>, IEquivalent<T> to the class declaration, zero method bodies. The default interface methods give you "merge only if already exactly equal, otherwise keep both" for free.
    • Wiring it into an actual merge is not automatic. Declaring the interfaces on a class doesn't make anything call it — some call site has to explicitly invoke MergeableListHelper.Merge(...) on that specific list property (either in Merge.cs's FlatMerge, or inside a container type's own MergeWith, e.g. Component.cs for Hashes/Properties/nested Components). This is exactly why Formulation/Citations fall through the cracks above — even if Formula/Citation grew the interfaces tomorrow, nothing would call MergeableListHelper.Merge on them until someone also adds that line to Merge.cs. So a genuinely new spec field needs both steps, not just inheritance.
    • Real reconciliation logic (beyond "merge if identical") is always custom code — Equivalent()/MergeWith() overrides, unavoidable since it's field-specific business logic.

    3. Do list merges reconcile, or just append?

    They reconcile, by default. MergeableListHelper.Merge<T> (used whenever strategy.UseEntityMerge, which MergeStrategy.Default() sets true) scans the existing merged list for each incoming item; if an entry .Equals() or .Equivalent() it, it calls .MergeWith() to fold the incoming data into the existing entry in place, and only appends as a new entry when nothing matched.

    Right now only Component, Hash, and Dependency have non-trivial fold logic (scope-aware squash, fill-in-missing-hash-content, union of dependsOn sub-lists respectively); the other dispatched classes (Tool, Service, Annotation, Property, etc.) fall back to the interface default, which is just "merge if byte-identical, else keep both" — functionally the same as the old exact-match-only behavior for those, just routed through the new mechanism so they're ready for a custom override later without touching call sites again.

    One loose end I noticed while checking this: the doc comment above result.Dependencies = MergeableListHelper.Merge(...) in Merge.cs still says "Dependency reconciliation beyond exact-match... is not yet implemented" — that's stale, left over from before Dependency.MergeWith got its real subset-union logic in a later commit. (Updating this one now)

  2. jkowalleck commented on Sep 17, 2026

    @jkowalleck
    Member
  3. locked and limited conversation to collaborators on Sep 17, 2026
  4. converted this issue into a discussion #452 on Sep 17, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions