-
Notifications
You must be signed in to change notification settings - Fork 8
Build an open SSDLC assessment benchmark #19
Copy link
Copy link
Open
Labels
area: standardsFramework packs, OSCAL, and conformanceFramework packs, OSCAL, and conformanceenhancementNew feature or requestNew feature or requesthelp wantedExtra attention is neededExtra attention is neededmaintainer-ledSecurity-critical or architecture-heavy work led by maintainersSecurity-critical or architecture-heavy work led by maintainerspriority: p0Critical path for the next milestoneCritical path for the next milestone
Description
Activity
Metadata
Metadata
Assignees
Labels
area: standardsFramework packs, OSCAL, and conformanceFramework packs, OSCAL, and conformanceenhancementNew feature or requestNew feature or requesthelp wantedExtra attention is neededExtra attention is neededmaintainer-ledSecurity-critical or architecture-heavy work led by maintainersSecurity-critical or architecture-heavy work led by maintainerspriority: p0Critical path for the next milestoneCritical path for the next milestone
Why this matters
DocSentinel can only be trusted if its security review output can be measured against repeatable ground truth. Today, SSDLC document review is hard to compare across LLM providers, prompts, retrieval settings, and policy packs because there is no shared benchmark.
From first principles: if the project cannot tell whether a change improves recall, precision, citation quality, and false-positive rate, the community cannot safely optimize the system.
Community help wanted
We need a public, synthetic, license-clean benchmark for SSDLC assessments across requirements, design, development, testing, deployment, and operations.
Suggested scope
Acceptance criteria
Progress (2026-08-18)
scorecard.jsonand human-readablescorecard.mdoutput.