From the July 2026 direction review (LLM-notes feature audit).
Carved out: the renderExamplesBlock example-escaping item originally listed under Scope shipped separately as #571 (prompt-injection hardening in the #540/#565 family — it didn't need the eval harness to land). This issue now tracks the eval harness itself.
Gap
LLM-enhanced notes are releasekit's clearest differentiator, but nothing measures whether output is good — only structurally valid. Prompt drift is invisible in review diffs (no prompt is snapshot-tested; prompts.spec.ts only covers override-append), and a wording tweak silently busts the entire response cache with no visible signal.
Scope
test/eval/ harness in packages/notes:
- Golden fixtures: real commit sets (e.g. from this repo's history) as inputs.
- Record/replay: recorded provider responses replayed via the existing
withContentHashCache on-disk format — the cache module already is a record/replay mechanism; deterministic in CI, no keys needed.
- Deterministic assertions: past tense, length bounds, category distribution sanity, no
<!-- releasekit- marker strings, no duplicated "Updated dependencies" churn.
- Opt-in live mode (
RELEASEKIT_EVAL=1): runs against Ollama or a cheap model; optional LLM-judge scoring.
- Prompt snapshot tests for the four task system prompts (
enhance, categorize, enhanceAndCategorize, releaseNotes) so drift is a visible, reviewed event.
From the July 2026 direction review (LLM-notes feature audit).
Gap
LLM-enhanced notes are releasekit's clearest differentiator, but nothing measures whether output is good — only structurally valid. Prompt drift is invisible in review diffs (no prompt is snapshot-tested;
prompts.spec.tsonly covers override-append), and a wording tweak silently busts the entire response cache with no visible signal.Scope
test/eval/harness inpackages/notes:withContentHashCacheon-disk format — the cache module already is a record/replay mechanism; deterministic in CI, no keys needed.<!-- releasekit-marker strings, no duplicated "Updated dependencies" churn.RELEASEKIT_EVAL=1): runs against Ollama or a cheap model; optional LLM-judge scoring.enhance,categorize,enhanceAndCategorize,releaseNotes) so drift is a visible, reviewed event.