Skip to content

LLM notes: eval harness — golden fixtures, cached-response replay, prompt snapshots #542

Description

@goosewobbler

From the July 2026 direction review (LLM-notes feature audit).

Carved out: the renderExamplesBlock example-escaping item originally listed under Scope shipped separately as #571 (prompt-injection hardening in the #540/#565 family — it didn't need the eval harness to land). This issue now tracks the eval harness itself.

Gap

LLM-enhanced notes are releasekit's clearest differentiator, but nothing measures whether output is good — only structurally valid. Prompt drift is invisible in review diffs (no prompt is snapshot-tested; prompts.spec.ts only covers override-append), and a wording tweak silently busts the entire response cache with no visible signal.

Scope

  • test/eval/ harness in packages/notes:
    • Golden fixtures: real commit sets (e.g. from this repo's history) as inputs.
    • Record/replay: recorded provider responses replayed via the existing withContentHashCache on-disk format — the cache module already is a record/replay mechanism; deterministic in CI, no keys needed.
    • Deterministic assertions: past tense, length bounds, category distribution sanity, no <!-- releasekit- marker strings, no duplicated "Updated dependencies" churn.
    • Opt-in live mode (RELEASEKIT_EVAL=1): runs against Ollama or a cheap model; optional LLM-judge scoring.
  • Prompt snapshot tests for the four task system prompts (enhance, categorize, enhanceAndCategorize, releaseNotes) so drift is a visible, reviewed event.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions