Skip to content

Publish a recurring cross-model Jury Report #46

Description

@cderinbogaz

Parent: #39
Depends on: #42

Context

Juror can generate research that individual-model reviewers cannot: marginal recall by juror, agreement patterns, model-family blind spots, and confirmed defects per dollar. A reproducible recurring report can become an enduring discovery channel rather than a one-time launch post.

Acceptance criteria

  • Define a repeatable report format backed by versioned benchmark data and raw artifacts.
  • Report defect-class performance by model family, marginal recall from each additional juror, consensus versus all-findings precision, duplicate rate, latency, and confirmed defects per dollar.
  • Where authorship is known, analyze same-family versus cross-family review while clearly documenting confounders.
  • Include model, harness, prompt/configuration, corpus, evaluator, and pricing versions.
  • Publish negative and inconclusive findings alongside favorable results.
  • Choose a sustainable cadence after the first report; do not imply live accuracy for stale model versions.
  • Produce shareable charts/tables and a machine-readable result artifact.

Metadata

Metadata

Assignees

Labels

benchmarkCorpus, adjudication, evaluationdocumentationImprovements or additions to documentation

Type

No type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions