Skip to content

Add evaluated visual verification routing and model harness#7

Draft
theclaymethod wants to merge 3 commits into
mainfrom
agent/visual-model-eval-harness
Draft

Add evaluated visual verification routing and model harness#7
theclaymethod wants to merge 3 commits into
mainfrom
agent/visual-model-eval-harness

Conversation

@theclaymethod

@theclaymethod theclaymethod commented Jul 25, 2026

Copy link
Copy Markdown
Owner

What changed

  • preserve the committed Artifacture, Impeccable, and Unslop routing boundaries
  • add provider-independent rendering, adapter, runner, aggregation, merge, telemetry, and policy-selection contracts
  • replace the repetitive 112-state fixture set with a curated 36-pair / 72-state seed across five Artifacture-owned families
  • add responsive viewports, visually distinct operating-model cases, and source-conditioned controls for claims whose validity depends on supplied evidence
  • bind evidence identity to rendered pixels plus visible/source evidence while independently gating distinct-image coverage
  • enforce smallest-first model ladders and corpus-declared batch sizes 1 and 2
  • fail closed unless live graduation has human review, representative Artifacture capture provenance, and 10 distinct images and source artifacts per label/family

Why

Artifacture needs empirical evidence before assigning visual verification to a smaller model or increasing screenshot batch size. The corpus is now useful for seed development and contract testing without pretending synthetic fixtures are enough to qualify production routing.

Explicit limitations

  • no provider measurements or model winner are claimed
  • this seed can compare batch 1 with batch 2 only; it does not establish that 2 is globally maximal
  • graduation still requires representative product captures, human adjudication, exact provider candidates, and complete cost/latency/cache telemetry

Validation

  • npm test — 89/89 tests passed
  • visual corpus render — 72 states
  • dry-run request matrix — 324 requests across batch 1 and 2
  • npm run ve:eval — 180 seeded violations, 7 clean fixtures, and 9 design-system cases passed
  • npm run ve:check
  • npm run check:manifests
  • follow-up corpus, specification, and standards nemesis reviews found no remaining material defects in this overhaul

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant