Skip to content

Add Shipwright results - #56

Closed
kpuru88 wants to merge 1 commit into
withmartian:mainfrom
kpuru88:add-shipwright-results
Closed

Add Shipwright results#56
kpuru88 wants to merge 1 commit into
withmartian:mainfrom
kpuru88:add-shipwright-results

Conversation

@kpuru88

@kpuru88 kpuru88 commented Aug 22, 2026

Copy link
Copy Markdown

Adds Shipwright to the 50-PR offline benchmark for Martian's independent judging and ranking.

What's included

  • one shipwright review entry for each of the 50 benchmark tasks
  • 142 findings copied directly from Shipwright review.json outputs
  • one task with an empty finding set, preserved as produced
  • one evaluated-tools row in offline/README.md

No judged outputs, TP/FP/FN labels, dedup groups, or golden-derived data are added. Existing tool data and golden comments are unchanged.

How the reviews were produced

The reviews were run on Modal through Shipwright's production review-agent path:

  • reviewer backend: Claude Agent SDK
  • reviewer model: Sonnet
  • tools: none
  • execution mode: full-review-json
  • diff context limit: 100,000 characters
  • snippet context limit: 0
  • findings cap: 8 per PR
  • prompt SHA-256: 4701de31de6b87b020429f35406426bf7ee72898a06717005de341c40c423433
  • run ID: staged-gepa-round2b-20260822T185655Z

Every submitted review maps to exactly one benchmark task. Candidate text, path, and line are the official candidate conversion of that task's review.json findings. The prompt was generic and passed the harness's benchmark-leak checks.

Shipwright is an offline review agent rather than a GitHub App, so pr_url records the public source task reviewed and repo_name records the Shipwright run identity.

Implementation and provenance work is reviewed in https://github.com/kpuru88/shipwright-agent/pull/20.

Preliminary result—not the requested official score

Our separate Martian/Kimi run over the 173-finding all profile produced TP=83, FP=60, FN=90, errors=0. That is included here only for provenance; it is not submitted as an official benchmark score. Please run Martian's normal dedup/judge pipeline and rank this submission with the public benchmark configuration.

Validation

  • exactly 50 Shipwright review entries
  • exactly 50 unique task URLs
  • 142 submitted findings
  • upstream offline tests: 28 passed
  • additive data diff only

@kpuru88

kpuru88 commented Aug 22, 2026

Copy link
Copy Markdown
Author

@zverianskii @ashleyzhang01 — the 50 raw Shipwright reviews are ready and both CI suites pass. Could you please run the normal Martian dedup/judge pipeline and report the official offline rank? The PR deliberately excludes our local judged outputs so the assessment remains independent.

@kpuru88 kpuru88 closed this Aug 22, 2026
@kpuru88
kpuru88 deleted the add-shipwright-results branch August 22, 2026 20:40
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants