Add Shipwright results - #56
Closed
kpuru88 wants to merge 1 commit into
Closed
Conversation
Author
|
@zverianskii @ashleyzhang01 — the 50 raw Shipwright reviews are ready and both CI suites pass. Could you please run the normal Martian dedup/judge pipeline and report the official offline rank? The PR deliberately excludes our local judged outputs so the assessment remains independent. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds Shipwright to the 50-PR offline benchmark for Martian's independent judging and ranking.
What's included
shipwrightreview entry for each of the 50 benchmark tasksreview.jsonoutputsoffline/README.mdNo judged outputs, TP/FP/FN labels, dedup groups, or golden-derived data are added. Existing tool data and golden comments are unchanged.
How the reviews were produced
The reviews were run on Modal through Shipwright's production review-agent path:
full-review-json4701de31de6b87b020429f35406426bf7ee72898a06717005de341c40c423433staged-gepa-round2b-20260822T185655ZEvery submitted review maps to exactly one benchmark task. Candidate text, path, and line are the official candidate conversion of that task's
review.jsonfindings. The prompt was generic and passed the harness's benchmark-leak checks.Shipwright is an offline review agent rather than a GitHub App, so
pr_urlrecords the public source task reviewed andrepo_namerecords the Shipwright run identity.Implementation and provenance work is reviewed in https://github.com/kpuru88/shipwright-agent/pull/20.
Preliminary result—not the requested official score
Our separate Martian/Kimi run over the 173-finding
allprofile produced TP=83, FP=60, FN=90, errors=0. That is included here only for provenance; it is not submitted as an official benchmark score. Please run Martian's normal dedup/judge pipeline and rank this submission with the public benchmark configuration.Validation