Skip to content

Correct stale annotations; write down the bar we were enforcing - #67

Open
xdotli wants to merge 1 commit into
mainfrom
fix/annotation-accuracy
Open

Correct stale annotations; write down the bar we were enforcing#67
xdotli wants to merge 1 commit into
mainfrom
fix/annotation-accuracy

Conversation

@xdotli

@xdotli xdotli commented Jul 30, 2026

Copy link
Copy Markdown
Member

Audit fallout from today's merges and closes.

The one that matters: README's Harbor entry claimed ~2.7k★ (actual: 3683). I quoted that stale number at a contributor in #34 as the §5a floor, when the real floor is benchflow-ai/benchflow at 305★ — our own project. That PR is reopened and corrected.

README

  • Harbor ~2.7k★~3.7k★
  • "146 deep reading notes" → 143 (two places; notes/ holds 143)
  • Tura (§6): add ⚠️ vendor-published — it benchmarks Ponytail and RTK on Tura's own harness. Disclosed in the PR body, not in the merged line, in the benchmark-integrity section.
  • truescore: add ⚠️ new and unproven, and move §8 → §5d. §8 is papers/posts about judge alignment; §5d is judge/verifier libraries.

whatbroke is intentionally left unmarked: at 10★ it ties patronus-ai/glider and beats agi-inc/REAL (8★), both listed uncaveated. Marking only the newest entry would be a double standard.

CONTRIBUTING

  • Write down the traction and disclosure bars. Five closes today cited rules this file doesn't contain. Stated with the floor honestly low and applied to benchflow-ai projects too.
  • Add the verbatim-number rule.
  • Fix the format block — it specified a backticked bare URL and matched 0 of 462 real entries — and say that eval-mentions go to MENTIONS.md, which the file never mentioned. That omission is a rejection a contributor only finds out about after doing the work.

Fallout from auditing today's merges and closes. Several entries assert
numbers that are no longer true, and one stale number caused me to give a
contributor a factually wrong reason for closing their PR.

README:
- Harbor's star count read ~2.7k★; it is 3683. That figure was quoted at a
  contributor as the §5a floor when closing #34. The actual §5a floor is
  benchflow-ai/benchflow at 305★ — our own project.
- "146 deep reading notes" in two places; notes/ holds 143.
- Tura (§6) gets a vendor-published caveat. The post benchmarks two competing
  plugins on Tura's own harness, which the PR body disclosed and the merged
  line did not — worth stating explicitly in the benchmark-integrity section.
- truescore gets a new-and-unproven caveat and moves from §8 to §5d. §8 is
  papers and blog posts about judge alignment; §5d is judge/verifier
  libraries, which is what truescore is.

whatbroke is deliberately left unmarked: at 10★ it ties patronus-ai/glider
and sits above agi-inc/REAL (8★), both listed without caveats. Marking it
alone would be a double standard.

CONTRIBUTING:
- Write down the traction and disclosure bars. Five closes today cited rules
  this file does not contain, including a star threshold that would exclude
  several currently-listed projects. Stated with the floor honestly low, and
  explicitly applying to benchflow-ai projects too.
- Add the verbatim-number rule the scan already follows.
- Fix the format block, which specified a backticked bare URL and matched
  none of the 462 real entries, and note that MENTIONS.md is where
  eval-mentions go — the file never said so, which is a rejection a
  contributor only discovers after doing the work.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant