Skip to content

chore(release): 0.7.6 — digest-pinned, honesty-classified confined baseline - #1501

Merged
hartsock merged 2 commits into
mainfrom
chore/release-0.7.6
Jul 31, 2026
Merged

chore(release): 0.7.6 — digest-pinned, honesty-classified confined baseline#1501
hartsock merged 2 commits into
mainfrom
chore/release-0.7.6

Conversation

@hartsock

@hartsock hartsock commented Jul 31, 2026

Copy link
Copy Markdown
Member

Summary

Release prep for 0.7.6: bumps the workspace version 0.7.5 → 0.7.6 (package + the 17 internal =0.7.x pins + Cargo.lock), cuts CHANGELOG [0.7.6], and banks the confined (OCAP-on) tb-30 scoreboard under the honesty-classified, digest-pinned methodology.

Confined champions: ornith 36.7%, qwen3.6 26.7%, qwen3-coder 13.3%, kimi-linear 10.0%, glm-4.7-flash 3.3%. Three models were quarantined for under-engagement (granite4.1, nemotron-3-nano canonical + bare) — tracked in #1500.

This is a re-baseline under stricter measurement, not a raw champion-defeat; tb-30 is a 30-task sample (~±10pp variance). Parity (OCAP-on ≈ off) is pursued forward via pre-granted permissions, not gated here.

Test plan

  • python3 scripts/eval/bench_scoreboard.py --self-test → OK
  • cargo update --workspace resolves cleanly (lock synced to 0.7.6)
  • CI: full just check on this PR

Out of scope

Refs #1500.


Note

Low Risk
Version and documentation/eval-data only; no application logic, auth, or runtime behavior changes.

Overview
Release 0.7.6 bumps the workspace and all lock-step internal crate pins from 0.7.5 to 0.7.6 (Cargo.toml, Cargo.lock) and adds a CHANGELOG [0.7.6] entry documenting the confined (OCAP-on) tb-30 re-baseline under stricter measurement (≥25/30 tool-call engagement, GGUF model_digest pinning, quarantine for under-engaged models).

Benchmark artifacts extend scripts/eval/bench-results.jsonl with model_digest on key OCAP-on runs and new confined results for kimi-linear_48b and glm-4.7-flash. The auto-generated README scoreboard and bench_scoreboard.py header now describe 0.7.6 as establishing that honesty-classified baseline and note that OCAP parity is pursued forward, not gated on this release.

Reviewed by Cursor Bugbot for commit 2ffb69b. Configure here.

hartsock and others added 2 commits July 31, 2026 06:50
…seline

Workspace version 0.7.5 -> 0.7.6 (package + 17 internal =0.7.x pins + Cargo.lock).

Banks the confined (OCAP-on) tb-30 scoreboard under a stricter methodology:
- honesty classifier: a run counts only if the model made a real tool-call
  attempt on >=25/30 tasks, else quarantined (3 quarantined this cycle:
  granite4.1, nemotron-3-nano canonical + bare — see #1500).
- each local score pinned to the served GGUF sha256 (a name is not an identity;
  cf. the Gemma 4 silent re-upload).

Confined champions: ornith 36.7%, qwen3.6 26.7%, qwen3-coder 13.3%,
kimi-linear 10.0%, glm-4.7-flash 3.3%. A re-baseline under stricter measurement,
not a raw champion-defeat; tb-30 is a 30-task sample (~+/-10pp variance).

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
@hartsock
hartsock merged commit 2b9a9a4 into main Jul 31, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant