Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,11 @@ changes, new features. **Patch:** fixes. Images are published to
a release is validated. Release process: [runbook](docs/operate/runbook.md#release-a-new-version). Gateway/CLI-only releases
(`make release-gateway`) reuse the previous serving images.

## Unreleased

- `bench-intents`: more than 26 candidate intents run as a 2-stage bracket (groups of ≤ 20, top 5 to a final round)
instead of being rejected by the server (`--dataset banking77` sends 30).

## v0.1.2 (2026-09-28)

Gateway and monitoring; serving images unchanged (v0.1.0).
Expand Down
8 changes: 5 additions & 3 deletions benchmarks/runs/20260928-v010-verification/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@ committed receipts in `benchmarks/` (earlier images and hardware) and with the i
| `bench-calibration` | 50 | 46 (92%) | 43–46 (image parity) | Pass |
| `bench` (triage, code review, security) | 30 | 22 (73.3%) | 23–24 on L4 receipts | Pass (±2 items is noise on 30) |
| `bench-intents` clinc150 (small set) | 30 | 29 (96.7%) after tie-break; domain 28/30 | Committed receipt 96.7% | Pass |
| `bench-intents` banking77 (small set) | 30 | 0: every request rejected, `at most 26 alternatives` | Committed receipt (Sep 20, older server) | **Harness issue**, see below |
| `bench-intents` banking77 (small set) | 30 | 0 at first (every request rejected, `at most 26 alternatives`); after the bracket fix: **23 (76.7%)** | Committed receipt 26 (Sep 20, older server allowed 30 options in one question) | Fixed (harness) |
| `bench-bbox` | 12 | mIoU 0.406 argmax / 0.422 expectation | 0.290 / 0.377 (Cloud Run L4) | Pass |
| `bench-rerank` | 30 queries | nDCG@10 0.824, 0.856, 0.856 (3 runs); 0% ties, 100% poison quarantine | 0.9265 Cloud Run L4, 0.8502 Vertex L4 | Pass (in the range of earlier receipts) |
| `bench-permutation` | 16 | 100% accuracy, 6.25% flip rate | EXP-13 | Pass |
Expand All @@ -21,9 +21,11 @@ committed receipts in `benchmarks/` (earlier images and hardware) and with the i
| `bench-ecotone` | — | Not run (needs the C++ ecotone sidecar) | — | Skipped |

**Two findings, both pre-existing, not regressions:**
- `bench-intents --dataset banking77` sends all 30 candidate intents in one `choice` question. The server has
- `bench-intents --dataset banking77` sent all 30 candidate intents in one `choice` question. The server has
rejected more than 26 options since before v0.1.0 (`ab208dd` had the same check); the committed receipt predates
it. The harness should split into brackets (as `dgem systemone serve` does) or use the 26-option slice.
it. **Fixed:** with more than 26 options the harness now runs a 2-stage bracket (groups of ≤ 20, top 5 of each
group to a final round), like `dgem systemone serve`. Result 23/30 (76.7%); keeping the top 3 per group gave 22,
the top 8 gave 23. The misses are confusable intent pairs (e.g. `card_arrival` vs `card_delivery_estimate`).
- `intent_banking77`, `intent_clinc150`, `intent_tiebreak`, `jevbench_generic` and `tn_disambiguation` are harness
templates: they take option lists from `bench-intents`, `bench-jev` or `bench-ecotone`, so the gateway's sample
variables leave them with fewer than two options (400). The harnesses themselves pass (above).
Expand Down
Loading
Loading