Hard-negative probe retrain on the shared Pillow path - #31
Merged
Merged
Conversation
Extension, Node harness, and probe-training extractor now compute identical DINO inputs. The extractor's local preprocess duplicate is replaced with an import of the shipped function, which also fixes its unguarded grayscale channel reads. Probe retrain on these features follows; the shipped dino-probe.json is unchanged until then.
- Sources already at target count are skipped on rerun, and existing files are never re-downloaded, so a rate-limit abort resumes cheaply - Requests to huggingface.co hosts attach HF_TOKEN or the CLI cached token when present, which raises the datasets-server rate limit - Five hard-negative real sources added (Unsplash stock, Amazon Berkeley product shots, LSUN bedrooms, DeepFashion catalog, Oxford flowers), rows strictly disjoint from the eval stress set, with nested URL columns, a shortest-side gate, and row-index file naming mirroring the exact fetch that built the shipped probe's negatives
The probe head is retrained on 11,405 rows: the standard training set plus 1,915 hard-negative reals (professional stock, product catalogs, interiors, high-saturation nature), all features extracted through the shared Pillow-exact resize. On a 240-image full-resolution stress set this removes every DINO-attributable false positive (shipped v1.1.0 measured 7 FPs on identical bytes; this build measures 4, all driven by CommunityForensics alone at >= 0.65 where it is authoritative). With the cleaner probe the rescue bands re-derive wider under the same stress guard. A conservative near-optimum was chosen over the grid maximum (keeps DINO_RESCUE_MIN 0.70 and the CF-hard-zero guard): DINO_CF_FLOOR 0.02, DINO_STRONG_RESCUE_FLOOR 0.10, DINO_STRONG_RESCUE_MIN 0.90, DINO_SUBFLOOR_MIN 0.995, DINO_SUBFLOOR_CF_MIN 0.0005. Bench (893 imgs, raw 0.65, harness-verified): 90.5% BA, 81.2% TPR, 99.8% TNR, vs 87.7 / 75.8 / 99.6 for v1.1.0. Browser-vs-Node DINO parity: max delta 0.000046 on a live-browser sample (was up to 0.22). Full suite: 213 tests green.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Retrains the DINO probe head on features from the shared Pillow-exact resize, with 1,915 hard-negative reals added to training, and re-derives the fusion rescue bands for the cleaner probe. Three commits:
Why hard negatives
The naive Pillow-path retrain regressed on professional stock photography: the training reals (COCO, Flickr, faces, food, scenes) contain no polished stock/product distribution, so the probe saturated on exactly the images that sank claim 1014-style submissions. Adding 1,915 such reals to training removed every DINO-attributable stress false positive.
Measured results (harness runs, not simulation)
893-image public bench at raw 0.65:
Per-generator recall: DALL-E 3 93.3% approx, MJHQ 96.7%, Flux 1.1 83.3%, GPT-4o image 60.0%, SD-ELSA 78.5% (winner-detail figures; the shipped conservative variant is within about a point of each).
240-image full-resolution stock/catalog/product stress set, identical bytes across configurations:
The 4 remaining stress FPs are CommunityForensics alone scoring >= 0.65 on full-resolution Unsplash originals, where CF is authoritative by design. Worth checking exactly that during the pre-claim live smoke.
Full suite: 213 tests green. Constants chosen conservatively: DINO_RESCUE_MIN stays 0.70 and the CF-hard-zero guard stays, giving up 1.4 bench BA versus the grid maximum in exchange for margin on unmeasured live distributions.
Caveats
Note
Medium Risk
Changes core scoring (probe weights and fusion constants) and DINO preprocessing parity; bench/stress metrics are documented but live-site smoke and full refetch validation are still caveated in the PR.
Overview
Retrains and ships a new DINO logistic probe (
models/probe/dino-probe.json, ~11.4k training rows) on features from the shared Pillow-exact resize, with ~1.9k hard-negative reals (stock, catalog, product, interiors) added via extendedeval/fetch-train.mjs(resumable fetches, optionalHF_TOKEN, disjoint stress-set rows).DINO inputs are unified end-to-end:
src/dino.jsandeval/extract-features.mjsnow usepillowResize+cropPackedPixels(optionalcropSizeinclip-preprocess.js) so extension, Node harness, and training match; the extractor’s duplicate preprocess is removed.Fusion rescue tiers are re-derived in
src/fuse.jsfor the new probe (e.g. strong band0.02–0.10/ DINO ≥0.90, sub-floor CF ≥0.0005/ DINO ≥0.995). Docs and tests reflect harness-reported 90.5% BA on the 893-image fixture and 0 DINO-attributable stress-set FPs vs v1.1.0’s 7 total / 3 DINO-driven on the same bytes.Reviewed by Cursor Bugbot for commit 5fcaba9. Bugbot is set up for automated code reviews on this repo. Configure here.