Skip to content

Hard-negative probe retrain on the shared Pillow path - #31

Merged
felirami merged 3 commits into
mainfrom
fable/dino-pillow-probe
Aug 16, 2026
Merged

felirami merged 3 commits into
mainfrom
fable/dino-pillow-probe

Conversation

@felirami

@felirami felirami commented Aug 16, 2026 •

Copy link
Copy Markdown
Collaborator

What

Retrains the DINO probe head on features from the shared Pillow-exact resize, with 1,915 hard-negative reals added to training, and re-derives the fusion rescue bands for the cleaner probe. Three commits:

  1. DINO preprocess unification (src/dino.js, eval/extract-features.mjs): extension, Node harness, and probe training now compute identical DINO inputs through src/pixel-resize.js. The extractor's private preprocess duplicate is deleted (it had drifted and carried an unguarded grayscale read). Browser-vs-Node DINO parity measured in a live browser: max delta 0.000046, down from 0.22.
  2. fetch-train hardening: resumable (rate-limit aborts resume cheaply), authenticated (HF token picked up automatically), and extended with the five hard-negative real sources (Unsplash stock, Amazon Berkeley product shots, LSUN bedrooms, DeepFashion catalog, Oxford flowers; rows strictly disjoint from the eval stress set).
  3. Probe + bands (models/probe/dino-probe.json, src/fuse.js): retrained on 11,405 rows; bands re-derived under the stress guard with a conservative near-optimum chosen over the grid maximum.

Why hard negatives

The naive Pillow-path retrain regressed on professional stock photography: the training reals (COCO, Flickr, faces, food, scenes) contain no polished stock/product distribution, so the probe saturated on exactly the images that sank claim 1014-style submissions. Adding 1,915 such reals to training removed every DINO-attributable stress false positive.

Measured results (harness runs, not simulation)

893-image public bench at raw 0.65:

Configuration BA TPR TNR
v1.1.0 (shipped) 87.7% 75.8% 99.6%
Naive Pillow retrain (not shipped) 87.2% 74.8% 99.6%
This PR 90.5% 81.2% 99.8%

Per-generator recall: DALL-E 3 93.3% approx, MJHQ 96.7%, Flux 1.1 83.3%, GPT-4o image 60.0%, SD-ELSA 78.5% (winner-detail figures; the shipped conservative variant is within about a point of each).

240-image full-resolution stock/catalog/product stress set, identical bytes across configurations:

Configuration Stress FPs Of which DINO-attributable
v1.1.0 (shipped) 7 3
Naive Pillow retrain 9 5
This PR 4 0

The 4 remaining stress FPs are CommunityForensics alone scoring >= 0.65 on full-resolution Unsplash originals, where CF is authoritative by design. Worth checking exactly that during the pre-claim live smoke.

Full suite: 213 tests green. Constants chosen conservatively: DINO_RESCUE_MIN stays 0.70 and the CF-hard-zero guard stays, giving up 1.4 bench BA versus the grid maximum in exchange for margin on unmeasured live distributions.

Caveats

  • The hard-negative additions to fetch-train mirror the exact datasets-server rows and rules that built the shipped probe's negatives, but an end-to-end refetch validation is pending a rate-limit cooldown (today's fetch volume exhausted even the authenticated tier).
  • Public fixture numbers are directional and are not a claim about the private benchmark. Live-site smoke remains required before any claim.

Note

Medium Risk
Changes core scoring (probe weights and fusion constants) and DINO preprocessing parity; bench/stress metrics are documented but live-site smoke and full refetch validation are still caveated in the PR.

Overview
Retrains and ships a new DINO logistic probe (models/probe/dino-probe.json, ~11.4k training rows) on features from the shared Pillow-exact resize, with ~1.9k hard-negative reals (stock, catalog, product, interiors) added via extended eval/fetch-train.mjs (resumable fetches, optional HF_TOKEN, disjoint stress-set rows).

DINO inputs are unified end-to-end: src/dino.js and eval/extract-features.mjs now use pillowResize + cropPackedPixels (optional cropSize in clip-preprocess.js) so extension, Node harness, and training match; the extractor’s duplicate preprocess is removed.

Fusion rescue tiers are re-derived in src/fuse.js for the new probe (e.g. strong band 0.02–0.10 / DINO ≥ 0.90, sub-floor CF ≥ 0.0005 / DINO ≥ 0.995). Docs and tests reflect harness-reported 90.5% BA on the 893-image fixture and 0 DINO-attributable stress-set FPs vs v1.1.0’s 7 total / 3 DINO-driven on the same bytes.

Reviewed by Cursor Bugbot for commit 5fcaba9. Bugbot is set up for automated code reviews on this repo. Configure here.

Extension, Node harness, and probe-training extractor now compute
identical DINO inputs. The extractor's local preprocess duplicate is
replaced with an import of the shipped function, which also fixes its
unguarded grayscale channel reads. Probe retrain on these features
follows; the shipped dino-probe.json is unchanged until then.
- Sources already at target count are skipped on rerun, and existing
  files are never re-downloaded, so a rate-limit abort resumes cheaply
- Requests to huggingface.co hosts attach HF_TOKEN or the CLI cached
  token when present, which raises the datasets-server rate limit
- Five hard-negative real sources added (Unsplash stock, Amazon
  Berkeley product shots, LSUN bedrooms, DeepFashion catalog, Oxford
  flowers), rows strictly disjoint from the eval stress set, with
  nested URL columns, a shortest-side gate, and row-index file naming
  mirroring the exact fetch that built the shipped probe's negatives
The probe head is retrained on 11,405 rows: the standard training set
plus 1,915 hard-negative reals (professional stock, product catalogs,
interiors, high-saturation nature), all features extracted through the
shared Pillow-exact resize. On a 240-image full-resolution stress set
this removes every DINO-attributable false positive (shipped v1.1.0
measured 7 FPs on identical bytes; this build measures 4, all driven
by CommunityForensics alone at >= 0.65 where it is authoritative).

With the cleaner probe the rescue bands re-derive wider under the same
stress guard. A conservative near-optimum was chosen over the grid
maximum (keeps DINO_RESCUE_MIN 0.70 and the CF-hard-zero guard):
DINO_CF_FLOOR 0.02, DINO_STRONG_RESCUE_FLOOR 0.10,
DINO_STRONG_RESCUE_MIN 0.90, DINO_SUBFLOOR_MIN 0.995,
DINO_SUBFLOOR_CF_MIN 0.0005.

Bench (893 imgs, raw 0.65, harness-verified): 90.5% BA, 81.2% TPR,
99.8% TNR, vs 87.7 / 75.8 / 99.6 for v1.1.0. Browser-vs-Node DINO
parity: max delta 0.000046 on a live-browser sample (was up to 0.22).
Full suite: 213 tests green.
@felirami
felirami merged commit fcac826 into main Aug 16, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant