Skip to content

fix(ml-inference): handle corrupt-image partial-failure predictions - #536

Merged
Chouffe merged 1 commit into
mainfrom
arthur/ml-better-error-handling
May 15, 2026
Merged

Chouffe merged 1 commit into
mainfrom
arthur/ml-better-error-handling

Conversation

@Chouffe

@Chouffe Chouffe commented May 15, 2026

Copy link
Copy Markdown
Collaborator

Summary

  • Root cause: SpeciesNet silently emits prediction records with a failures field and no prediction when an image can't be loaded (e.g. corrupt JPEG). The Zod validator rejected those, insertModelOutput threw, and the catch block cascaded the failure to every other job in the same batch — fanning a single corrupt file into up to batchSize permanent failures and leaving the progress bar stuck at <100% with no way to resume.
  • Python fix: Added VideoCapableLitAPI._normalize_failed_predictions in utils.py to rewrite partial-failure records into the standard prediction: "error" format used by the existing exception handler. Also backfilled model_version: "unknown" in that handler so DeepFaune/Manas error paths (which already raised on bad images) now pass validation too.
  • JS fix: Added an early-exit guard in InferenceConsumer.processBatch that detects prediction === 'error' or a failures field, marks the job complete (corrupt files aren't retriable), and skips observation creation. Media rows stay observation-less and surface naturally in the blank section via the existing notExists(realObservations) query.

Test plan

  • New unit tests in tests/test_utils.py for _normalize_failed_predictions (7 cases: partial-failure, success passthrough, fallback filepath/model_version, mixed batch, empty, missing key)
  • make lint + make format clean in python-environments/common/
  • Existing JS tests pass (queue-consumer, model-output validator — 45 total)
  • Verified end-to-end on a fresh md-test-images study: 48/48 jobs completed on first attempt, both corrupt-images/ JPEGs end up with 0 observations and surface in the blank section, no batch errors in logs.

SpeciesNet silently emits prediction records with a `failures` field
and no `prediction` when an image can't be loaded, which the Zod
validator rejected and the catch block then propagated to every other
job in the same batch.

- Normalize partial-failure predictions to the standard `prediction: "error"` shape inside `VideoCapableLitAPI.predict_with_video_support` so all ML servers emit validator-compatible output.
- Skip insertPrediction on failed predictions in InferenceConsumer; the media row stays observation-less and surfaces in the blank section.
- Add `model_version: "unknown"` to the existing exception handler so DeepFaune/Manas error paths also pass validation.
@netlify

netlify Bot commented May 15, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for lucent-yeot-0cb408 canceled.

Name Link
🔨 Latest commit 71ca05c
🔍 Latest deploy log https://app.netlify.com/projects/lucent-yeot-0cb408/deploys/6a072678504a4300081664a3

@Chouffe
Chouffe merged commit 8b96d4f into main May 15, 2026
15 checks passed
@Chouffe
Chouffe deleted the arthur/ml-better-error-handling branch May 15, 2026 14:07
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant