Add results.box.image_metrics and errors/ + correct/ visualize folders - #902
Merged
Merged
Conversation
EHxuban11
added a commit
that referenced
this pull request
Sep 26, 2026
Add results.box.image_metrics and errors/ + correct/ visualize folders
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What:
val()results carrybox.image_metrics(per-image precision, recall, F1, TP, FP, FN), andvisualize=Truesorts its images intovisualize/errors/andvisualize/correct/.Why: #887 asks to see only the samples the model got wrong. #890 added the ecosystem's
visualizebut it draws every image; the ecosystem's way to find the wrong ones isresults.box.image_metrics(docs.ultralytics.com/modes/val, docs only).image_metrics: dict, image filename ->{"precision", "recall", "f1", "tp", "fp", "fn"}, same keys as the ecosystem. Always computed for detect and segment (the ecosystem computes it on every run too), with thevisualizematching: confidence 0.25 or the run'sconfif higher, IoU 0.5, class-aware. Precision/recall are 0.0 when their denominator is 0. An image whose filename was already seen (same name in another folder) is keyed by its full path, so no entry is overwritten. Ground truth outsideclasses=is dropped, matching the prediction filter.val()still returns the metrics dict:ValidationMetricsis adictsubclass that adds.box.metrics["..."],==andjson.dumpsare unchanged. Wrapped inmodel.val()and exported-backendval()only; the trainer keeps the plain dict. Other tasks return a plain dict as before.visualizeimages go tovisualize/errors/(any FP or FN; wrong top-1 for classification) orvisualize/correct/, so the mistakes are one folder. Unreleased layout from Add val(visualize=True) error-analysis images #890, so nothing shipped moves. A rerun clears old images from both subfolders.Deviations from the ecosystem, on purpose:
image_metricscount boxes and live on.box; there is no mask-based.seg.image_metrics. Pose and OBB have none yet.visualizeone so the numbers match the drawn images.errors/and filteringimage_metricscover it.Check:
_score_imagesinvalidation/detection_validator.py;with_image_metricsinvalidation/base.py.Verified (macOS, CPU): unit suite 7692 passed, 3 failed (the
test_detr_cpu_export_matrixONNX parity cases that fail on clean dev on this machine and pass in CI).LibreYOLO9ton coco8: mAP identical with and withoutvisualize, 1 image incorrect/and 3 inerrors/, matchingimage_metrics, and each image's TP/FP/FN header matches its entry.Not verified: CUDA, DDP (each rank keeps its own images' entries; not gathered), exported-backend
val()live.Opened by an agent.
Closes #887
Code provenance
Original code written for this PR against LibreYOLO's own first-party code; no third-party code ported, adapted, or introduced; no GPL/AGPL/LGPL/non-commercial/unknown-license material involved. Names follow the ecosystem's public documentation only.
The PR appears safe to merge based on the changes since the previous review.
Summary
The PR adds per-image box metrics to detection and segmentation validation results and sorts validation visualizations into errors and correct folders. Since the previous review, it changes duplicate-filename handling so a later image uses its full path as a metrics key.
Reviews (2) · Last reviewed commit: "Key a repeated image filename by its ful..."