Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
34 changes: 34 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,11 @@ before 1.4.0 are documented in the

### Added

- **`LibreRFDETRm-ui.pt`: class-agnostic UI element detector (#896).**
UI-DETR-1 (racineai, MIT), an RF-DETR-M fine-tune for screenshots, hosted
as an RF-DETR dataset variant with one class, `object`. Auto-downloads from
`LibreYOLO/LibreRFDETRm-ui`.

- **`val(visualize=True)` draws every validated image for error analysis
(#887).** Detection and segmentation images show true positives (green),
false positives (red) and false negatives (orange), matched class-aware at
Expand Down Expand Up @@ -148,6 +153,35 @@ before 1.4.0 are documented in the

### Fixed

- **`Results.plot()` draws boxes, masks, OBB, keypoints and every other
predict output (#896).** It raised `NotImplementedError` for detection,
segmentation, pose, OBB, classification, points, OCR, semantic and panoptic
results. It now returns the annotated image as an `HxWx3` uint8 BGR array,
pixel-identical to `predict(save=True)` (both use the new
`drawing.draw_results`). Accepts `conf`, `labels`, `boxes`, `masks`,
`probs`, `line_width`, `pil`, `img`, `show`, `save` and `filename`.
Classification plots and `predict(save=True)` write the top-5 labels.
Tracked results show their track IDs. Exported-model backends (ONNX and the
other runtimes) save and plot through the same renderer. Dense maps, 3D
cuboids and action chunks keep returning a PIL image unless `pil=False`.
Predict keeps in-memory and URL inputs on the result as `Results.orig_img`
(BGR) so they plot without a path; local files are not kept and `plot()`
reopens them, so directory predictions stay light. A finite video or GIF
collected without `stream=True` keeps no frames; `plot()` decodes the
result's frame on demand.

- **RF-DETR predict and val resize without antialiasing (#896).** The PIL
bilinear resize antialiased on downscale, unlike RF-DETR training (cv2
bilinear) and upstream `predict()` since rf-detr 1.9.0. Boxes drifted on
inputs much larger than the canvas (median IoU 0.71 to 0.88 against upstream
on UI screenshots); they now match upstream on screenshots and COCO images,
in PyTorch and ONNX. COCO mAP50-95 of RF-DETR-M on a 200-image val subset
moves from 0.6195 to 0.6179.

- **Single-class upstream RF-DETR checkpoints convert as `nc=1` (#896).** A
one-output class head (a one-category training set) was converted as
`nc=80` with 79 placeholder names. Predictions were unaffected.

- **Classification `val()`, INT8 calibration and exports use the model's own
eval pipeline (#886).** The validator now takes the transform from the model
(`_get_eval_transform`), the classification counterpart of
Expand Down
3 changes: 2 additions & 1 deletion docs/nomenclature.md
Original file line number Diff line number Diff line change
Expand Up @@ -533,7 +533,7 @@ Detector-factory family support follows:
| `rtmdet` | `("detect", "segment")` (default: detect) | detect | RTMDet-Ins uses `-seg`; detect training is implemented and directly callable, segment training is not implemented |
| `picodet` | `("detect",)` (default) | detect | detect-only |
| `ppyoloe` | `("detect",)` (default) | detect | detect-only; no task suffix; pretrained weights linked from the source CDN, not mirrored |
| `rfdetr` | `("detect", "segment", "pose", "obb")` | detect | seg uses smaller sizes; pose/OBB use detect sizes |
| `rfdetr` | `("detect", "segment", "pose", "obb")` | detect | seg uses smaller sizes; pose/OBB use detect sizes; dataset-variant weights `-ui` (class-agnostic UI elements, nc=1) |
| `dinov2` | `("semantic", "classify", "embed")` | semantic | DINOv2 backbone + task head; embed bypasses heads and returns the 384-d final CLS token at 224 (all sizes share DINOv2-S); no text tower |
| `eomt` | `("semantic", "segment", "panoptic")` | semantic | DINOv2 backbone; sizes s/b/l. Semantic: ADE20K 150-class at 512. Instance segment: COCO 80-class at 640 (l also at 1280). Panoptic: COCO 133-class at 640. Upstream ships no COCO instance checkpoint at s/b. DINOv3 variants excluded |
| `pidnet` | `("semantic",)` | semantic | real-time PIDNet semantic segmentation; s/m/l at 1024; Cityscapes 19-class checkpoints; inference + `val`; not trainable in LibreYOLO |
Expand Down Expand Up @@ -649,6 +649,7 @@ LibreRFDETRn.pt # detect
LibreRFDETRn-seg.pt # segment
LibreRFDETRx-pose.pt # pose (preview; only size x ships)
LibreRFDETRn-obb.pt # obb
LibreRFDETRm-ui.pt # detect; UI element variant (UI-DETR-1, 1 class)

# dinov2 — DINOv2 backbone + task head (NOT the RF-DETR detector)
LibreDINOv2n.pt # semantic (default task; dense head at 518)
Expand Down
128 changes: 33 additions & 95 deletions libreyolo/backends/base.py
Original file line number Diff line number Diff line change
Expand Up @@ -44,14 +44,7 @@
)
from ..preprocess.yolox import preprocess_image as yolox_preprocess_image
from ..tasks import normalize_supported_tasks, normalize_task, resolve_task
from ..utils.drawing import (
draw_boxes,
draw_keypoints,
draw_masks,
draw_obb,
draw_points,
draw_semantic_mask,
)
from ..utils.drawing import draw_results
from ..utils.general import (
COCO_CLASSES,
get_safe_stem,
Expand All @@ -62,6 +55,7 @@
from ..utils.model_info import build_model_info, format_model_info
from ..utils.predict_args import normalize_predict_kwargs
from ..utils.results import (
keep_source,
Boxes,
DepthMap,
EdgeMap,
Expand Down Expand Up @@ -4091,80 +4085,12 @@ def _build_result(

def _save_annotated(self, result, original_img, image_path, output_path):
"""Save annotated image to disk."""
annotated_img = original_img
forced_ext = None
if result.boxes is None and getattr(result, "probs", None) is not None:
pass
elif result.boxes is None and getattr(result, "restored", None) is not None:
annotated_img = Image.fromarray(result.restored.array, mode="RGB")
elif result.boxes is None and getattr(result, "depth_map", None) is not None:
from ..utils.drawing import draw_depth_map

depth_data = result.depth_map.data
if isinstance(depth_data, torch.Tensor):
depth_data = depth_data.cpu().numpy()
annotated_img = draw_depth_map(original_img, depth_data)
elif result.boxes is None and getattr(result, "normal_map", None) is not None:
from ..utils.drawing import draw_normal_map

normal_data = result.normal_map.data
if isinstance(normal_data, torch.Tensor):
normal_data = normal_data.cpu().numpy()
annotated_img = draw_normal_map(original_img, normal_data)
elif result.boxes is None and getattr(result, "edges", None) is not None:
from ..utils.drawing import draw_edge_map

edge_data = result.edges.data
if isinstance(edge_data, torch.Tensor):
edge_data = edge_data.cpu().numpy()
annotated_img = draw_edge_map(original_img, edge_data)
elif (
result.boxes is None and getattr(result, "semantic_mask", None) is not None
):
mask_data = result.semantic_mask.data
if isinstance(mask_data, torch.Tensor):
mask_data = mask_data.cpu().numpy()
annotated_img = draw_semantic_mask(original_img, mask_data)
elif result.boxes is None and getattr(result, "matte", None) is not None:
if result.boxes is None and getattr(result, "matte", None) is not None:
annotated_img = Image.fromarray(result.cutout(original_img), mode="RGBA")
forced_ext = "png"
elif result.boxes is None and getattr(result, "points", None) is not None:
if len(result.points) > 0:
annotated_img = draw_points(
original_img,
result.points.xy.tolist(),
result.points.conf.tolist(),
result.points.cls.tolist(),
class_names=result.names,
)
elif len(result) > 0:
if result.masks is not None:
annotated_img = draw_masks(
annotated_img,
result.masks.data.numpy(),
result.boxes.cls.tolist(),
)
if result.obb is not None:
annotated_img = draw_obb(
annotated_img,
result.obb.xywhr.tolist(),
result.obb.conf.tolist(),
result.obb.cls.tolist(),
class_names=self.names,
)
else:
annotated_img = draw_boxes(
annotated_img,
result.boxes.xyxy.tolist(),
result.boxes.conf.tolist(),
result.boxes.cls.tolist(),
class_names=self.names,
)
if result.keypoints is not None:
kpts_np = result.keypoints.data
if isinstance(kpts_np, torch.Tensor):
kpts_np = kpts_np.cpu().numpy()
annotated_img = draw_keypoints(annotated_img, kpts_np)
else:
annotated_img = draw_results(result, original_img)

ext = forced_ext or (
"png" if isinstance(getattr(self, "input_profile", None), dict)
Expand Down Expand Up @@ -4558,6 +4484,11 @@ def val(
# Inference pipeline
# =========================================================================

@staticmethod
def _keep_source(result, original_img, source=None):
"""Keep the decoded source image on a result so ``plot()`` works."""
return keep_source(result, original_img, source)

def _predict_single(
self,
image: Union[str, Path, Image.Image, np.ndarray],
Expand Down Expand Up @@ -4602,12 +4533,16 @@ def _predict_single(
image_path if image_path is not None else save_stem,
output_path,
)
return result
return self._keep_source(result, original_img, image_path)
if self.task == "embed":
return self._build_embedding_result(
all_outputs,
orig_shape=orig_shape,
image_path=image_path,
return self._keep_source(
self._build_embedding_result(
all_outputs,
orig_shape=orig_shape,
image_path=image_path,
),
original_img,
image_path,
)
if self.task == "restore":
result = self._build_restore_result(
Expand All @@ -4623,7 +4558,7 @@ def _predict_single(
image_path if image_path is not None else save_stem,
output_path,
)
return result
return self._keep_source(result, original_img, image_path)
if self.task == "depth":
result = self._build_depth_result(
all_outputs,
Expand All @@ -4638,7 +4573,7 @@ def _predict_single(
image_path if image_path is not None else save_stem,
output_path,
)
return result
return self._keep_source(result, original_img, image_path)
if self.task == "normal":
result = self._build_normal_result(
all_outputs,
Expand All @@ -4653,7 +4588,7 @@ def _predict_single(
image_path if image_path is not None else save_stem,
output_path,
)
return result
return self._keep_source(result, original_img, image_path)
if self.task == "edge":
result = self._build_edge_result(
all_outputs,
Expand All @@ -4668,7 +4603,7 @@ def _predict_single(
image_path if image_path is not None else save_stem,
output_path,
)
return result
return self._keep_source(result, original_img, image_path)
if self.task == "matte":
result = self._build_matte_result(
all_outputs,
Expand All @@ -4683,7 +4618,7 @@ def _predict_single(
image_path if image_path is not None else save_stem,
output_path,
)
return result
return self._keep_source(result, original_img, image_path)
if self.task == "gaze":
result = self._build_gaze_result(
all_outputs,
Expand All @@ -4697,7 +4632,7 @@ def _predict_single(
image_path if image_path is not None else save_stem,
output_path,
)
return result
return self._keep_source(result, original_img, image_path)
if self.task == "semantic":
result = self._build_semantic_result(
all_outputs,
Expand All @@ -4714,7 +4649,7 @@ def _predict_single(
image_path if image_path is not None else save_stem,
output_path,
)
return result
return self._keep_source(result, original_img, image_path)
if self.task == "point":
result = self._build_point_result(
all_outputs,
Expand All @@ -4732,7 +4667,7 @@ def _predict_single(
image_path if image_path is not None else save_stem,
output_path,
)
return result
return self._keep_source(result, original_img, image_path)

parsed = self._parse_outputs(
all_outputs,
Expand Down Expand Up @@ -4769,7 +4704,7 @@ def _predict_single(
output_path,
)

return result
return self._keep_source(result, original_img, image_path)

def _supports_batched_inference(self) -> bool:
"""Whether ``_run_inference`` accepts stacked (N, C, H, W) blobs.
Expand Down Expand Up @@ -5063,7 +4998,7 @@ def _predict_batch(

if save:
self._save_annotated(result, original_img, save_name, output_path)
results.append(result)
results.append(self._keep_source(result, original_img, image_path))
return results

# =========================================================================
Expand Down Expand Up @@ -5243,7 +5178,7 @@ def _predict_video(
source_label = str(source) if source_label is None else source_label
effective_imgsz = self._resolve_predict_imgsz(imgsz)

def predict_frame(pil_img):
def predict_frame_result(pil_img):
input_tensor, original_img, original_size, ratio = self._preprocess(
pil_img, effective_imgsz, "rgb"
)
Expand Down Expand Up @@ -5349,6 +5284,9 @@ def predict_frame(pil_img):
max_det=max_det,
)

def predict_frame(pil_img):
return self._keep_source(predict_frame_result(pil_img), pil_img)

yield from run_video_inference(
source,
predict_frame,
Expand Down
2 changes: 1 addition & 1 deletion libreyolo/backends/tensorrt.py
Original file line number Diff line number Diff line change
Expand Up @@ -522,7 +522,7 @@ def _process_in_batches(
if save:
self._save_annotated(result, orig_img, save_name, output_path)

results.append(result)
results.append(self._keep_source(result, orig_img, image_path))

return results

Expand Down
8 changes: 8 additions & 0 deletions libreyolo/models/autoconvert.py
Original file line number Diff line number Diff line change
Expand Up @@ -645,6 +645,14 @@ def _rfdetr_class_metadata(
# COCO arch-classes (91 outputs incl. background) -> LibreYOLO's COCO-80.
return 80, _checkpoint_names(loaded, 80)

if raw_nc == 0:
# RF-DETR scores classes with independent sigmoids (focal loss); there
# is no background logit. The usual extra output is an unused index
# slot for COCO-style ids, but upstream sizes the head from the
# dataset's category count, so a one-category dataset yields a single
# output and logit 0 is that class (upstream predicts class_id 0).
return 1, _checkpoint_names(loaded, 1)

nc = raw_nc if raw_nc else 80
return nc, _checkpoint_names(loaded, nc)

Expand Down
Loading
Loading