Skip to content

Run RF-DETR at the settings its authors used for RF100-VL - #14

Merged
EHxuban11 merged 1 commit into
rf100vl-harnessfrom
rfdetr-roboflow-protocol
Aug 13, 2026
Merged

EHxuban11 merged 1 commit into
rf100vl-harnessfrom
rfdetr-roboflow-protocol

Conversation

@EHxuban11

@EHxuban11 EHxuban11 commented Aug 13, 2026 •

Copy link
Copy Markdown
Contributor

Our RF-DETR recipe came from the rf-detr library defaults. Roboflow did not use those defaults for their own RF100-VL numbers.

Source: rf-detr#266, maintainer isaacrob-roboflow, 2025-07-23:

The RF-DETR numbers for the current release are using the default arguments but with batch size 16 and grad accum 1

Posted 3 hours after the 1.2.0 release, so "default arguments" = 1.2.0 defaults. Checked against the 1.2.0 tag, not develop, because develop drifted (batch_size now takes "auto", amp_dtype became a selector).

Changes

field was now why
physical_batch 4 16 Roboflow's one override. accum 4 -> 1
precision fp32 bfloat16 engine.py::get_autocast_args hardcodes dtype=torch.bfloat16, amp defaults True
cache disk removed RF-DETR takes the PRE-resize cache point (full-res pixels, decode-skip only): big disk, small win
workers 0 2 upstream num_workers. Other families run 4
backbone_lr_mult 0.1 removed dead key, see below
dense_oom_fallback 2 4 keep the rescue proportional to the new base

Batch is not cosmetic. Effective batch was already 16, but multi_scale draws a fresh resolution per forward, so accum 4 averaged 4 scales into one optimizer step where Roboflow averaged 1.

backbone_lr_mult was never applied. LibreYOLO hardcodes Roboflow's lr_encoder 1.5e-4, lr_vit_layer_decay 0.8, lr_component_decay 0.7 in models/rfdetr/nn.py, and the fallback that would read the key is unreachable because the DINOv2 backbone exposes get_named_param_lr_pairs. Recipes get published as provenance, so a key we do not apply is a false claim about how the numbers were made.

Recorded as unmatchable, in the recipe

  • Roboflow initialised from private Objects365 checkpoints, never released. They estimate ~0.1 mAP.
  • Their published table predates their own fix for an incorrectly initialised head. Not reproducible by any current code.
  • Size l stays at batch 2. 704 px at batch 16 does not fit 16 GB.
  • Resize is cv2 INTER_LINEAR, no antialias.

Verification

  • 2-epoch 8-GPU probe of rfdetr-s on this exact recipe: 0 OOM, peak 15524 MiB of 16 GB.
  • 167 tests pass.
  • New test pins batch 16 / bf16 / no-cache so a copy-paste from another family cannot silently undo it.
  • rfdetr added to NON_FP32_FAMILIES with the upstream citation, which is the gate that caught the precision change.

No published results are affected. RF-DETR has 0 datasets on the Hub.

https://claude.ai/code/session_01Gy5rL27krK8eMZGCP7xozD

Our RF-DETR recipe was assembled from the rf-detr library defaults. Roboflow
did not use those defaults for their own RF100-VL numbers. Their maintainer
states the protocol in rf-detr#266: "default arguments but with batch size 16
and grad accum 1", posted three hours after the 1.2.0 release, so "default
arguments" means the 1.2.0 defaults.

Checked against the 1.2.0 tag rather than develop, because develop has drifted
(batch_size now accepts "auto", amp_dtype became a selector):

- physical_batch 4 -> 16. Effective batch was already 16, but via grad accum 4.
  That is not cosmetic: multi_scale draws a fresh resolution per forward, so
  accum 4 averages four scales into one optimizer step where Roboflow averaged
  one.
- precision fp32 -> bfloat16. rfdetr/engine.py::get_autocast_args hardcodes
  dtype=torch.bfloat16 and ModelConfig.amp defaults True. fp32 matched no
  upstream setting.
- cache "disk" dropped. A post-resize cache pins one resolution per image,
  which quietly defeats the multi_scale this recipe asks for.
- workers 0 -> 2, matching upstream num_workers. Every other family already
  runs 4.
- backbone_lr_mult removed. It was dead: LibreYOLO hardcodes Roboflow's
  lr_encoder 1.5e-4, lr_vit_layer_decay 0.8 and lr_component_decay 0.7 in
  models/rfdetr/nn.py, and the fallback that would have read this key is
  unreachable because the DINOv2 backbone exposes get_named_param_lr_pairs.
  A recipe is published as provenance, so a key we do not apply is a false
  claim about how the numbers were produced.
- dense_oom_fallback 2 -> 4, keeping the rescue path proportional to the new
  base batch.

Recorded in the recipe what cannot be matched: Roboflow initialised from
private Objects365 checkpoints that were never released (they estimate ~0.1
mAP), and their published table predates their own fix for an incorrectly
initialised head, so those numbers are not reproducible by any current code.
Size l still steps down to batch 2 because 704 px at batch 16 does not fit
16 GB.

Verified by a 2-epoch 8-GPU probe of rfdetr-s on this exact recipe: no OOM,
peak 15524 MiB of 16 GB.

Claude-Session: https://claude.ai/code/session_01Gy5rL27krK8eMZGCP7xozD
@EHxuban11
EHxuban11 merged commit be9ce56 into rf100vl-harness Aug 13, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant