Skip to content

Add custom training fitness callbacks - #900

Merged
EHxuban11 merged 3 commits into
devfrom
feat/custom-fitness
Sep 25, 2026
Merged

EHxuban11 merged 3 commits into
devfrom
feat/custom-fitness

Conversation

@EHxuban11

@EHxuban11 EHxuban11 commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor
  • Add optional fitness(metrics) on a callbacks= object; one finite higher-is-better score controls best.pt, patience and average_best. Refs Allow choosing the early-stopping / best.pt metric (e.g. F1) for detection and classification training #852.
  • Preserve raw validation metrics and existing scorer-free training/callback behavior; propagate scorer failures across DDP ranks.
  • Limits: Python-only; custom-fitness resume is rejected (fresh fine-tuning from saved weights is supported); VLM/VLA reject the hook. No CUDA/NCCL or convergence claim.
  • Verified: 7,666 CPU tests passed; 17 distributed tests passed; isolated wheel smoke passed. Four CPU failures reproduce on unchanged dev (three ONNX parity cases and a D-FINE random-augmentation case at seed 537).
  • Six public-API two-epoch CPU runs compared baseline/default/custom scoring for YOLO9-t and RF-DETR-n: weights, optimizer states, losses and accuracy metrics were bit-identical; RF-DETR custom scoring changed best epoch from 1 to 2. Reload/predict passed. Independent first-party review found no outstanding issues.
  • Pre-push review: fixed FOMO metadata overriding the selected key; 82 focused tests passed, including default/custom FOMO best/last/average checkpoint creation and reload. NumPy scalar rejection did not reproduce; its acceptance test passes. The existing best_mAP50_95 result remains a documented legacy alias for the selected score.
  • Opened by an agent; no approval or merge is implied.

Continue from GitHub

  • Draft handoff: latest pre-push review ad89a4bc-9b8b-4d98-b484-b32edcf1e5f6 is still pending. Resume with greptile review show ad89a4bc-9b8b-4d98-b484-b32edcf1e5f6 where the CLI account is available; otherwise use the GitHub bot review. Do not merge until current CI/review findings are assessed.

  • Prior review follow-ups: preserve FOMO point-task metadata for direct and wrapped trainers; checkpoint/average tests pass. Alleged averaged-scorer deadlock did not reproduce in the new two-process Gloo regression.

  • Branch: feat/custom-fitness; initial implementation commit: cb84b2d5; FOMO follow-ups: 5eae5aa4, 7b0f0295. Based on dev 6884d5eb.

  • Context: Allow choosing the early-stopping / best.pt metric (e.g. F1) for detection and classification training #852's proposed metric aliases were declined; classification metrics landed separately in Add macro precision, recall and F1 to the classification validator #883. This implements Xuban's September 12 callback-based alternative.

  • Read docs/training_loggers.md, docs/checkpoint_schema.md and tests/unit/test_train_fitness.py for the implemented contract. Inspect current CI and Greptile results before deciding to merge; preserve the documented resume restriction unless separately implementing and validating a replacement.

  • Full CPU gate: LIBREYOLO_PR_GATE=1 python -m pytest tests/unit -m "unit and not external_data and not network and not distributed" -n 4 --dist loadfile -q. Distributed gate: same command with -m "unit and not external_data and not network and distributed" and without xdist. Wheel smoke: python tests/smoke/run_install_smoke.py --mode wheel.

  • Implementation session label: implement_fitness_clean (Codex). This description is the portable handoff; no local transcript is required.

Code provenance

Original implementation and tests written from LibreYOLO's own code and the user's requirements in a fresh agent context; no third-party code copied or adapted. An earlier investigation encountered AGPL snippets in search results and stopped before implementation. The maintainer authorized recovery through the fresh implementation context, which consulted no external source; the exposed investigator did not contribute implementation code.

RetriggerConfidence Score: 5/5

The PR appears safe to merge based on the reviewed changes.

Summary

Adds a Python callback hook for custom validation fitness in shared training.

  • Uses the selected score for best-checkpoint selection, patience, and average_best while retaining raw validation metrics.
  • Marks custom-fitness checkpoints and rejects unsupported resume and VLM/VLA usage.
  • Adds focused checkpoint, callback, and distributed tests.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  V[Validation metrics] --> F{Fitness callback?}
  F -->|Yes| S[Score on rank zero]
  F -->|No| D[Use default task metric]
  S --> B[Best state and patience]
  D --> B
  B --> C[Checkpoint and epoch event]
  B --> A[average_best snapshot ranking]
Loading

Reviews (1) · Last reviewed commit: "Keep point metadata for direct FOMO trai..."

@EHxuban11
EHxuban11 marked this pull request as ready for review September 25, 2026 19:33
@EHxuban11
EHxuban11 merged commit 53ebdb9 into dev Sep 25, 2026
14 checks passed
@EHxuban11
EHxuban11 deleted the feat/custom-fitness branch September 26, 2026 15:36
EHxuban11 added a commit that referenced this pull request Sep 26, 2026
Add custom training fitness callbacks
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant