Add custom training fitness callbacks - #900
Merged
Merged
Conversation
EHxuban11
marked this pull request as ready for review
September 25, 2026 19:33
EHxuban11
added a commit
that referenced
this pull request
Sep 26, 2026
Add custom training fitness callbacks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
fitness(metrics)on acallbacks=object; one finite higher-is-better score controlsbest.pt, patience andaverage_best. Refs Allow choosing the early-stopping / best.pt metric (e.g. F1) for detection and classification training #852.best_mAP50_95result remains a documented legacy alias for the selected score.Continue from GitHub
Draft handoff: latest pre-push review
ad89a4bc-9b8b-4d98-b484-b32edcf1e5f6is still pending. Resume withgreptile review show ad89a4bc-9b8b-4d98-b484-b32edcf1e5f6where the CLI account is available; otherwise use the GitHub bot review. Do not merge until current CI/review findings are assessed.Prior review follow-ups: preserve FOMO point-task metadata for direct and wrapped trainers; checkpoint/average tests pass. Alleged averaged-scorer deadlock did not reproduce in the new two-process Gloo regression.
Branch:
feat/custom-fitness; initial implementation commit:cb84b2d5; FOMO follow-ups:5eae5aa4,7b0f0295. Based on dev6884d5eb.Context: Allow choosing the early-stopping / best.pt metric (e.g. F1) for detection and classification training #852's proposed metric aliases were declined; classification metrics landed separately in Add macro precision, recall and F1 to the classification validator #883. This implements Xuban's September 12 callback-based alternative.
Read
docs/training_loggers.md,docs/checkpoint_schema.mdandtests/unit/test_train_fitness.pyfor the implemented contract. Inspect current CI and Greptile results before deciding to merge; preserve the documented resume restriction unless separately implementing and validating a replacement.Full CPU gate:
LIBREYOLO_PR_GATE=1 python -m pytest tests/unit -m "unit and not external_data and not network and not distributed" -n 4 --dist loadfile -q. Distributed gate: same command with-m "unit and not external_data and not network and distributed"and without xdist. Wheel smoke:python tests/smoke/run_install_smoke.py --mode wheel.Implementation session label:
implement_fitness_clean(Codex). This description is the portable handoff; no local transcript is required.Code provenance
Original implementation and tests written from LibreYOLO's own code and the user's requirements in a fresh agent context; no third-party code copied or adapted. An earlier investigation encountered AGPL snippets in search results and stopped before implementation. The maintainer authorized recovery through the fresh implementation context, which consulted no external source; the exposed investigator did not contribute implementation code.
The PR appears safe to merge based on the reviewed changes.
Summary
Adds a Python callback hook for custom validation fitness in shared training.
average_bestwhile retaining raw validation metrics.Diagram
%%{init: {'theme': 'neutral'}}%% flowchart LR V[Validation metrics] --> F{Fitness callback?} F -->|Yes| S[Score on rank zero] F -->|No| D[Use default task metric] S --> B[Best state and patience] D --> B B --> C[Checkpoint and epoch event] B --> A[average_best snapshot ranking]Reviews (1) · Last reviewed commit: "Keep point metadata for direct FOMO trai..."