Repository navigation
Allow choosing the early-stopping / best.pt metric (e.g. F1) for detection and classification training #852
Description
Activity
Already made a pull request following the instructions present in the repo. PR here: #853
Also addressed greptile comment which has since been auto resolved.Unsure how to assign myself to the Kanban entry though, can't see any options to pick up the open ticket linked to this issue.
Side note: Thanks for the awesome repo, appreciate the effort 👍
Hello @velu-97 ,
Thank you for the issue and the PR.
I agree custom early stopping is useful, but I don't think it's the best approach, for two reasons:
- The flagship use case doesn't hold. mAP is already macro-averaged over classes, so it isn't skewed by imbalance. The only F1 the library has is micro-averaged at IoU 0.50, which handles imbalance worse than the default.
- A per-task alias table is a lot of permanent surface. It has to be maintained for every task and metric, and it's inert for three of the five tasks.
Where you're right: the classification validator only has top-1/top-5, and that's a real gap. If you want to open a separate PR adding macro precision, recall and F1 from a confusion matrix in the classify validator, I'd be happy to review and merge that.
For custom stopping and best.pt selection, I have something else in mind...
Thanks again and very open to the separate PR adding macro precision, recall and F1 from a confusion matrix in the classify validator
Reacted by VeluHi @EHxuban11 !
For custom stopping and best.pt selection, I have something else in mind...
Noted 👍
PR for classifier changes here
The callback alternative is implemented and available in draft PR #900: #900
- Branch:
feat/custom-fitness, latest commit7b0f0295. Another agent can clone/fetch from GitHub and continue directly; no local transcript is needed. - Optional
fitness(metrics)callback controlsbest.pt, patience and checkpoint averaging. See the PR description anddocs/training_loggers.mdfor usage and the contract. - The PR includes the verification summary, baseline-reproduced failures, remaining limitations, prior review fixes and continuation commands. Implementation session label:
implement_fitness_clean(Codex). - Status: draft, not approved or merge-ready. GitHub CI and the latest pre-push Greptile review (
ad89a4bc-9b8b-4d98-b484-b32edcf1e5f6) are pending. Next agent should assess current review/check results before recommending merge.
- Branch:
Hello @velu-97, this is what I had in mind! Custom fitness is merged into
devin #900 and will ship in the next release (v1.6.0).You can pass a callback with a
fitness(metrics)method totrain(). Its return value (higher is better) selectsbest.ptand drivespatience. With your classifier metrics from #883, you can use F1 directly:class F1Fitness: def fitness(self, metrics): return metrics["metrics/f1"] model.train(data="data.yaml", patience=10, callbacks=F1Fitness())
Thanks again for the idea and for the classifier PR! Let us know how it goes 🙏
Reacted by VeluReacted by VeluHi @EHxuban11 ! Thanks for the support on my PR and merging it in and also for adding in the feature! Looking forward to using it on release 👍
Reacted by Xuban@velu-97 Perfect!, thanks for suggesting. v1.6.0 should be released this weekend if everything goes well!
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsDone
Goal
Let the user pick which validation metric drives early stopping and
best.ptselection (for example F1 on an imbalanced detection dataset) instead of the trainer's fixed choice pluspatience.Current behavior (dev @ 5d1d7f6)
best.ptshare one comparison inBaseTrainer(libreyolo/training/trainer.py:1968-1977,2307-2323,3812-3820). The only user-facing knob ispatience(libreyolo/training/config.py:268, CLI--patience).best_metric_key = "metrics/mAP50-95"(trainer.py:155). The detect validation branch reads that attribute (trainer.py:2929-2934), but nothing user-facing can set it._run_classify_validationhard-codes top-1 (trainer.py:3007-3016) and ignoresbest_metric_key. The semantic, depth and restore branches hard-code mIoU, delta1 and PSNR the same way.metrics/precisionandmetrics/recallare deprecated aliases of mAP50-95 and AR100 (libreyolo/validation/coco_evaluator.py:205-225), not a precision/recall pair. The only real F1 ismetrics/best_conf_f1, the maximum F1 over the confidence sweep at IoU 0.50 (detection_validator.py:1105-1145), which is already computed every epoch and can be NaN.libreyolo/validation/classify_validator.py:153-204). No precision, recall or F1.best_metric=is warned about and silently dropped byTrainConfig.from_kwargs(config.py:316-326).Proposal
best_metricon Python and CLI (bothkey=valueand--key value), sitting next topatience. DefaultNonekeeps each trainer's current key, so existing runs are unchanged.map50-95(default),map50,map75,f1. Classification:top1(default),top5,f1,precision,recall. Other tasks accept only their current default name.f1maps to the existingmetrics/best_conf_f1. NaN counts as no improvement.metrics/precision,metrics/recall,metrics/f1.best_metricthrough the configured key instead of hard-coding it.best_metric_keykeeps going into checkpoint metadata as today, so resume keeps refusing to inherit a best value across a key change.mode=minandmin_delta(loss-driven stopping). All selectable metrics are higher-is-better.Environment
Code-level finding on
devat 5d1d7f6 (1.5.0.dev0). No runtime needed to reproduce; the paths above are the evidence.