From f026a4c94bce0d29df048c1883f82340d63bbe1e Mon Sep 17 00:00:00 2001 From: Quentin Geissmann Date: Thu, 3 Sep 2026 07:18:19 +0200 Subject: [PATCH 1/3] Fix the inpainting halo, which marked every kept instance in train AND val telea_inpaint_polys dilated the union of the polygons to inpaint and the polygons to exclude, then subtracted only the undilated exclusions. The dilation ring around each EXCLUDED polygon therefore stayed in the mask and was overwritten with inpainted content, painting a band roughly `downscale_factor` px wide that traces the outline of every instance that survived the crop. FixInstances calls this from the VALIDATION pipeline as well as the training one, so the band sat on both sides of the train/val split. A ring hugging every labelled instance is a cue a model can learn in place of the animal and then be rewarded for at validation time, so metrics computed before this fix were measured against crops carrying an outline of their own ground truth. Measured on real crops: 0.2-1.1% of each crop altered, concentrated entirely on instance boundaries. Nothing on disk was affected - the function mutates the in-memory crop, so the prepared dataset is unchanged and only the consumed crops carried it. Fix: subtract a dilated exclude mask, using the same kernel and iterations that created the ring. Regression check goes from 43.9% to 0.0% of pixels altered in a 25px band outside kept instances, while the instances being removed are still fully inpainted. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01SyKzVebNKNYj82VZWTZU6Y --- src/flat_bug/augmentations.py | 26 ++++++++++++++++++++++---- 1 file changed, 22 insertions(+), 4 deletions(-) diff --git a/src/flat_bug/augmentations.py b/src/flat_bug/augmentations.py index 38d125f..41bdf4a 100644 --- a/src/flat_bug/augmentations.py +++ b/src/flat_bug/augmentations.py @@ -163,14 +163,32 @@ def telea_inpaint_polys( flags=cv2.INPAINT_TELEA ) - # Remove the exclude polygons from the inpaint bitmap, so that the original image is not inpainted under them + # Remove the exclude polygons from the inpaint bitmap, so that the original image is not + # inpainted under them. This must undo the DILATION as well as the polygon: the dilate + # above grew every contour drawn, excluded ones included, so subtracting only the bare + # polygon left a dilated ring - 1 low-res px, about `downscale_factor` px at full + # resolution - still marked for inpainting, painting a smeared halo tight around every + # KEPT instance. + # + # This runs in the validation pipeline as well as the training one, so the marker sat on + # both sides of the train/val split: a ring hugging every labelled instance is a cue a + # model can learn instead of the animal, and be rewarded for at validation time. Every + # metric computed before this fix was measured against crops carrying it. + exclude_bitmap = np.zeros_like(inpaint_bitmap) for p in exclude_polys: - inpaint_bitmap = cv2.drawContours( - inpaint_bitmap, + exclude_bitmap = cv2.drawContours( + exclude_bitmap, [p // downscale_factor], - color=0, + color=1, **kwargs ) + cv2.dilate( + src=exclude_bitmap, + dst=exclude_bitmap, + kernel=np.ones((3, 3), np.uint8), + iterations=1 + ) + inpaint_bitmap[exclude_bitmap == 1] = 0 # Upsample the inpainted image and bitmap inpaint_bitmap = cv2.resize(inpaint_bitmap, orig_shape) From 055b96e077cdcf19a8fe9f581f4c28563d764bef Mon Sep 17 00:00:00 2001 From: Quentin Geissmann Date: Thu, 3 Sep 2026 07:18:44 +0200 Subject: [PATCH 2/3] Add the 50-epoch A/B that measures the inpainting halo One config, two source trees: ARM=before runs develop with the halo present, ARM=after runs the patched tree. Everything else is held fixed, including the pyramid bug, so the pair attributes any difference to the halo alone. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01SyKzVebNKNYj82VZWTZU6Y --- scripts/training/fb_config_halo50_GHPC.yaml | 19 +++++++++++ scripts/training/train_halo50_AU_GHPC.sh | 35 +++++++++++++++++++++ 2 files changed, 54 insertions(+) create mode 100644 scripts/training/fb_config_halo50_GHPC.yaml create mode 100755 scripts/training/train_halo50_AU_GHPC.sh diff --git a/scripts/training/fb_config_halo50_GHPC.yaml b/scripts/training/fb_config_halo50_GHPC.yaml new file mode 100644 index 0000000..6c68675 --- /dev/null +++ b/scripts/training/fb_config_halo50_GHPC.yaml @@ -0,0 +1,19 @@ +# 50-epoch A/B for the inpainting-halo fix. Two runs share this config exactly; the only +# difference is whether the code they run against carries the patch. +# +# Stock settings on purpose - no synthetic scenes, no magnified crops, no thin weighting, +# default mask_ratio, and the pyramid left as it is - so the pair isolates the halo and +# nothing else. Mirrors fb_config_control_GHPC.yaml apart from the epoch count. + +batch: 8 +model: "yolo26m-seg.pt" +epochs: 50 +device: 0 +patience: 9999 +lr0: 0.01 +lrf: 0.00001 +optimizer: 'SGD' +workers: 16 +fb_max_instances: 9999 +fb_max_images: -1 +fb_synth_prob: 0.0 diff --git a/scripts/training/train_halo50_AU_GHPC.sh b/scripts/training/train_halo50_AU_GHPC.sh new file mode 100755 index 0000000..2a5b320 --- /dev/null +++ b/scripts/training/train_halo50_AU_GHPC.sh @@ -0,0 +1,35 @@ +#!/bin/bash +# 50-epoch A/B for the inpainting-halo fix. +# +# sbatch --export=ALL,ARM=before train_halo50_AU_GHPC.sh # develop, halo present +# sbatch --export=ALL,ARM=after train_halo50_AU_GHPC.sh # patched, halo removed +# +# Both arms read the same prepared data and the same config; ARM only selects which source +# tree is on PYTHONPATH. The pyramid bug is deliberately left in place in both. + +#SBATCH -p ghpc_gpu +#SBATCH -N 1 +#SBATCH -n 16 +#SBATCH --mem=96000 +#SBATCH -t 48:00:00 +#SBATCH --gres=gpu:1 +#SBATCH -J fb_halo +#SBATCH -o /usr/home/qgg/qgeiss/flatbug-dir/logs/fb_halo_%x_%j.out +#SBATCH -e /usr/home/qgg/qgeiss/flatbug-dir/logs/fb_halo_%x_%j.err + +source ~/.venv/bin/activate +export PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True + +case "$ARM" in + before) SRC=$HOME/flat-bug-halo-before ;; + after) SRC=$HOME/flat-bug-halo-after ;; + *) echo "set ARM=before or ARM=after"; exit 1 ;; +esac +export PYTHONPATH=$SRC/src + +ROOT=$HOME/flatbug-dir +CONFIG=$SRC/scripts/training/fb_config_halo50_GHPC.yaml +NAME=fb_halo_${ARM}_$(date +"%Y-%m-%d_%H-%M-%S") + +cd $SRC/scripts/training +fb_train -c ${CONFIG} -d ${ROOT}/flat-bug-data/yolo/ --name ${NAME} From d41b184ada205a55db69a8515c26fbb5646d0dc0 Mon Sep 17 00:00:00 2001 From: Quentin Geissmann Date: Thu, 3 Sep 2026 21:08:22 +0200 Subject: [PATCH 3/3] Record the halo defect and its A/B result in KNOWN_ISSUES.md The controlled 50-epoch pair (1724716/1724717, one file different) settles what the cropped benchmark could not: the fix looks like an 8-point mAP regression on validation crops and is a decisive improvement in real inference - end-to-end F1 0.849 -> 0.878, recall 0.850 -> 0.895, touching-instance recall 0.749 -> 0.819, 1,244 more ground-truth instances matched. The cropped benchmark carried an outline of its own ground truth, so it was measuring the artefact. Also documents the pyramid native-scale bug, which is deliberately NOT fixed on this branch so the halo change stays attributable on its own. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01SyKzVebNKNYj82VZWTZU6Y --- KNOWN_ISSUES.md | 89 +++++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 89 insertions(+) create mode 100644 KNOWN_ISSUES.md diff --git a/KNOWN_ISSUES.md b/KNOWN_ISSUES.md new file mode 100644 index 0000000..0df8989 --- /dev/null +++ b/KNOWN_ISSUES.md @@ -0,0 +1,89 @@ +# Known issues in the flat-bug stack + +Defects found in `src/flat_bug/` that affect the main pipeline, with how they were measured +so a fix can be verified rather than assumed. Prototype-only problems belong in the +prototype's own docstrings. + +--- + +## 1. Inpainting drew an outline of the ground truth into every crop + +**Status:** fixed on `fix/inpaint-halo` (`telea_inpaint_polys`, 2026-09-03). Every model and +every metric produced before this date is affected. + +**What happened.** The inpaint mask is built from the polygons to erase *and* the polygons to +keep — the latter deliberately, so `cv2.inpaint` cannot rebuild a deleted animal out of its +neighbour's body. The mask is then dilated so a removed instance's edge cannot bleed into the +fill. Both steps are correct. The defect is that only the *undilated* excluded polygons were +subtracted afterwards, leaving the dilation ring — about `downscale_factor` px at full +resolution — still marked for inpainting. Every instance that SURVIVED the crop therefore +acquired a smeared band tracing its own outline. + +**Why it mattered.** `FixInstances` runs in the validation pipeline as well as the training +one, so the marker sat on both sides of the train/val split. A ring hugging every labelled +instance is a cue a model can learn instead of the animal, and then be rewarded for at +validation time. + +**Measured.** 0.2-1.1% of each crop altered, concentrated on instance boundaries: on a +regenerated 126 Mpx tile, 12.06% of pixels differed inside a 25px ring outside kept instances +against 0.008% elsewhere, on a 0.000% JPEG re-encode noise floor. Nothing on disk was ever +affected - the function mutates the in-memory crop. + +**Fix.** Track the exclusions in their own mask, dilate it with the same kernel, and subtract +that. A 25px band outside kept instances goes from 43.9% of pixels altered to 0.0%, while +instances being removed stay 100% inpainted. + +**Verified by a controlled A/B** (jobs 1724716 / 1724717, yolo26m-seg, 50 epochs, stock +settings, source trees differing in one file): + +| measured on | before (halo) | after (fix) | +|-------------------------------|---------------|-------------| +| box mAP50-95, val CROPS | 0.859 | 0.781 | +| mask mAP50-95, val CROPS | 0.750 | 0.655 | +| F1, end-to-end whole images | 0.849 | **0.878** | +| recall, end-to-end | 0.850 | **0.895** | +| touching-instance recall | 0.749 | **0.819** | +| merge rate (lower better) | 0.156 | **0.142** | +| GT instances matched | 23,440 | **24,684** | + +The fix looks like an 8-point mAP regression on the cropped benchmark and is a decisive +improvement in real inference, because the cropped benchmark carried the artefact. F1 rises +on 26 of 31 sub-datasets; the largest gains are on the weakest ones (broto2025 +0.187, +BugNet +0.099, AMI-traps +0.081). PeMaToEuroPep (-0.024) and Diopsis (-0.014) regress on +precision. + +**Consequences.** No checkpoint trained before this is a valid baseline, and no metric +computed before it is comparable with one computed after. + +--- + +## 2. The scale pyramid silently skipped native resolution + +**Status:** NOT fixed on this branch, deliberately - it is an inference-time change and was +kept out so the halo fix could be attributed alone. Fixed on +`fix/inpaint-halo-and-pyramid-scale`. + +**What happened.** `pyramid_predictions` decides whether to add native scale by testing the +ladder's exit value rather than the ladder itself: + +```python +while s <= 0.9: + scales.append(s) + s /= scale_increment # s grows by 1.5x +if s != 1: # tests `s`, not `scales` + scales.append(1.0) +``` + +The loop can only append values <= 0.9, so 1.0 is never already present and should always be +added. Where `max_dim == TILE_SIZE * 1.5**k` the exit value lands exactly on 1.0 and full +resolution was never run - **1024, 1536, 2304, 3456 and 5184 px** at the default 1024 tile, +and 5184x3456 is a stock DSLR frame. + +**Consequence.** The pyramid is the only way to find objects larger than a tile and helps in +every size bucket (recall 0.864 vs 0.847 single-scale). For affected sizes the finest level +was missing, so those benchmarks are pessimistic by an unknown amount. + +**Fix.** `if 1.0 not in scales: scales.append(1.0)`. Verified across image sizes: 1536, 2304, +3456 and 5184 px go from missing native scale to including it, all others unchanged. + +**History.** Already fixed in the M2F prototype's predictor but never ported back.