diff --git a/KNOWN_ISSUES.md b/KNOWN_ISSUES.md new file mode 100644 index 0000000..0df8989 --- /dev/null +++ b/KNOWN_ISSUES.md @@ -0,0 +1,89 @@ +# Known issues in the flat-bug stack + +Defects found in `src/flat_bug/` that affect the main pipeline, with how they were measured +so a fix can be verified rather than assumed. Prototype-only problems belong in the +prototype's own docstrings. + +--- + +## 1. Inpainting drew an outline of the ground truth into every crop + +**Status:** fixed on `fix/inpaint-halo` (`telea_inpaint_polys`, 2026-09-03). Every model and +every metric produced before this date is affected. + +**What happened.** The inpaint mask is built from the polygons to erase *and* the polygons to +keep — the latter deliberately, so `cv2.inpaint` cannot rebuild a deleted animal out of its +neighbour's body. The mask is then dilated so a removed instance's edge cannot bleed into the +fill. Both steps are correct. The defect is that only the *undilated* excluded polygons were +subtracted afterwards, leaving the dilation ring — about `downscale_factor` px at full +resolution — still marked for inpainting. Every instance that SURVIVED the crop therefore +acquired a smeared band tracing its own outline. + +**Why it mattered.** `FixInstances` runs in the validation pipeline as well as the training +one, so the marker sat on both sides of the train/val split. A ring hugging every labelled +instance is a cue a model can learn instead of the animal, and then be rewarded for at +validation time. + +**Measured.** 0.2-1.1% of each crop altered, concentrated on instance boundaries: on a +regenerated 126 Mpx tile, 12.06% of pixels differed inside a 25px ring outside kept instances +against 0.008% elsewhere, on a 0.000% JPEG re-encode noise floor. Nothing on disk was ever +affected - the function mutates the in-memory crop. + +**Fix.** Track the exclusions in their own mask, dilate it with the same kernel, and subtract +that. A 25px band outside kept instances goes from 43.9% of pixels altered to 0.0%, while +instances being removed stay 100% inpainted. + +**Verified by a controlled A/B** (jobs 1724716 / 1724717, yolo26m-seg, 50 epochs, stock +settings, source trees differing in one file): + +| measured on | before (halo) | after (fix) | +|-------------------------------|---------------|-------------| +| box mAP50-95, val CROPS | 0.859 | 0.781 | +| mask mAP50-95, val CROPS | 0.750 | 0.655 | +| F1, end-to-end whole images | 0.849 | **0.878** | +| recall, end-to-end | 0.850 | **0.895** | +| touching-instance recall | 0.749 | **0.819** | +| merge rate (lower better) | 0.156 | **0.142** | +| GT instances matched | 23,440 | **24,684** | + +The fix looks like an 8-point mAP regression on the cropped benchmark and is a decisive +improvement in real inference, because the cropped benchmark carried the artefact. F1 rises +on 26 of 31 sub-datasets; the largest gains are on the weakest ones (broto2025 +0.187, +BugNet +0.099, AMI-traps +0.081). PeMaToEuroPep (-0.024) and Diopsis (-0.014) regress on +precision. + +**Consequences.** No checkpoint trained before this is a valid baseline, and no metric +computed before it is comparable with one computed after. + +--- + +## 2. The scale pyramid silently skipped native resolution + +**Status:** NOT fixed on this branch, deliberately - it is an inference-time change and was +kept out so the halo fix could be attributed alone. Fixed on +`fix/inpaint-halo-and-pyramid-scale`. + +**What happened.** `pyramid_predictions` decides whether to add native scale by testing the +ladder's exit value rather than the ladder itself: + +```python +while s <= 0.9: + scales.append(s) + s /= scale_increment # s grows by 1.5x +if s != 1: # tests `s`, not `scales` + scales.append(1.0) +``` + +The loop can only append values <= 0.9, so 1.0 is never already present and should always be +added. Where `max_dim == TILE_SIZE * 1.5**k` the exit value lands exactly on 1.0 and full +resolution was never run - **1024, 1536, 2304, 3456 and 5184 px** at the default 1024 tile, +and 5184x3456 is a stock DSLR frame. + +**Consequence.** The pyramid is the only way to find objects larger than a tile and helps in +every size bucket (recall 0.864 vs 0.847 single-scale). For affected sizes the finest level +was missing, so those benchmarks are pessimistic by an unknown amount. + +**Fix.** `if 1.0 not in scales: scales.append(1.0)`. Verified across image sizes: 1536, 2304, +3456 and 5184 px go from missing native scale to including it, all others unchanged. + +**History.** Already fixed in the M2F prototype's predictor but never ported back. diff --git a/scripts/training/fb_config_halo50_GHPC.yaml b/scripts/training/fb_config_halo50_GHPC.yaml new file mode 100644 index 0000000..6c68675 --- /dev/null +++ b/scripts/training/fb_config_halo50_GHPC.yaml @@ -0,0 +1,19 @@ +# 50-epoch A/B for the inpainting-halo fix. Two runs share this config exactly; the only +# difference is whether the code they run against carries the patch. +# +# Stock settings on purpose - no synthetic scenes, no magnified crops, no thin weighting, +# default mask_ratio, and the pyramid left as it is - so the pair isolates the halo and +# nothing else. Mirrors fb_config_control_GHPC.yaml apart from the epoch count. + +batch: 8 +model: "yolo26m-seg.pt" +epochs: 50 +device: 0 +patience: 9999 +lr0: 0.01 +lrf: 0.00001 +optimizer: 'SGD' +workers: 16 +fb_max_instances: 9999 +fb_max_images: -1 +fb_synth_prob: 0.0 diff --git a/scripts/training/train_halo50_AU_GHPC.sh b/scripts/training/train_halo50_AU_GHPC.sh new file mode 100755 index 0000000..2a5b320 --- /dev/null +++ b/scripts/training/train_halo50_AU_GHPC.sh @@ -0,0 +1,35 @@ +#!/bin/bash +# 50-epoch A/B for the inpainting-halo fix. +# +# sbatch --export=ALL,ARM=before train_halo50_AU_GHPC.sh # develop, halo present +# sbatch --export=ALL,ARM=after train_halo50_AU_GHPC.sh # patched, halo removed +# +# Both arms read the same prepared data and the same config; ARM only selects which source +# tree is on PYTHONPATH. The pyramid bug is deliberately left in place in both. + +#SBATCH -p ghpc_gpu +#SBATCH -N 1 +#SBATCH -n 16 +#SBATCH --mem=96000 +#SBATCH -t 48:00:00 +#SBATCH --gres=gpu:1 +#SBATCH -J fb_halo +#SBATCH -o /usr/home/qgg/qgeiss/flatbug-dir/logs/fb_halo_%x_%j.out +#SBATCH -e /usr/home/qgg/qgeiss/flatbug-dir/logs/fb_halo_%x_%j.err + +source ~/.venv/bin/activate +export PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True + +case "$ARM" in + before) SRC=$HOME/flat-bug-halo-before ;; + after) SRC=$HOME/flat-bug-halo-after ;; + *) echo "set ARM=before or ARM=after"; exit 1 ;; +esac +export PYTHONPATH=$SRC/src + +ROOT=$HOME/flatbug-dir +CONFIG=$SRC/scripts/training/fb_config_halo50_GHPC.yaml +NAME=fb_halo_${ARM}_$(date +"%Y-%m-%d_%H-%M-%S") + +cd $SRC/scripts/training +fb_train -c ${CONFIG} -d ${ROOT}/flat-bug-data/yolo/ --name ${NAME} diff --git a/src/flat_bug/augmentations.py b/src/flat_bug/augmentations.py index 38d125f..41bdf4a 100644 --- a/src/flat_bug/augmentations.py +++ b/src/flat_bug/augmentations.py @@ -163,14 +163,32 @@ def telea_inpaint_polys( flags=cv2.INPAINT_TELEA ) - # Remove the exclude polygons from the inpaint bitmap, so that the original image is not inpainted under them + # Remove the exclude polygons from the inpaint bitmap, so that the original image is not + # inpainted under them. This must undo the DILATION as well as the polygon: the dilate + # above grew every contour drawn, excluded ones included, so subtracting only the bare + # polygon left a dilated ring - 1 low-res px, about `downscale_factor` px at full + # resolution - still marked for inpainting, painting a smeared halo tight around every + # KEPT instance. + # + # This runs in the validation pipeline as well as the training one, so the marker sat on + # both sides of the train/val split: a ring hugging every labelled instance is a cue a + # model can learn instead of the animal, and be rewarded for at validation time. Every + # metric computed before this fix was measured against crops carrying it. + exclude_bitmap = np.zeros_like(inpaint_bitmap) for p in exclude_polys: - inpaint_bitmap = cv2.drawContours( - inpaint_bitmap, + exclude_bitmap = cv2.drawContours( + exclude_bitmap, [p // downscale_factor], - color=0, + color=1, **kwargs ) + cv2.dilate( + src=exclude_bitmap, + dst=exclude_bitmap, + kernel=np.ones((3, 3), np.uint8), + iterations=1 + ) + inpaint_bitmap[exclude_bitmap == 1] = 0 # Upsample the inpainted image and bitmap inpaint_bitmap = cv2.resize(inpaint_bitmap, orig_shape)