Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 14 additions & 1 deletion .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -204,4 +204,17 @@ __marimo__/

# Generated HF model cards — build output of hf_modelcards/generate_hf_model_zoo.py;
# source of truth is hf_modelcards/template/ + hf_modelcards/model_metrics/
hf_modelcards/generated/
hf_modelcards/

# PyTorch model checkpoints
*.ckpt
*.pt
*.pth
*.ptl

# Experiments
experiments/
cabinet_github_files/
node_modules/
outputs/
app.py
81 changes: 69 additions & 12 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -61,6 +61,23 @@ All numbers below are test-split mIoU (not val — see [`train_yolo.py`](src/scr

CABiNet (both backbones) outperforms **every** YOLO26-sem variant on UAVid, including the largest (YOLO26x) — and does so at a fraction of the compute: CABiNet-Large (54.8 GFLOPs) beats YOLO26x (430.9 GFLOPs, ~8× more) and YOLO26m (152.3 GFLOPs, ~3× more) on both mIoU and FLOPs simultaneously. CABiNet-Small also improves substantially over the numbers originally reported in the [CABiNet paper](#citation) on this dataset.

## VDD Model Zoo

CABiNet and YOLO26-sem are also trained and evaluated under one shared [VDD (Varied Drone Dataset)](https://github.com/RussRobin/VDD) pipeline (`images/`+`masks/` format, 7 classes). VDD is a much smaller dataset than UAVid (280 train images) spanning varied altitudes and viewpoints; heavier augmentation (mosaic/mixup/copy-paste) is used to help offset the small training set — see the [VDD Dataset](#vdd-dataset) section below.

All numbers below are test-split mIoU, same methodology as the [UAVid Model Zoo](#uavid-model-zoo) table above (Params/FLOPs are architecture-only and identical across datasets, since both are measured at 1024×1024 regardless of what the model was finetuned on).

| Model | mIoU (%) | Params (M) | FLOPs (GFLOPs) | HF Weights |
| --------------------------- | -------- | ---------- | -------------- | --------------------------------------------------------------------------- |
| YOLO26x-sem | 78.83 | 40.16 | 430.9 | [HF Model](https://huggingface.co/dronefreak/vdd-yolo26x-sem) |
| YOLO26l-sem | 78.57 | 17.87 | 192.4 | [HF Model](https://huggingface.co/dronefreak/vdd-yolo26l-sem) |
| CABiNet (MobileNetV3-Large) | 77.76 | 9.17 | 54.8 | [HF Model](https://huggingface.co/dronefreak/cabinet-mobilenetv3-large-vdd) |
| YOLO26m-sem | 77.02 | 14.32 | 152.3 | [HF Model](https://huggingface.co/dronefreak/vdd-yolo26m-sem) |
| YOLO26s-sem | 76.35 | 6.50 | 44.4 | [HF Model](https://huggingface.co/dronefreak/vdd-yolo26s-sem) |
| YOLO26n-sem | 73.99 | 1.63 | 11.4 | [HF Model](https://huggingface.co/dronefreak/vdd-yolo26n-sem) |

Unlike UAVid, the two largest YOLO26-sem variants (x, l) edge out CABiNet-Large on raw mIoU here, by a small margin (≤1.1 pts) — VDD's much smaller training set and different 7-class scheme make this a genuine dataset effect, not a regression. CABiNet-Large still beats YOLO26m on mIoU while using ~64% less compute (54.8 vs 152.3 GFLOPs), and outperforms YOLO26s/YOLO26n outright on both mIoU and FLOPs.

## Installation

### Prerequisites
Expand Down Expand Up @@ -470,11 +487,18 @@ All 8 UAVid classes are valid and active — none are mapped to the ignore label
> | `mosaic` / `mixup` / `copy_paste` | Aerial-tuned augmentation for small-object diversity (cars, humans) |
> | EMA | Always on — `best.pt` / `last.pt` are EMA-averaged; no separate flag needed |

5. **Evaluate**:
5. **Evaluate** (val split by default; UAVid also has a real test split) — use the Hydra wrapper, not
the raw `yolo semantic val cfg=...` CLI (the `semantic` task isn't reliably invokable that way;
`train_yolo.py` drives it correctly via the Ultralytics Python API):

```bash
# val split
python src/scripts/train_yolo.py mode=val
python src/scripts/train_yolo.py mode=val validation_config.weights=path/to/best.pt

# test split
python src/scripts/train_yolo.py mode=val validation_config.split=test \
validation_config.weights=path/to/best.pt
```

6. **Inference & showcase mosaic**:
Expand Down Expand Up @@ -552,18 +576,32 @@ All 12 AeroScapes classes are valid and active — none are mapped to the ignore
yolo semantic train cfg=configs/yolo/aeroscapes_train.yaml
```

or the repo's Hydra wrapper:
or the repo's Hydra wrapper, which also takes the same dotted-path overrides as UAVid's:

```bash
# Default (yolo26n-sem):
python src/scripts/train_yolo.py --config-name train_yolo_aeroscapes

# Swap size — only override the model group:
python src/scripts/train_yolo.py --config-name train_yolo_aeroscapes 'yolo/model@model=yolo26s-sem'

# Override hyperparameters:
python src/scripts/train_yolo.py --config-name train_yolo_aeroscapes \
training_config.epochs=200 training_config.batch_size=8

# Resume an interrupted run:
python src/scripts/train_yolo.py --config-name train_yolo_aeroscapes training_config.resume=true

# Multi-GPU (DDP):
python src/scripts/train_yolo.py --config-name train_yolo_aeroscapes runtime.device="0,1"
```

4. **Evaluate**:
4. **Evaluate** (val split only — AeroScapes has no test split, see above) — use the Hydra wrapper, not
the raw `yolo semantic val cfg=...` CLI (the `semantic` task isn't reliably invokable that way):

```bash
yolo semantic val cfg=configs/yolo/aeroscapes_val.yaml model=runs/aeroscapes/yolo26n/weights/best.pt
# or
python src/scripts/train_yolo.py --config-name train_yolo_aeroscapes mode=val
python src/scripts/train_yolo.py --config-name train_yolo_aeroscapes mode=val \
validation_config.weights=runs/aeroscapes/yolo26n/weights/best.pt
```

### VDD → YOLO Format
Expand Down Expand Up @@ -623,18 +661,37 @@ All 7 VDD classes are valid and active — none are mapped to the ignore label:
yolo semantic train cfg=configs/yolo/vdd_train.yaml
```

or the repo's Hydra wrapper:
or the repo's Hydra wrapper, which also takes the same dotted-path overrides as UAVid's:

```bash
# Default (yolo26n-sem):
python src/scripts/train_yolo.py --config-name train_yolo_vdd

# Swap size — only override the model group:
python src/scripts/train_yolo.py --config-name train_yolo_vdd 'yolo/model@model=yolo26s-sem'

# Override hyperparameters:
python src/scripts/train_yolo.py --config-name train_yolo_vdd \
training_config.epochs=200 training_config.batch_size=8

# Resume an interrupted run:
python src/scripts/train_yolo.py --config-name train_yolo_vdd training_config.resume=true

# Multi-GPU (DDP):
python src/scripts/train_yolo.py --config-name train_yolo_vdd runtime.device="0,1"
```

4. **Evaluate**:
4. **Evaluate** (val split by default; VDD also has a real test split) — use the Hydra wrapper, not
the raw `yolo semantic val cfg=...` CLI (the `semantic` task isn't reliably invokable that way):

```bash
yolo semantic val cfg=configs/yolo/vdd_val.yaml model=runs/vdd/yolo26n/weights/best.pt
# or
python src/scripts/train_yolo.py --config-name train_yolo_vdd mode=val
# val split
python src/scripts/train_yolo.py --config-name train_yolo_vdd mode=val \
validation_config.weights=runs/vdd/yolo26n/weights/best.pt

# test split
python src/scripts/train_yolo.py --config-name train_yolo_vdd mode=val \
validation_config.split=test validation_config.weights=runs/vdd/yolo26n/weights/best.pt
```

### Evaluation
Expand Down Expand Up @@ -704,7 +761,7 @@ print(f"Peak Memory: {results['memory']['peak_mb']:.2f} MB")

Pretrained MobileNetV3 backbone weights: [`src/models/pretrained_backbones/`](src/models/pretrained_backbones/).

Full CABiNet + YOLO26-sem checkpoints are being published to a Hugging Face UAVid model zoo — model cards are generated from [`hf_modelcards/`](hf_modelcards/) (one Jinja template + per-model `metrics.json`; run `python hf_modelcards/generate_hf_model_zoo.py` to regenerate all cards). See [UAVid Model Zoo](#uavid-model-zoo) above for current numbers.
Full CABiNet + YOLO26-sem checkpoints are being published to Hugging Face UAVid and VDD model zoos — model cards are generated from [`hf_modelcards/`](hf_modelcards/) (one Jinja template shared across datasets + per-model `metrics.json`; run `python hf_modelcards/generate_hf_model_zoo.py` to regenerate all cards). See [UAVid Model Zoo](#uavid-model-zoo) / [VDD Model Zoo](#vdd-model-zoo) above for current numbers.

## Testing

Expand Down
72 changes: 72 additions & 0 deletions configs/VDD_info.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,72 @@
[
{
"hasInstances": false,
"category": "void",
"catid": 0,
"name": "Other",
"ignoreInEval": false,
"id": 0,
"color": [0, 0, 0],
"trainId": 0
},
{
"hasInstances": false,
"category": "construction",
"catid": 1,
"name": "Wall",
"ignoreInEval": false,
"id": 1,
"color": [128, 64, 0],
"trainId": 1
},
{
"hasInstances": false,
"category": "flat",
"catid": 2,
"name": "Road",
"ignoreInEval": false,
"id": 2,
"color": [128, 64, 128],
"trainId": 2
},
{
"hasInstances": false,
"category": "vegetation",
"catid": 3,
"name": "Vegetation",
"ignoreInEval": false,
"id": 3,
"color": [0, 128, 0],
"trainId": 3
},
{
"hasInstances": false,
"category": "vehicle",
"catid": 4,
"name": "Vehicle",
"ignoreInEval": false,
"id": 4,
"color": [64, 0, 128],
"trainId": 4
},
{
"hasInstances": false,
"category": "construction",
"catid": 1,
"name": "Roof",
"ignoreInEval": false,
"id": 5,
"color": [192, 0, 0],
"trainId": 5
},
{
"hasInstances": false,
"category": "water",
"catid": 5,
"name": "Water",
"ignoreInEval": false,
"id": 6,
"color": [0, 128, 192],
"trainId": 6
}
]
2 changes: 1 addition & 1 deletion configs/train.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -56,7 +56,7 @@ training_config:
ema_tau: 2000 # decay ramp time constant, in optimizer steps

# Early stopping on mIoU (evaluated on EMA weights). 0 = disabled.
patience: 30
patience: 100

# Validation Configuration
validation_config:
Expand Down
Loading