This repository extends YOLOv5 to jointly perform object detection and semantic segmentation of drivable area on the BDD100K dataset. It adds a lightweight segmentation head to the YOLOv5s model, a custom dataset/dataloader that returns both detection and segmentation targets, and training/validation changes to compute and log segmentation loss alongside standard detection losses.
- Build an end-to-end pipeline to train on BDD100K with both detection (10 classes) and a drivable-area segmentation task.
- Keep model changes minimal and compatible with YOLOv5 style and utilities.
- Provide reproducible data preprocessing, a clear dataset layout, and simple commands to train and validate.
-
New segmentation head in YOLOv5s
- What: Introduced
SemanticSegmentationOutinmodels/common.pyand wired a 2‑channel head (background, drivable) intomodels/yolov5s.yaml. The forward inmodels/yolo.pyreturns both detection outputs and the segmentation map. - Why: To jointly learn drivable-area segmentation using a light head that reuses backbone features, enabling multi-task learning with minimal overhead.
- What: Introduced
-
BDD-specific dataset and dataloader
- What: Implemented
BDDDatasetinBDD.pyto read BDD images, YOLO-format detection labels, and segmentation masks; applies letterbox/augmentations consistently to images, boxes, and masks; returns both detection targets and a 2‑channel segmentation target. Addedcreate_bdd_dataloader(...)inutils/dataloaders.pyplus a customcollate_fn. - Why: To guarantee alignment between images, boxes, and masks under resizing and augmentation while keeping the API similar to YOLOv5’s.
- What: Implemented
-
Training/validation integration
- What:
train.pyandval.pynow use the BDD dataloader and compute a segmentation loss usingBCEWithLogitsLossin addition to box/objectness/class losses. Segmentation loss is logged and aggregated like other losses. - Why: Seamless multi-task training using familiar YOLOv5 training scripts and logging.
- What:
-
Data preparation utilities
- What:
bdd_preprocessing.ipynbconverts original BDD annotations to YOLO detection labels and generates/fixes segmentation masks (including polygon patching).voxel.pyprovides quick dataset visualization. - Why: Reproducible preprocessing to build a clean, aligned dataset for training/validation.
- What:
-
BDD.py— DefinesBDDDataset:- Builds an index of images with matching detection
.txtand segmentation mask files. - Loads and letterboxes image and mask to a consistent size, applies augmentations (HSV, random perspective, flips), and converts detection labels between
xyxyand normalizedxywhas required. - Produces a 2‑channel segmentation target
[background, drivable]and returns(image, [det_labels, seg_mask], path, shapes). - Provides
collate_fnto batch both detection and segmentation targets.
- Builds an index of images with matching detection
-
utils/dataloaders.py— Addscreate_bdd_dataloader(...):- Mirrors the YOLOv5
create_dataloaderlogic but instantiatesBDDDatasetand uses itscollate_fn.
- Mirrors the YOLOv5
-
models/common.py— AddsSemanticSegmentationOut:- A small conv + BN + activation head that outputs a 2‑channel segmentation map.
-
models/yolov5s.yaml— Modifies the head:- Adds upsampling and C3 layers from P3 features into a final
SemanticSegmentationOut(2, ...)layer that predicts[background, drivable].
- Adds upsampling and C3 layers from P3 features into a final
-
models/yolo.py— Model parsing/forward updates:- Recognizes
SemanticSegmentationOutinparse_modeland returns(detections, seg_out)fromBaseModel._forward_once.
- Recognizes
-
train.py— Training loop updates:- Uses
create_bdd_dataloaderfor train/val. - Computes segmentation loss with
BCEWithLogitsLossand logs it alongside detection losses.
- Uses
-
val.py— Validation loop updates:- Validates with the BDD dataloader, computes segmentation loss for reporting, and uses the usual NMS/metrics for detection.
-
data/bdd.yaml— Dataset config:- Defines 10 detection classes and the dataset root (default
BDDDataset/).
- Defines 10 detection classes and the dataset root (default
-
bdd_preprocessing.ipynb— Preprocessing steps:- Converts detections to YOLO format, generates masks, and patches polygons for better segmentation masks.
Place your dataset as follows (or update data/bdd.yaml:path):
BDDDataset/
images/
train/ ...
val/ ...
test/ ...
labels/
det/
train/ ... # .txt in YOLO format
val/ ...
test/ ...
seg/
train/ ... # segmentation masks (image files)
val/ ...
test/ ...
python -m venv .venv && source .venv/bin/activate # or use your preferred env
pip install -r requirements.txtpython train.py \
--weights yolov5s.pt \
--cfg models/yolov5s.yaml \
--data data/bdd.yaml \
--epochs 100 \
--batch-size 16 \
--imgsz 640Tips:
- Ensure
data/bdd.yaml:pathpoints to yourBDDDataset/root. train.pydefaults are already set for this project; you can omit flags if you accept defaults.
python val.py --weights runs/train/exp29/weights/best.pt --data data/bdd.yaml --imgsz 640The following images come from recent training runs (runs/train/exp29). They illustrate batch visuals and label statistics during training.
Note: The training visualizations primarily show detection labels/predictions. The segmentation head produces a drivable-area map (2 channels). If desired, you can add utilities to overlay the predicted mask on images for richer qualitative results.
- Minimal, modular change: Rather than forking into a large seg architecture, the project keeps YOLOv5 intact and adds a small head (
SemanticSegmentationOut) that can be parsed from YAML. This preserves most of the ecosystem while enabling segmentation. - Aligned augmentations: All image transforms (resize, letterbox, perspective, flips) are applied consistently to both the image and the segmentation mask to keep labels aligned.
- Loss choice: A simple
BCEWithLogitsLossis used for segmentation for its stability and speed on binary masks. It can be swapped for Dice, Focal, or Combo losses if you emphasize mask quality. - Data efficiency: Reuses the strong YOLOv5 backbone + head structure for feature reuse and efficient joint training.
Oldest → newest highlights of changes on master:
- 2edfa4cc: Preprocess data, convert detections to YOLO, create segmentation masks (
bdd_preprocessing.ipynb). - 3a87a2a2: Ignore dataset folder (
.gitignore). - c1426c53: Remove extra images lacking annotations; fix segmentation name bug (
bdd_preprocessing.ipynb). - af5496db: Add
voxel.pyfor dataset viewing. - bb98d4b9: Add
BDD.py,data/bdd.yaml; start dataset constructor. - 672cb355: Add database build logic (update
BDD.py,data/bdd.yaml). - 00c2b315: Fix dataloading bug (update
BDD.py,data/bdd.yaml). - 360aab9b: Use scalabel poly patching for better masks (
bdd_preprocessing.ipynb). - 90319d5a: Augment seg mask, add generator (
BDD.py,utils/augmentations.py). - 7ce18f16: Create BDD dataloader (
utils/dataloaders.py). - 418d1785: Use BDD dataloader in training (
train.py). - 85a43046: Add segmentation head to model (
models/yolov5s.yaml). - 4041a75e: Define
SemanticSegmentationOut(models/common.py). - 5e912242: Parse seg output from YAML (
models/yolo.py). - e4275c56: Return segmentation result from model forward (
models/yolo.py). - d9517402: Bug fixes across model and dataloader.
- ba2f05d0: Finalize training: bug fixes, compute/log seg loss (
train.py,val.py,utils/dataloaders.py, others).
- Multi-class segmentation: Extend from binary drivable-area to multiple scene classes (lanes, sidewalks, vegetation, etc.).
- Better losses: Try Dice, Tversky, or BCE+Dice combos; class-imbalance handling for finer edges.
- Visualization: Add scripts to overlay predicted masks on images and log to
runs/during train/val. - Post-processing: CRFs or morphology for cleaner boundaries if required.
- Inference pipeline: Export segmentation-friendly ONNX/TensorRT and add a CLI to run detection+segmentation on videos.
- Data quality: Additional polygon fixes and mask smoothing during preprocessing.
- Built on top of Ultralytics YOLOv5. See
LICENSEand original project documentation for details.



