Skip to content
 
 

Repository files navigation

BDD100K Drivable-Area Segmentation with YOLOv5 (Detection + Seg)

This repository extends YOLOv5 to jointly perform object detection and semantic segmentation of drivable area on the BDD100K dataset. It adds a lightweight segmentation head to the YOLOv5s model, a custom dataset/dataloader that returns both detection and segmentation targets, and training/validation changes to compute and log segmentation loss alongside standard detection losses.

Goals

  • Build an end-to-end pipeline to train on BDD100K with both detection (10 classes) and a drivable-area segmentation task.
  • Keep model changes minimal and compatible with YOLOv5 style and utilities.
  • Provide reproducible data preprocessing, a clear dataset layout, and simple commands to train and validate.

Key Changes (What and Why)

  • New segmentation head in YOLOv5s

    • What: Introduced SemanticSegmentationOut in models/common.py and wired a 2‑channel head (background, drivable) into models/yolov5s.yaml. The forward in models/yolo.py returns both detection outputs and the segmentation map.
    • Why: To jointly learn drivable-area segmentation using a light head that reuses backbone features, enabling multi-task learning with minimal overhead.
  • BDD-specific dataset and dataloader

    • What: Implemented BDDDataset in BDD.py to read BDD images, YOLO-format detection labels, and segmentation masks; applies letterbox/augmentations consistently to images, boxes, and masks; returns both detection targets and a 2‑channel segmentation target. Added create_bdd_dataloader(...) in utils/dataloaders.py plus a custom collate_fn.
    • Why: To guarantee alignment between images, boxes, and masks under resizing and augmentation while keeping the API similar to YOLOv5’s.
  • Training/validation integration

    • What: train.py and val.py now use the BDD dataloader and compute a segmentation loss using BCEWithLogitsLoss in addition to box/objectness/class losses. Segmentation loss is logged and aggregated like other losses.
    • Why: Seamless multi-task training using familiar YOLOv5 training scripts and logging.
  • Data preparation utilities

    • What: bdd_preprocessing.ipynb converts original BDD annotations to YOLO detection labels and generates/fixes segmentation masks (including polygon patching). voxel.py provides quick dataset visualization.
    • Why: Reproducible preprocessing to build a clean, aligned dataset for training/validation.

Repository Guide (What I Wrote/Touched)

  • BDD.py — Defines BDDDataset:

    • Builds an index of images with matching detection .txt and segmentation mask files.
    • Loads and letterboxes image and mask to a consistent size, applies augmentations (HSV, random perspective, flips), and converts detection labels between xyxy and normalized xywh as required.
    • Produces a 2‑channel segmentation target [background, drivable] and returns (image, [det_labels, seg_mask], path, shapes).
    • Provides collate_fn to batch both detection and segmentation targets.
  • utils/dataloaders.py — Adds create_bdd_dataloader(...):

    • Mirrors the YOLOv5 create_dataloader logic but instantiates BDDDataset and uses its collate_fn.
  • models/common.py — Adds SemanticSegmentationOut:

    • A small conv + BN + activation head that outputs a 2‑channel segmentation map.
  • models/yolov5s.yaml — Modifies the head:

    • Adds upsampling and C3 layers from P3 features into a final SemanticSegmentationOut(2, ...) layer that predicts [background, drivable].
  • models/yolo.py — Model parsing/forward updates:

    • Recognizes SemanticSegmentationOut in parse_model and returns (detections, seg_out) from BaseModel._forward_once.
  • train.py — Training loop updates:

    • Uses create_bdd_dataloader for train/val.
    • Computes segmentation loss with BCEWithLogitsLoss and logs it alongside detection losses.
  • val.py — Validation loop updates:

    • Validates with the BDD dataloader, computes segmentation loss for reporting, and uses the usual NMS/metrics for detection.
  • data/bdd.yaml — Dataset config:

    • Defines 10 detection classes and the dataset root (default BDDDataset/).
  • bdd_preprocessing.ipynb — Preprocessing steps:

    • Converts detections to YOLO format, generates masks, and patches polygons for better segmentation masks.

Dataset Layout

Place your dataset as follows (or update data/bdd.yaml:path):

BDDDataset/
  images/
    train/ ...
    val/   ...
    test/  ...
  labels/
    det/
      train/ ...  # .txt in YOLO format
      val/   ...
      test/  ...
    seg/
      train/ ...  # segmentation masks (image files)
      val/   ...
      test/  ...

Installation

python -m venv .venv && source .venv/bin/activate  # or use your preferred env
pip install -r requirements.txt

Training

python train.py \
  --weights yolov5s.pt \
  --cfg models/yolov5s.yaml \
  --data data/bdd.yaml \
  --epochs 100 \
  --batch-size 16 \
  --imgsz 640

Tips:

  • Ensure data/bdd.yaml:path points to your BDDDataset/ root.
  • train.py defaults are already set for this project; you can omit flags if you accept defaults.

Validation

python val.py --weights runs/train/exp29/weights/best.pt --data data/bdd.yaml --imgsz 640

Results (Samples)

The following images come from recent training runs (runs/train/exp29). They illustrate batch visuals and label statistics during training.

Train Batch 0 Train Batch 1 Train Batch 2 Label Distribution

Note: The training visualizations primarily show detection labels/predictions. The segmentation head produces a drivable-area map (2 channels). If desired, you can add utilities to overlay the predicted mask on images for richer qualitative results.

Design Notes and Rationale

  • Minimal, modular change: Rather than forking into a large seg architecture, the project keeps YOLOv5 intact and adds a small head (SemanticSegmentationOut) that can be parsed from YAML. This preserves most of the ecosystem while enabling segmentation.
  • Aligned augmentations: All image transforms (resize, letterbox, perspective, flips) are applied consistently to both the image and the segmentation mask to keep labels aligned.
  • Loss choice: A simple BCEWithLogitsLoss is used for segmentation for its stability and speed on binary masks. It can be swapped for Dice, Focal, or Combo losses if you emphasize mask quality.
  • Data efficiency: Reuses the strong YOLOv5 backbone + head structure for feature reuse and efficient joint training.

Commit Summary (My Contributions)

Oldest → newest highlights of changes on master:

  • 2edfa4cc: Preprocess data, convert detections to YOLO, create segmentation masks (bdd_preprocessing.ipynb).
  • 3a87a2a2: Ignore dataset folder (.gitignore).
  • c1426c53: Remove extra images lacking annotations; fix segmentation name bug (bdd_preprocessing.ipynb).
  • af5496db: Add voxel.py for dataset viewing.
  • bb98d4b9: Add BDD.py, data/bdd.yaml; start dataset constructor.
  • 672cb355: Add database build logic (update BDD.py, data/bdd.yaml).
  • 00c2b315: Fix dataloading bug (update BDD.py, data/bdd.yaml).
  • 360aab9b: Use scalabel poly patching for better masks (bdd_preprocessing.ipynb).
  • 90319d5a: Augment seg mask, add generator (BDD.py, utils/augmentations.py).
  • 7ce18f16: Create BDD dataloader (utils/dataloaders.py).
  • 418d1785: Use BDD dataloader in training (train.py).
  • 85a43046: Add segmentation head to model (models/yolov5s.yaml).
  • 4041a75e: Define SemanticSegmentationOut (models/common.py).
  • 5e912242: Parse seg output from YAML (models/yolo.py).
  • e4275c56: Return segmentation result from model forward (models/yolo.py).
  • d9517402: Bug fixes across model and dataloader.
  • ba2f05d0: Finalize training: bug fixes, compute/log seg loss (train.py, val.py, utils/dataloaders.py, others).

Future Work

  • Multi-class segmentation: Extend from binary drivable-area to multiple scene classes (lanes, sidewalks, vegetation, etc.).
  • Better losses: Try Dice, Tversky, or BCE+Dice combos; class-imbalance handling for finer edges.
  • Visualization: Add scripts to overlay predicted masks on images and log to runs/ during train/val.
  • Post-processing: CRFs or morphology for cleaner boundaries if required.
  • Inference pipeline: Export segmentation-friendly ONNX/TensorRT and add a CLI to run detection+segmentation on videos.
  • Data quality: Additional polygon fixes and mask smoothing during preprocessing.

Acknowledgements

  • Built on top of Ultralytics YOLOv5. See LICENSE and original project documentation for details.

About

YOLOv5 Detection model with Segmentation head

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages