ROD-Dataset: A Pedestrian-Perspective Benchmark of 25 Urban Obstacle Classes for Real-Time Detection
- Paper (coming soon)
- arXiv (coming soon)
- DOI (coming soon)
- Code (coming soon)
- Dataset (coming soon)
Abstract
ROD-Dataset is a pedestrian-perspective benchmark for real-time obstacle detection. It contains more than 24,000 eye-level images and over 40,000 human-verified obstacle annotations across 25 classes, from vehicles and people to stairs, manholes, traffic cones and benches. It unifies 29 public collections under a single taxonomy and adds new captures from Tehran and Toronto. We benchmark six nano-scale YOLO detectors, all of which exceed 0.86 recall: YOLO12n achieves the best recall (0.889), YOLO26n the best precision (0.925), and YOLOv9t the best mAP@0.50 (0.882).
Highlights
- 24,000+ eye-level images and 40,000+ human-verified obstacle annotations.
- 25 obstacle classes, from vehicles and people to stairs, manholes, traffic cones and benches.
- 29 public collections unified onto one taxonomy, plus new captures from Tehran and Toronto.
Baseline results
Six nano-scale YOLO detectors were benchmarked; all exceed 0.86 recall.
| Detector | Best at | Score |
|---|---|---|
| YOLO12n | Recall | 0.889 |
| YOLO26n | Precision | 0.925 |
| YOLOv9t | mAP@0.50 | 0.882 |
Cite this work
@misc{zandi2026roddataset,
title = {ROD-Dataset: A Pedestrian-Perspective Benchmark of 25 Urban Obstacle Classes for Real-Time Detection},
author = {Zandi, Abtin and Azami, Ariyan and Sabbagh Kermani, Bardia and Abbasian, Parsa and Ganjipour, Roza and Nami, Sarvin and Farbeh, Hamed},
year = {2026},
note = {Under review}
}