TY - GEN
T1 - Depth Augmented Semantic Segmentation Networks for Automated Driving
AU - Rashed, Hazem
AU - Yogamani, Senthil
AU - El-Sallab, Ahmad
AU - Das, Arindam
AU - El-Helw, Mohamed
N1 - Publisher Copyright:
© Springer Nature Singapore Pte Ltd, 2019.
PY - 2019
Y1 - 2019
N2 - In this paper, we explore the augmentation of depth maps to improve the performance of semantic segmentation motivated by the geometric structure in automotive scenes. Typically depth is already computed in an automotive system to localize objects and path planning and thus can be leveraged for semantic segmentation. We construct two networks that serve as a baseline for comparison which are “RGB only” and “Depth only”, and we investigate the impact of fusion of both cues using another two networks which are “RGBD concat”, and “Two Stream RGB+D”. We evaluate these networks on two automotive datasets namely Virtual KITTI using synthetic depth and Cityscapes using a standard stereo depth estimation algorithm. Additionally, we evaluate our approach using monoDepth unsupervised estimator [10]. Two-stream architecture achieves the best results with an improvement of 5.7% IoU in Virtual KITTI and 1% IoU in Cityscapes. There is a large improvement for certain classes like trucks, building, van and cars which have an increase of 29%, 11%, 9% and 8% respectively in Virtual KITTI. Surprisingly, CNN model is able to produce good semantic segmentation from depth images only. The proposed network runs at 4 fps on TitanX GPU, Maxwell architecture.
AB - In this paper, we explore the augmentation of depth maps to improve the performance of semantic segmentation motivated by the geometric structure in automotive scenes. Typically depth is already computed in an automotive system to localize objects and path planning and thus can be leveraged for semantic segmentation. We construct two networks that serve as a baseline for comparison which are “RGB only” and “Depth only”, and we investigate the impact of fusion of both cues using another two networks which are “RGBD concat”, and “Two Stream RGB+D”. We evaluate these networks on two automotive datasets namely Virtual KITTI using synthetic depth and Cityscapes using a standard stereo depth estimation algorithm. Additionally, we evaluate our approach using monoDepth unsupervised estimator [10]. Two-stream architecture achieves the best results with an improvement of 5.7% IoU in Virtual KITTI and 1% IoU in Cityscapes. There is a large improvement for certain classes like trucks, building, van and cars which have an increase of 29%, 11%, 9% and 8% respectively in Virtual KITTI. Surprisingly, CNN model is able to produce good semantic segmentation from depth images only. The proposed network runs at 4 fps on TitanX GPU, Maxwell architecture.
KW - Automated driving
KW - Semantic segmentation
KW - Visual perception
UR - https://www.scopus.com/pages/publications/85076696935
U2 - 10.1007/978-981-15-1387-9_1
DO - 10.1007/978-981-15-1387-9_1
M3 - Conference contribution
AN - SCOPUS:85076696935
SN - 9789811513862
T3 - Communications in Computer and Information Science
SP - 1
EP - 13
BT - Computer Vision Applications - 3rd Workshop, WCVA 2018, held in Conjunction with ICVGIP 2018, Revised Selected Papers
A2 - Arora, Chetan
A2 - Mitra, Kaushik
PB - Springer
T2 - 3rd Workshop on Computer Vision Applications, WCVA 2018, held in conjunction with the 11th Indian Conference on Computer Vision, Graphics and Image Processing, ICVGIP 2018
Y2 - 18 December 2018 through 18 December 2018
ER -