Dense Pixel Labelling of Road Scenes Using a Deep Supervised Encoder-Decoder Network with ResNet101

Authors:
S. Vaenkatakrishnan, N. Sankar Ram, M. Siva Subramanian, B. Sai Saran, R. Vinoth, K. Daniel Jasper

Addresses:
Department of Computer Science and Engineering in AIML, SRM Institute of Science and Technology, Ramapuram, Chennai, Tamil Nadu, India.  School of Electronics, Electrical Engineering and Computer Science, Queens University Belfast, Belfast, Northern Ireland, United Kingdom.

Abstract:

Scene understanding at the pixel-level has long been the essential enabling function of autonomous navigation systems and smart transportation infrastructures that operate in real-world and diverse environments. Researchers propose an end-to-end, fully convolutional deep learning framework for the challenging task of semantic segmentation of urban street images: DS-UNet-ResNet101. The network employs a pre-trained ResNet101 residual deep learning network as the encoder, and its corresponding U-Net-like symmetry-progressive decoder progressively reconstructs spatial resolution in 4 steps. A novel multi-stage auxiliary supervision scheme embedded in each intermediate decoder stage adds additional gradient paths to supplement the primary training signal, improving stable convergence and better intermediate features. Experiments have been conducted on the urban driving benchmark, CamVid, which contains 32 semantic classes across various environments, illumination conditions, and traffic situations, and is a richly annotated dataset. Geometric data augmentation, including random horizontal flipping, controlled random rotation, illumination variations, and synchronous mask transformations, has improved distributional variance and reduced overfitting in the tiny training division. With 26M parameters, DS-UNet-ResNet101 surpasses FCN-8s, SegNet, U-Net, and DeepLabV3+ in all three measures (mIoU: 0.749, mPA: 0.755, MAE: 0.068 Qualitative results demonstrate accurate segmentation of dominant features (roads, sky, trees) but difficulty with minor, largely obscured targets (“pedestrians and cyclists). Researchers built a Flask application with inference, dataset exploration and viewing, augmentation preview, and training tracking to connect prototype research and applications. Intelligent and reliable autonomous transportation systems benefit from our findings. 

Keywords: Semantic Segmentation; Autonomous Navigation; Deep Learning (DL); Urban Street Images; Flask Application; Training Convergence; Data Augmentation; Urban Driving.

Received on: 18/05/2025, Revised on: 29/07/2025, Accepted on: 24/09/2025, Published on: 05/06/2026

DOI: 10.69888/FTSIN.2026.000708

FMDB Transactions on Sustainable Intelligent Networks, 2026 Vol. 3 No. 2, Pages: 124-136

  • Views : 75
  • Downloads : 7
Download PDF