Rich semantic extractor network for real-time semantic segmentation
ZHAO Shan
TIAN Kaiwen
SUN Junding
Abstract:Objectives The inference speed of the real-time semantic segmentation network is limited,the depth of the network is shallow,which lead to insufficient semantic feature information extracted.Addition-ally,the shallow network depth restricts the capability of feature extraction networks,reducing their robust-ness and adaptability.In order to solve such the problems,Methods a rich semantic extractor network(RSENet)for real-time semantic segmentation was proposed.Firstly,aiming at the problem of inadequate se-mantic feature information extraction,a rich semantic extractor(RSE)was introduced,which included a multi-scale global semantic extraction module(MGSEM)and a semantic fusion module(SFM).MGSEM was used to extract rich multi-scale global semantics and expand the effective receptive field of the network.At the same time,SFM efficiently fused multi-scale local semantics and multi-scale global semantics,so that the network had more comprehensive and rich semantic information.Finally,according to the characteristics of the detailed branch and the semantic branch,a space reconstruction aggregation module(SRAM)was de-signed to model the context information of the detailed features and enhanced the feature representation,so that the two branches could be efficiently aggregated.Results Comprehensive experiments were conducted on Cityscapes and ADE20K datasets,and the proposed RSENet achieved mIoU of 75.6%and 35.7%at in-ference speed of 76 frames/s and 67 frames/s,respectively.Conclusions The experimental results suggested that in the extraction of semantic information within complex scenes,the network proposed in this paper was able to deeply explore and accurately capture such semantic information in images.Furthermore,outstanding performance was demonstrated in achieving a balance between accuracy and speed,with the network not only capable of achieving high-precision semantic segmentation but also exhibiting very fast inference speeds.This efficient image segmentation capability endowed the network with high practicality and operabil-ity in real-world application scenarios.
Keywords:semantic segmentationmulti-scale featurevision Transformerfeature fusion
Publication Date:2024-12-28
Online Publishing Date:2025-08-15(First online date of this platform, not the publication date of the document)
Pages:10( 146-155 )
Journal of Henan Polytechnic University(Natural Science)

Journal of Henan Polytechnic University(Natural Science)

ISTICPKU
ISSN:1673-9787
Year, Vol.(Issue):2024,43(6)