Saliency Object Detection Method Fusing Global and Local Features Based on Transformer
GE Yichen
ZHANG Ming
Abstract:To address the inefficiency of existing saliency object detection methods in global and local feature fusion,a saliency object detection method,namely,vision transformer-deep transformer decoder(ViT-DTD)was proposed.Firstly,the method chunked the input image by means of a visual transformer which first captured local features at a shallow level and then aggregated global features at a deep level.Secondly,feature information was ensured to be fully utilized in the decoding process by fusing shallow local features and deep global features layer by layer,combined with a dense connectivity strategy.Finally,the fused features were decoded by layer-by-layer upsampling to recover the high-resolution saliency map of the image.The results showed that the mean absolute error(MAE)of the method on the two benchmark test datasets was reduced by an average of 14.0%compared with the label decoupling framework(LDF)method,the weighted F metric was improved by an average of 3.7%,and the structure-measure(S)index was improved by an average of 2.7%.Compared with the other methods,the ViT-DTD has better performance in complex scenarios,and it provides a more accurate and stable saliency detection scheme.
Keywords:saliency object detectiontransformerlocal featureglobal featureViT-DTD
Publication Date:2024-12-20
Online Publishing Date:2025-08-15(First online date of this platform, not the publication date of the document)
Pages:6( 464-469 )