Unsupervised Video Object Segmentation Based on Discrete Cosine Transform Feature Fusion
WANG Yuchen
FAN Jiaqing
SONG Huihui
Abstract:Unsupervised Video Object Segmentation(UVOS)aims to localize and segment foreground objects in videos with-out manually providing the ground-truth object segmentation mask for the first frame.Existing methods mainly focus on improving segmentation accuracy while ignoring memory and computational cost.Generally,the existing methods only enhance the fusion fea-tures of appearance and motion in the spatial domain according to their significance,ignoring the particularity of features in the fre-quency domain.In addition,the existing methods do not make full use of global semantic information to guide video object segmenta-tion.To solve the problems above,this paper proposes a lightweight UVOS network based on discrete cosine transform feature fu-sion.Firstly,a lightweight backbone network is used to extract appearance and motion features simultaneously.Secondly,the dis-crete cosine transform feature fusion module is designed to fuse and enhance the appearance and motion features.Then,the large kernel convolution global semantic guidance module is used to integrate the large kernel volume,which can reduce the computation-al complexity of large kernel convolution and keep the ability of extracting global semantic information.Finally,under the guidance of global semantic information,the multi-level features enhanced in frequency domain are aggregated progressively,and finally the accurate segmentation results are obtained.Through the aforementioned designs,the presented method has only 14.7 M parameters.A large number of experimental evaluations are conducted on DAVIS2016,FBMS and DAVSOD datasets,showing that the method achieves favorable performance on J&F,MAE and Fm also keeps high reasoning speed.
Keywords:unsupervised video object segmentationdiscrete cosine transformattention mechanismfrequency do-main analysis
Publication Date:2025-02-20
Online Publishing Date:2025-08-15(First online date of this platform, not the publication date of the document)
Pages:8( 395-402 )
