Behavior Recognition Algorithm Based on Two-stream Transformer Network
TAI Aibing
ZHA Keke
Abstract:In recent years,most of the research works are based on 3DConvNet to build models,extract temporal features,and optimize them step by step.However,3DConv also poses the problems of insufficient computational demand and high deploy-ment cost.In view of the advantages of VIT in sequential data feature extraction,this paper proposes a Two-stream Vision Trans-former Network(TVTN)model based on Transformer.In order to adapt to videos of different lengths,TVTN obtains RGB frames and corresponding optical stream images by segmented sparse sampling of the videos,and obtains the temporal features of the ac-tions by processing the optical stream images with the VIT Hybrid algorithm,and the spatial features of the actions by processing the RGB frames.The prediction results of the two networks are fused to complete the task of recognizing actions at the video level.This paper discusses the performance of different VIT and VIT Hybrid in Flow and video,respectively,and the effect of different weight values in the result fusion and different number of frames of input on the experimental results of TVTN.The experimental results show that the TVTN network,pre-trained only on ImageNet,achieves 96.1%accuracy on the UCF-101 dataset and also performs well on the HMDB-51 dataset in terms of improvement,demonstrating the effectiveness of the model in video behavior classifica-tion.
Keywords:behavior recognitionVision Transformercomputer vision
Publication Date:2025-08-20
Online Publishing Date:2025-12-12(First online date of this platform, not the publication date of the document)
Pages:6( 2089-2094 )
