Few-shot action recognition in video method based on continuous fame information fusion modeling
Zhang Bingbing
Li Haibo
Ma Yuanchen
Zhang Jianxin
Abstract:Objectives To overcome the limitations of existing few-shot video action recognition methods in capturing global spatiotemporal information and modeling complex behaviors,a new network architecture was developed to significantly enhances the accuracy and robustness of few-shot learning in video action recognition tasks.Methods A network architecture was presented integrating a continuous frame information fusion module and a multi-dimensional attention modeling module.The continuous frame information fusion module was positioned at the input end of the network,primarily responsible for capturing and transforming low-level information into richer high-level semantic information.The multi-dimensional attention modeling module was set in the middle layer of the network.The entire network was designed based on a 2D convolu-tional model,effectively reducing computational complexity.Results Experiments on four mainstream action recognition datasetsshowed that,on the Something-Something V2 dataset,the accuracy rates for 1-shot and 5-shot tasks reached 50.8%and 68.5%,respectively;on the Kinetics-100 dataset,the 1-shot and 5-shot tasks achieved accuracy rates of 68.5%and 83.8%,respectively,showing significant improvement over ex-isting methods;on the UCF101 dataset,the method achieved an accuracy rate of 81.3%for the 1-shot task and 93.8%for the 5-shot task,both markedly superior to baseline methods.Additionally,on the HMDB51 dataset,the method demonstrated good generalization performance,with accuracy rates of 56.0%for the 1-shot task and 74.4%for the 5-shot task.Conclusions The continuous frame integration modeling network has shown significant advantages in improving the model's ability to process complex spatiotemporal infor-mationThe solutions presented in this study could introduce effective new methods to the field of few-shot action recognition,demonstrating their efficiency and practicality.
Keywords:few-shot learningvideo action recognitionspatiotemporal modelingspatiotemporal representa-tion learningcontinuous frame information
Publication Date:2025-07-31
Online Publishing Date:2025-08-15(First online date of this platform, not the publication date of the document)
Pages:10( 11-20 )
