Fusion prompt information to enhance multimodal video captioning
LIU Weiguang
ZUO Zaoxue
Abstract:Video caption task faces challenges in capturing critical local features and higher-order relation-ships among features.In the commonly used model SwinBERT,the Video Swin Transformer generates many irrelevant features and noises.To address these issues,a multimodal video caption model Prompt Vid was pro-posed to enhance the framework of Video Swin Transformer,which incorporated an adaptive multilayer per-ceptron module and an adaptive sparse self-attention mechanism.Experimental results demonstrate that the model can generate accurate and semantically rich video descriptions with effectiveness and robustness.
Keywords:video captionmultimodalityadapt perceptronattention mechanism
Publication Date:2025-06-25
Online Publishing Date:2025-08-15(First online date of this platform, not the publication date of the document)
Pages:7( 13-19 )