Gaze Estimation Method Using ResNet18 and Transformer Hybrid
WANG Yanxia
LIU Dan
Abstract:CNN has a powerful feature extraction capability.However,it is difficult to capture global information due to the limited receptive field.In contrast,Transformer can effectively obtain global features due to its long-distance modeling capability.Moreover,aiming at the problem that minor changes in the eye area in the gaze estimation task can easily lead to significant devia-tions in the gaze angle,this paper proposes the GazeTR-LM model.By combining the advantages of CNN and Transformer,and in-troducing the attention mechanism and MAD(Multi-Attention Dropping),the model's attention to the key areas of the eyes is en-hanced,achieving adaptive exploration of local areas,effectively suppressing the interference of redundant information,thereby im-proving the overall performance.To verify the validity of the model,experiments are conducted on three datasets,which are EyeDi-ap,MPIIFaceGaze and Gaze360.Experimental results show that the performance of GazeTR-LM improves by 1.4%,3.5%and 2.8%on EyeDiap,MPIIFaceGaze and Gaze360 datasets,compared to the currently better GazeTR-Hybrid model.
Keywords:Transformergaze estimationattention mechanismregularization
Publication Date:2025-09-20
Online Publishing Date:2025-12-23(First online date of this platform, not the publication date of the document)
Pages:4( 2461-2464 )
