Image Inpainting Model Based on TLIAM-Net
TENG Lin
ZHANG Qian
BAI Wuer
Abstract:To address the issues of insufficient utilization of spatial information and semantic ambiguity in the inpainting of large missing areas of the existing Transformer-based image inpainting models,a Transformer-local importance-based attention and Mish activation function encoder-decoder network(TLIAM-Net)image inpainting model was proposed.Firstly,the TLIAM-Net model was designed with an encoder-decoder architecture,where Transformer blocks were progressively connected to comprehensively exploit the hierarchical feature information of the image.Subsequently,the local importance-based attention(LIA)mechanism was introduced following each Transformer block,to enhance the model′s spatial information utilization capability.Finally,the Mish activation function was employed within the Transformer blocks,to ensure smoother feature transitions and improved capture of subtle details.The results demonstrated that the TLIAM-Net model achievedthe Fréchet inception distance(FID)value of 12.0116 under mask rates of(0.5,0.6]on Flickr-faces-high quality dataset(FFHQ),representing a 51.17%reduction compared to the multi-level interactive siamese filtering(MISF)model.The accuracy and deblurring performance of image inpainting tasks were significantly improved by the TLIAM-Net model,which could be successfully applied to multiple subtasks including deblurring,denoising,and defect completion,and the cost of manual inpainting was substantially reduced.
Keywords:image inpaintingspatial attentionactivation functionTransformerencoder-decoder
Publication Date:2025-12-30
Online Publishing Date:2025-12-10(First online date of this platform, not the publication date of the document)
Pages:7( 504-510 )