Compositional Zero-shot Learning Model Based on Pixel-level Feature Modulation and Text-guided Refinement
ZHAO Wei
BAO Xianglin
DU Wenlong
XU Xiaofeng
Abstract:To address the insufficient generalization to unseen attribute-object compositions in compositional zero-shot learning(CZSL),a pixel-level feature modulation and text-guided refinement for compositional zero-shot learning(PFMTR)model was proposed to boost the recognition of novel compositions.Firstly,a pixel-level feature modulation(PLFM)module was devised,which employed a dual attention mechanism operating at both pixel-level and patch-level to enable fine-grained reassembly and semantic enhancement of image features.Secondly,a text-guided refinement(TGR)module was proposed,where textual features were used as queries and visual features as keys/values.This module leveraged cross-modal attention to compute semantic relevance weights,thereby dynamically guiding visual features with linguistic semantics and achieving cross-modal alignment.The results showed that,compared with other state-of-the-art models,the PFMTR model achieved outstanding performance on the University of Texas Zappos(UT-Zappos)dataset under the open-world setting,attaining 35.7%in the area under curve(AUC)and 49.7%in the harmonic mean(HM).This study demonstrated that the recognition of unseen compositions could be effectively enhanced by integrating pixel-wise local modulation with cross-modal semantic guidance,offering a viable technical route for CZSL in complex scenarios.
Keywords:compositional zero-shot learningpixel-levelcross-modal alignmentattention mechanismvision-language model
Publication Date:2026-03-20
Online Publishing Date:2026-03-26(First online date of this platform, not the publication date of the document)
Pages:7( 75-81 )
