GBI-CLIP for spatial and lexical evidence explanation: Preprint
Research Output:
Other contribution
Other contribution
Open access
Abstract
CLIP aligns images and texts through a global similarity score in a shared embedding space, but this scalar scoredoes not reveal which image regions or textual units support the alignment. This paper proposes Gradient-BasedInterpretable Contrastive Language–Image Pre-training (GBI-CLIP), a paired cross-modal explanation method alignedwith CLIP’s image–text similarity objective. Unlike attention-only or gradient-only explanations, GBI-CLIP combinesloosened Query–Key spatial weighting, Value features, and gradient-based channel sensitivity under the same similarityobjective. On the image side, it generates text-conditioned spatial attribution heatmaps; on the text side, it estimatesimage-conditioned token importance. Experiments on selected image–caption pairs from a validation subset of MSCOCO Val 2017 and selected examples from a domain-specific ExplainCars road-scene set show that GBI-CLIPproduces more coherent foreground localisation, clearer target boundaries, and reduced background noise comparedwith existing visual explanation methods. Ablation studies further confirm the role of the loosened spatial weight andattention-head selection, while runtime analysis shows that GBI-CLIP achieves approximately 13 ms per image–textpair on RTX 4080. These results demonstrate that GBI-CLIP provides a localisable, comparable, and efficient spatialand lexical evidence explanation method for CLIP-like vision-language models.
Publication Information
Output type
Research Output:
Other contribution
Other contribution
Original language
EnglishPublication milestones
- Accepted/In press - 01/01/2026
- Published - 15/08/2026
Publication status
Published - 15/08/2026
Publisher
SSRNPublication IDs
- ORCID: /0000-0003-4617-595X/work/223955839
