Skip to search boxSkip to navigationSkip to main content

Open access

Abstract

CLIP aligns images and texts through a global similarity score in a shared embedding space, but this scalar scoredoes not reveal which image regions or textual units support the alignment. This paper proposes Gradient-BasedInterpretable Contrastive Language–Image Pre-training (GBI-CLIP), a paired cross-modal explanation method alignedwith CLIP’s image–text similarity objective. Unlike attention-only or gradient-only explanations, GBI-CLIP combinesloosened Query–Key spatial weighting, Value features, and gradient-based channel sensitivity under the same similarityobjective. On the image side, it generates text-conditioned spatial attribution heatmaps; on the text side, it estimatesimage-conditioned token importance. Experiments on selected image–caption pairs from a validation subset of MSCOCO Val 2017 and selected examples from a domain-specific ExplainCars road-scene set show that GBI-CLIPproduces more coherent foreground localisation, clearer target boundaries, and reduced background noise comparedwith existing visual explanation methods. Ablation studies further confirm the role of the loosened spatial weight andattention-head selection, while runtime analysis shows that GBI-CLIP achieves approximately 13 ms per image–textpair on RTX 4080. These results demonstrate that GBI-CLIP provides a localisable, comparable, and efficient spatialand lexical evidence explanation method for CLIP-like vision-language models.

Publication Information

Output type

Research Output:
Other contribution
Other contribution

Original language

English

Publication milestones

  • Accepted/In press - 01/01/2026
  • Published - 15/08/2026

Publication status

Published - 15/08/2026

Publisher

SSRN

Publication IDs

  • ORCID: /0000-0003-4617-595X/work/223955839