#image-text alignment
try —
3 papers match
cs.CV2026
Visual Credit Audit for Multimodal Spatial Reasoning
Feixiang Liu, Qiang Qiu, Lanbo Sun +3
The paper introduces Visual Credit Audit (VCA), a method to quantify how much an image actually contributes to a multimodal model’s answer on spatial reasoning tasks, separating co…
#visual credit audit#multimodal spatial reasoning#benchmark evaluation#large language models
cs.CV2026
AspectCLIP: Optimizing CLIP Representation Space via Aspect-Guided Consistency Regularization
Yiyang Yao, Shanglin Liu, Jianming Lv +4
The paper introduces AspectCLIP, a method that groups image-text pairs by shared textual aspects and applies consistency regularization within these groups to avoid forcing unrelat…
#contrastive learning#image-text alignment#representation learning#aspect-guided regularization
cs.CV2026
SOLAR: Self-supervised Joint Learning for Symmetric Multimodal Retrieval
Wenjie Yang, Hang Yu, Yuyu Guo +1
The paper introduces SOLAR, a self‑supervised two‑stage framework for symmetric multimodal‑to‑multimodal retrieval that learns intersection masks from large unlabeled image‑text pa…
#multimodal retrieval#self-supervised learning#symmetric retrieval#image-text alignment