5 papers · 1 filter
GRAPE: Let GRPO Supervise Query Rewriting by Ranking for Retrieval
Zhaohua Zhang, Jianhuan Zhuo, Muxi Chen +8
The CLIP model has established itself as a cornerstone of large-scale retrieval systems. However, its performance often degrades under distributional shifts such as multilingual, l…
FBCIR: Balancing Cross-Modal Focuses in Composed Image Retrieval
Chenchen Zhao, Jianhuan Zhuo, Muxi Chen +6
Composed image retrieval (CIR) requires multi-modal models to jointly reason over visual content and semantic modifications presented in text-image input pairs. While current CIR m…
\textit{FocaLogic}: Logic-Based Interpretation of Visual Model Decisions
Chenchen Zhao, Muxi Chen, Qiang Xu
Interpretability of modern visual models is crucial, particularly in high-stakes applications. However, existing interpretability methods typically suffer from either reliance on w…
FailureAtlas:Mapping the Failure Landscape of T2I Models via Active Exploration
Muxi Chen, Zhaohua Zhang, Chenchen Zhao +8
Static benchmarks have provided a valuable foundation for comparing Text-to-Image (T2I) models. However, their passive design offers limited diagnostic power, struggling to uncover…
HiBug2: Efficient and Interpretable Error Slice Discovery for Comprehensive Model Debugging
Muxi Chen, Chenchen Zhao, Qiang Xu
Despite the significant success of deep learning models in computer vision, they often exhibit systematic failures on specific data subsets, known as error slices. Identifying and…