2 papers
cs.CV2025
Cross-modal Full-mode Fine-grained Alignment for Text-to-Image Person Retrieval
Hao Yin, Xin Man, Feiyu Chen +2
Text-to-Image Person Retrieval (TIPR) is a cross-modal matching task designed to identify the person images that best correspond to a given textual description. The key difficulty…
cs.CV2025
LISA: A Layer-wise Integration and Suppression Approach for Hallucination Mitigation in Multimodal Large Language Models
Zhihui Guo, Xin Man, Hui Xu +4
Multimodal Large Language Models (MLLMs) excel in vision-language tasks such as image captioning but remain prone to object hallucinations, where they describe objects that do not…