2 papers
cs.IR2026
ReCQR: Incorporating conversational query rewriting to improve Multimodal Image Retrieval
Yuan Hu, ZhiYu Cao, PeiFeng Li +1
With the rise of multimodal learning, image retrieval plays a crucial role in connecting visual information with natural language queries. Existing image retrievers struggle with p…
cs.CV2025
GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing
Ruizhe Ou, Yuan Hu, Fan Zhang +2
Multi-modal large language models (MLLMs) have achieved remarkable success in image- and region-level remote sensing (RS) image understanding tasks, such as image captioning, visua…