3 papers
cs.CV2025
Reasoning Guided Embeddings: Leveraging MLLM Reasoning for Improved Multimodal Retrieval
Chunxu Liu, Jiyuan Yang, Ruopeng Gao +4
Multimodal embeddings are widely used in downstream tasks such as multimodal retrieval, enabling alignment of interleaved modalities in a shared representation space. While recent…
cs.CV2025
SORCE: Small Object Retrieval in Complex Environments
Chunxu Liu, Chi Xie, Xiaxu Chen +4
Text-to-Image Retrieval (T2IR) is a highly valuable task that aims to match a given textual query to images in a gallery. Existing benchmarks primarily focus on textual queries des…
cs.CV2025
On the Suitability of Reinforcement Fine-Tuning to Visual Tasks
Xiaxu Chen, Wei Li, Chunxu Liu +5
Reinforcement Fine-Tuning (RFT) is proved to be greatly valuable for enhancing the reasoning ability of LLMs. Researchers have been starting to apply RFT to MLLMs, hoping it will a…