9 papers
Comprehensive language-image pre-training for 3D medical image understanding
Tassilo Wald, Ibrahim Ethem Hamamci, Yuan Gao +14
In the 3D medical image domain, vision-language pre-training is used to create vision-language encoders (VLEs) that can support radiologists by retrieving patients with similar abn…
Beyond Pixels: Benchmarking and Reward-Based Assessing Framework for Visual Spatial Aesthetics
Yuan Gao, Jin Song, Yiyun Fei +2
In recent years, Image Quality Assessment (IQA) for AI-generated images (AIGI) has advanced rapidly; however, existing methods primarily target portraits and artistic images, lacki…
MERIT: Multilingual Semantic Retrieval with Interleaved Multi-Condition Query
Wei Chow, Yuan Gao, Linfeng Li +15
Semantic retrieval is crucial for modern applications yet remains underexplored in current research. Existing datasets are limited to single languages, single images, or singular r…
VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models
Fufangchen Zhao, Liao Zhang, Daiqi Shi +5
We propose VideoPerceiver, a novel video multimodal large language model (VMLLM) that enhances fine-grained perception in video understanding, addressing VMLLMs' limited ability to…
InfoCLIP: Bridging Vision-Language Pretraining and Open-Vocabulary Semantic Segmentation via Information-Theoretic Alignment Transfer
Muyao Yuan, Yuanhong Zhang, Weizhan Zhang +4
Recently, the strong generalization ability of CLIP has facilitated open-vocabulary semantic segmentation, which labels pixels using arbitrary text. However, existing methods that…
On the Faithfulness of Visual Thinking: Measurement and Enhancement
Zujing Liu, Junwen Pan, Qi She +2
Recent large vision-language models (LVLMs) can generate vision-text multimodal chain-of-thought (MCoT) traces after reinforcement fine-tuning (RFT). However, we observe that the v…