8 papers
Stage-adaptive Token Selection for Efficient Omni-modal LLMs
Zijie Xin, Jie Yang, Ruixiang Zhao +4
Omni-modal large language models (om-LLMs) achieve unified audio-visual understanding by encoding video and audio into temporally aligned token sequences interleaved at the window…
OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding
Ruixiang Zhao, Jie Yang, Zijie Xin +4
Omni-proactive streaming video understanding, i.e., autonomously deciding when to speak and what to say from continuous audio-visual streams, is an emerging capability of omni-moda…
DAPL: Integration of Positive and Negative Descriptions in Text-Based Person Search
Yuchuan Deng, Zhanpeng Hu, Zijie Xin +2
Text-based person search (TBPS) aims to retrieve specific images of individuals from large datasets using textual descriptions. Existing TBPS methods focus primarily on identifying…
Fundus-R1: Training a Fundus-Reading MLLM with Knowledge-Aware Reasoning on Public Data
Yuchuan Deng, Qijie Wei, Kaiheng Qian +6
Fundus imaging such as CFP, OCT and UWF is crucial for the early detection of retinal anomalies and diseases. Fundus image understanding, due to its knowledge-intensive nature, pos…
SAVE: Speech-Aware Video Representation Learning for Video-Text Retrieval
Ruixiang Zhao, Zhihao Xu, Bangxiang Lan +3
For video-text retrieval, the use of CLIP has been a de facto choice. Since CLIP provides only image and text encoders, this consensus has led to a biased paradigm that entirely ig…
Learning Partially-Decorrelated Common Spaces for Ad-hoc Video Search
Fan Hu, Zijie Xin, Xirong Li
Ad-hoc Video Search (AVS) involves using a textual query to search for multiple relevant videos in a large collection of unlabeled short videos. The main challenge of AVS is the vi…