activity
20242026
collaborators

20 papers

cs.CV2026

Benchmarking Foundation and Large Language Models for Few-Shot Medical Image Segmentation

Jinghong Liu, Yuchuan Deng, Fanping Liu +2

The paper presents FAME, a unified benchmark for evaluating few-shot medical image segmentation methods across multiple anatomical sites, imaging modalities, and settings, and anal…

cs.CV2026

Stage-adaptive Token Selection for Efficient Omni-modal LLMs

Zijie Xin, Jie Yang, Ruixiang Zhao +4

Omni-modal large language models (om-LLMs) achieve unified audio-visual understanding by encoding video and audio into temporally aligned token sequences interleaved at the window…

cs.CV2026

OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding

Ruixiang Zhao, Jie Yang, Zijie Xin +4

Omni-proactive streaming video understanding, i.e., autonomously deciding when to speak and what to say from continuous audio-visual streams, is an emerging capability of omni-moda…

cs.CV2026

Fundus-R1: Training a Fundus-Reading MLLM with Knowledge-Aware Reasoning on Public Data

Yuchuan Deng, Qijie Wei, Kaiheng Qian +6

Fundus imaging such as CFP, OCT and UWF is crucial for the early detection of retinal anomalies and diseases. Fundus image understanding, due to its knowledge-intensive nature, pos…

cs.CV2026

EI: Early Intervention for Multimodal Imaging based Disease Recognition

Qijie Wei, Hailan Lin, Xirong Li

Current methods for multimodal medical imaging based disease recognition face two major challenges. First, the prevailing "fusion after unimodal image embedding" paradigm cannot fu…

cs.CV2026

SAVE: Speech-Aware Video Representation Learning for Video-Text Retrieval

Ruixiang Zhao, Zhihao Xu, Bangxiang Lan +3

For video-text retrieval, the use of CLIP has been a de facto choice. Since CLIP provides only image and text encoders, this consensus has led to a biased paradigm that entirely ig…