7 papers
Empowering Long-form Omni-modal Understanding with Robust Audio Perception
Kaiying Yan, Luoyi Sun, Xiao Zhou +1
Recent advances in large-scale multimodal models have drivenremarkable progress in vision-language tasks; however, comprehensiveomni-modal understanding remains under-explored, lar…
POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management
Yikun Liu, Yuan Liu, Le Tian +6
Large Multimodal Models (LMMs) excel at visual perception but struggle with real-time, knowledge-intensive queries due to their reliance on static parametric knowledge. While multi…
SpotSound: Enhancing Large Audio-Language Models with Fine-Grained Temporal Grounding
Luoyi Sun, Xiao Zhou, Zeqian Li +3
Large Audio-Language Models (ALMs) have recently demonstrated remarkable capabilities in holistic audio understanding, yet they remain unreliable for temporal grounding, i.e., the…
POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs
Haicheng Wang, Yuan Liu, Yikun Liu +9
Multimodal Large Language Models (MLLMs) have recently demonstrated remarkable capabilities in cross-modal understanding and generation. However, the rapid growth of visual token s…
Knowledge-enhanced Pretraining for Vision-language Pathology Foundation Model on Cancer Diagnosis
Xiao Zhou, Luoyi Sun, Dexuan He +10
Vision-language foundation models have shown great promise in computational pathology but remain primarily data-driven, lacking explicit integration of medical knowledge. We introd…
Boosting Pathology Foundation Models via Few-shot Prompt-tuning for Rare Cancer Subtyping
Dexuan He, Xiao Zhou, Wenbin Guan +11
Rare cancers comprise 20-25% of all malignancies but face major diagnostic challenges due to limited expert availability-especially in pediatric oncology, where they represent over…