3 citations · 14 across the 33 of their papers we have counts for
10 papers · 1 filter
Struct-Searcher: Agentic Structural Thinking Advances Multimodal Deep Information Seeking
Fan Zhang, Vireo Zhang, Shengju Qian +7
Deep research agents have attracted increasing attention for their ability to collect large-scale online information to acquire target knowledge, with recent efforts shifting from…
SurgLQA: Scalable Long-Horizon Surgical Video Question Answering
Diandian Guo, Xikai Yang, Ruiyang Li +2
Surgical Video Question Answering (VideoQA) provides a promising paradigm for dynamic intraoperative interpretation, enabling real-time decision support and context-aware retrieval…
Med-Evo: Test-time Self-evolution for Medical Multimodal Large Language Models
Dunyuan Xu, Xikai Yang, Juzheng Miao +3
Medical Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities across diverse healthcare tasks. However, current post-training strategies, such as super…
REAR: Rethinking Visual Autoregressive Models via Generator-Tokenizer Consistency Regularization
Qiyuan He, Yicong Li, Haotian Ye +6
Visual autoregressive (AR) generation offers a promising path toward unifying vision and language models, yet its performance remains suboptimal against diffusion models. Prior wor…
MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language Models
Fan Zhang, Zebang Cheng, Chong Deng +18
Recent advances in multimodal large language models (MLLMs) have catalyzed transformative progress in affective computing, enabling models to exhibit emergent emotional intelligenc…
Point Cloud Understanding via Attention-Driven Contrastive Learning
Yi Wang, Jiaze Wang, Ziyu Guo +5
Recently Transformer-based models have advanced point cloud understanding by leveraging self-attention mechanisms, however, these methods often overlook latent information in less…