5 papers · 1 filter
Towards Unified Surgical Scene Understanding:Bridging Reasoning and Grounding via MLLMs
Jincai Huang, Shihao Zou, Yuchen Guo +5
Surgical scene understanding is a cornerstone of computer-assisted intervention. While recent advances, particularly in surgical image segmentation, have driven progress, real-worl…
UniM: A Unified Any-to-Any Interleaved Multimodal Benchmark
Yanlin Li, Minghui Guo, Kaiwen Zhang +13
In real-world multimodal applications, systems usually need to comprehend arbitrarily combined and interleaved multimodal inputs from users, while also generating outputs in any in…
VimoRAG: Video-based Retrieval-augmented 3D Motion Generation for Motion Language Models
Haidong Xu, Guangwei Xu, Zhedong Zheng +7
This paper introduces VimoRAG, a novel video-based retrieval-augmented motion generation framework for motion large language models (LLMs). As motion LLMs face severe out-of-domain…
What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities
Wendong Bu, Yang Wu, Qifan Yu +10
As multimodal large language models (MLLMs) advance, MLLM-based virtual agents have demonstrated remarkable performance. However, existing benchmarks face significant limitations,…
De-fine: Decomposing and Refining Visual Programs with Auto-Feedback
Minghe Gao, Juncheng Li, Hao Fei +7
Visual programming, a modular and generalizable paradigm, integrates different modules and Python operators to solve various vision-language tasks. Unlike end-to-end models that ne…