2 papers
cs.CL2026
MedUAG: Unified Understanding and Generation for Medical Multimodal Models
Zijie Meng, Yuncheng Zhang, Hualiang Wang +8
Recent Multimodal Large Language Models (MLLMs) are rapidly evolving into unified understanding and generation (UAG) frameworks. However, extending these unified paradigms to the m…
cs.CV2025
OmniGround: A Comprehensive Spatio-Temporal Grounding Benchmark for Real-World Complex Scenarios
Hong Gao, Jingyu Wu, Xiangkai Xu +5
Spatio-Temporal Video Grounding (STVG) aims to localize target objects in videos based on natural language descriptions. Despite recent advances in Multimodal Large Language Models…