7 papers
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA
Ce Zhang, Ziyang Wang, Yulu Pan +6
Grounded long-video question answering (Grounded LVQA) requires answering a question about a long video while localizing the short evidence interval that supports the answer. Recen…
Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models
Yifu Yuan, Yaoting Huang, Xianze Yao +20
We introduce Embodied-R1.5, a unified Embodied Foundation Model (EFM) that integrates comprehensive embodied reasoning capabilities, spanning embodied cognition, task planning, cor…
Mixture of Probes: Learning from Privileged Modalities in Multimodal LLMs Through Probing
Dominick Reilly, Qiyu Wu, Hiromi Wakaki +2
Multimodal Large Language Models (MLLMs) are typically designed under the assumption that all modalities available during training will also be accessible at inference. However, ma…
DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics
Silin Gao, Hao Zhao, Zeming Chen +8
Multimodal LLMs struggle to systematically model the temporal evolution of visual scenes in videos or multi-image sequences. Such inputs require models to predict or simulate multi…
MusTBENCH: Benchmarking and Advancing Temporal Grounding in Music LLMs
Daeyong Kwon, Qiyu Wu, Shinobu Kuriya +6
Recent Large Audio-Language Models (LALMs) have demonstrated promising abilities in understanding musical content. However, whether their responses are grounded in the correct temp…
GRC: Unifying Reasoning-Driven Generation, Retrieval and Compression
Zhongtao Miao, Qiyu Wu, Yoshimasa Tsuruoka
Text embedding and generative tasks are usually trained separately based on large language models (LLMs) nowadays. This causes a large amount of training cost and deployment effort…