1 citations · 4 across the 21 of their papers we have counts for
23 papers
MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations
Jongsuk Kim, Qiyu Wu, Zhuoyuan Mao +3
Pretrained vision-language models such as CLIP excel at zero-shot recognition but often fail at compositionality, particularly attribute-object and relational structures. Recent st…
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA
Ce Zhang, Ziyang Wang, Yulu Pan +6
Grounded long-video question answering (Grounded LVQA) requires answering a question about a long video while localizing the short evidence interval that supports the answer. Recen…
Mixture of Probes: Learning from Privileged Modalities in Multimodal LLMs Through Probing
Dominick Reilly, Qiyu Wu, Hiromi Wakaki +2
Multimodal Large Language Models (MLLMs) are typically designed under the assumption that all modalities available during training will also be accessible at inference. However, ma…
DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics
Silin Gao, Hao Zhao, Zeming Chen +8
Multimodal LLMs struggle to systematically model the temporal evolution of visual scenes in videos or multi-image sequences. Such inputs require models to predict or simulate multi…
Learning to Route Languages for Multilingual Policy Optimization
Geyang Guo, Hiromi Wakaki, Yuki Mitsufuji +2
Large language models~(LLMs) are trained on heterogeneous multilingual corpora, yet existing policy optimization methods often implicitly restrict each training question to a singl…
MusTBench: Benchmarking and Advancing Temporal Grounding in Music LLMs
Daeyong Kwon, Qiyu Wu, Shinobu Kuriya +6
Recent Large Audio-Language Models (LALMs) have demonstrated promising abilities in understanding musical content. However, whether their responses are grounded in the correct temp…