From the 1 of 63 linked papers with an AI index.
10 papers · 1 filter
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey
Bingzheng Qu, Kehai Chen, Xuefeng Bai +1
Recent progress in multimodal large language models (MLLMs) is reshaping video translation from a cascaded pipeline of automatic speech recognition, machine translation, text-to-sp…
Beyond Rigid: Benchmarking Non-Rigid Video Editing
Bingzheng Qu, Xuefeng Bai, Kehai Chen +1
As video generation models are increasingly expected to manipulate physical dynamics, there is a growing need to move evaluation beyond appearance fidelity and semantic alignment.…
LoCoT2V-Bench: Benchmarking Long-Form and Complex Text-to-Video Generation
Xiangqing Zheng, Chengyue Wu, Kehai Chen +1
Recent advances in text-to-video generation have achieved impressive performance on short clips, yet evaluating long-form generation under complex textual inputs remains a signific…
Decoupling Skeleton and Flesh: Efficient Multimodal Table Reasoning with Disentangled Alignment and Structure-aware Guidance
Yingjie Zhu, Xuefeng Bai, Kehai Chen +4
Reasoning over table images remains challenging for Large Vision-Language Models (LVLMs) due to complex layouts and tightly coupled structure-content information. Existing solution…
Through the Lens of Character: Resolving Modality-Role Interference in Multimodal Role-Playing Agent
Yihong Tang, Kehai Chen, Xuefeng Bai +1
The advancement of Multimodal Large Language Models (MLLMs) has expanded Role-Playing Agents (RPAs) into visually grounded environments. However, human vision is inherently subject…
Mitigating Multimodal Hallucination via Phase-wise Self-reward
Yu Zhang, Chuyang Sun, Kehai Chen +3
Large Vision-Language Models (LVLMs) still struggle with vision hallucination, where generated responses are inconsistent with the visual input. Existing methods either rely on lar…