From the 1 of 4 linked papers with an AI index.
4 papers
AVSCap: Orchestrating Audio-Visual Synergy for Omni-modal Video Captioning
Yanghai Wang, Jiahao Wang, Jiafu Tang +9
The paper introduces AVSCap, a system for omni-modal video captioning that explicitly binds visual and audio events, using a large tri-modal dataset and a two-stage training with r…
OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
Caorui Li, Yu Chen, Yiyan Ji +40
Recent advances in multimodal large language models (MLLMs) have demonstrated substantial potential in video understanding. However, existing benchmarks fail to comprehensively eva…
Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs
Zhiyu Pan, Yizheng Wu, Jiashen Hua +5
Reasoning has emerged as a key capability of large language models. In linguistic tasks, this capability can be enhanced by self-improving techniques that refine reasoning paths fo…
MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues
Yaning Pan, Qianqian Xie, Guohui Zhang +13
The recent development of Multimodal Large Language Models (MLLMs) has significantly advanced AI's ability to understand visual modalities. However, existing evaluation benchmarks…