9 papers
AVSCap: Orchestrating Audio-Visual Synergy for Omni-modal Video Captioning
Yanghai Wang, Jiahao Wang, Jiafu Tang +9
Omni-modal video captioning is not merely combining visual captioning with audio transcription: a useful caption must describe how visual actions, speech, music, and sound effects…
OmniCap-IF: Benchmarking and Improving Instruction Following Abilities for Omni-Video Captioning
Jiahao Wang, An Ping, Yanghai Wang +13
While Omni-modal Large Language Models (OLLMs) have demonstrated impressive capabilities in jointly processing audio and visual streams, their ability to strictly adhere to complex…
OmniHalluc-L: Counterfactual Benchmarking and Modality-Perturbation Reliability Calibration for Long-Form Omni Hallucination
Zixuan Dong, Jiafu Tang, Zhide Lei +7
Long-video Omni assistants often fail not by inventing content, but by misbinding real evidence: they hear the right utterance and see the right event, yet attach it to the wrong s…
DR-Eval: Towards Realistic and Reproducible Deep Research Evaluation
Qianqian Xie, Qingheng Xiong, He Zhu +16
Deep Research Agents (DRAs) aim to solve complex, long-horizon research tasks involving planning, retrieval, multimodal understanding, and report generation, yet their evaluation r…
T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation
Zhe Cao, Tao Wang, Jiaming Wang +10
Text-to-Audio-Video (T2AV) generation aims to synthesize temporally coherent video and semantically synchronized audio from natural language, yet its evaluation remains fragmented,…
MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
Tianhao Peng, Haochen Wang, Yuanxing Zhang +13
The advent of Multimodal Large Language Models (MLLMs) has expanded AI capabilities to visual modalities, yet existing evaluation benchmarks remain limited to single-video understa…