4 papers
GameLogicBench: Evaluating Coding Agents on Runtime Game Logic with Tick-Level State Assertions
Xinyu Che, Yunfei Ge, Shihao Li +9
Coding agents can modify and test code across large software projects. Game development is a domain where agents must implement gameplay rules. A game can end in a valid state even…
AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities
Yuqing Wen, Yukai Huang, Qianqian Xie +6
While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modality often requires coordinated…
IF-VidCap: Can Video Caption Models Follow Instructions?
Shihao Li, Yuanxing Zhang, Jiangtao Wu +20
Although Multimodal Large Language Models (MLLMs) have demonstrated proficiency in video captioning, practical applications require captions that follow specific user instructions…
MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues
Yaning Pan, Qianqian Xie, Guohui Zhang +13
The recent development of Multimodal Large Language Models (MLLMs) has significantly advanced AI's ability to understand visual modalities. However, existing evaluation benchmarks…