9 papers
Homer: Understanding Long-form Videos with Hierarchical Memory and Agentic Reasoning
Yixin Ji, Fanghua Ye, Juntao Li +5
Multimodal large language models excel on short clips but struggle on hour-long videos in an online setting, where frames are processed incrementally under limited memory. Existing…
LivingArena: Do LLMs Know What Other LLMs Don't? Peer-Probing as Scalable Evaluation
Xingyu Chen, Rui Wang, Zhaopeng Tu +1
Evaluating frontier LLMs is challenging: static benchmarks suffer from contamination and saturation -- leaving users unable to distinguish top models and developers blind to specif…
MMAE: A Massive Multitask Audio Editing Benchmark
Ziyang Ma, Ruiqi Yan, Ruiyang Xu +35
We introduce MMAE, a Massive Multitask Audio Editing benchmark, serving as the first comprehensive evaluation testbed designed for general-purpose instruction-based audio editing.…
HYDRA: Unifying Multi-modal Generation and Understanding via Representation-Harmonized Tokenization
Xuerui Qiu, Yutao Cui, Guozhen Zhang +9
Unified Multimodal Models struggle to bridge the fundamental gap between the abstract representations needed for visual understanding and the detailed primitives required for gener…
Too Good to be Bad: On the Failure of LLMs to Role-Play Villains
Zihao Yi, Qingxuan Jiang, Ruotian Ma +8
Large Language Models (LLMs) are increasingly tasked with creative generation, including the simulation of fictional characters. However, their ability to portray non-prosocial, an…
EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion Generation
Xiangyue Zhang, Jianfang Li, Jiaxu Zhang +3
Masked modeling has shown promise in co-speech gesture generation. However, it struggles to identify semantically significant frames for effective motion masking. In this work, we…