15 papers
InteracVid: Building a Real Interactive Audio-Visual Response Dataset from Live-Chat Videos
Chi Zhang, Haoyang Shi, Yueyi Liu +4
Large language models have made text the default medium for human--AI interaction, buttext alone cannot express the full range of responses required by multimodal assistants,avatar…
ElasticTTT: Prior-Preserving Test-Time Tuning for Video Editing
Yueyi Liu, Chi Zhang, Sen Cui +1
Test-Time Tuning (TTT) on pretrained diffusion models has emerged as a powerful paradigm for video editing. However, there exists a foundational mismatch between the distribution-m…
MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings
Ziyi Wang, Yuhang Wu, Dongxu Piao +3
Theory of Mind (ToM), the ability to infer other's beliefs, intentions, and states of knowledge, is central to social interaction, yet remains challenging for current Multimodal La…
PIPBench: A Profile-Inclusive Framework for Personalized Image Generation Evaluation
Yuhang Wu, Shuxiang Zhang, Wee Hian Ching +2
Recent text-to-image models such as DALLE-3 excel at following diverse prompts yet remain blind to individual aesthetic preferences. We study personalized image generation, where m…
EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding
Yijia Lei, Jinzhao Li, Yichi Zhang +3
We introduce EgoSAT, the first comprehensive benchmark for egocentric video reasoning in streaming settings, designed to evaluate the capabilities of modern vision-language models…
PIVOTSBench: Evaluating Fine-Grained Interpersonal Relationship Reasoning in Multimodal Large Language Models
Shuxiang Zhang, Yiting Yin, Wenxuan Song +2
Humans possess an innate ability to understand fine-grained interpersonal relationships, which is central to everyday social interactions. Although such reasoning is inherently mul…