collaborators

15 papers

cs.CV2026

InteracVid: Building a Real Interactive Audio-Visual Response Dataset from Live-Chat Videos

Chi Zhang, Haoyang Shi, Yueyi Liu +4

Large language models have made text the default medium for human--AI interaction, buttext alone cannot express the full range of responses required by multimodal assistants,avatar…

cs.CV2026

ElasticTTT: Prior-Preserving Test-Time Tuning for Video Editing

Yueyi Liu, Chi Zhang, Sen Cui +1

Test-Time Tuning (TTT) on pretrained diffusion models has emerged as a powerful paradigm for video editing. However, there exists a foundational mismatch between the distribution-m…

cs.CL2026

MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings

Ziyi Wang, Yuhang Wu, Dongxu Piao +3

Theory of Mind (ToM), the ability to infer other's beliefs, intentions, and states of knowledge, is central to social interaction, yet remains challenging for current Multimodal La…

cs.CV2026

PIPBench: A Profile-Inclusive Framework for Personalized Image Generation Evaluation

Yuhang Wu, Shuxiang Zhang, Wee Hian Ching +2

Recent text-to-image models such as DALLE-3 excel at following diverse prompts yet remain blind to individual aesthetic preferences. We study personalized image generation, where m…

cs.CV2026

EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding

Yijia Lei, Jinzhao Li, Yichi Zhang +3

We introduce EgoSAT, the first comprehensive benchmark for egocentric video reasoning in streaming settings, designed to evaluate the capabilities of modern vision-language models…

cs.CL2026

PIVOTSBench: Evaluating Fine-Grained Interpersonal Relationship Reasoning in Multimodal Large Language Models

Shuxiang Zhang, Yiting Yin, Wenxuan Song +2

Humans possess an innate ability to understand fine-grained interpersonal relationships, which is central to everyday social interactions. Although such reasoning is inherently mul…