3 papers
cs.CV2026
InteracVid: Building a Real Interactive Audio-Visual Response Dataset from Live-Chat Videos
Chi Zhang, Haoyang Shi, Yueyi Liu +4
Large language models have made text the default medium for human--AI interaction, buttext alone cannot express the full range of responses required by multimodal assistants,avatar…
cs.CL2026
MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings
Ziyi Wang, Yuhang Wu, Dongxu Piao +3
Theory of Mind (ToM), the ability to infer other's beliefs, intentions, and states of knowledge, is central to social interaction, yet remains challenging for current Multimodal La…
cs.CV2026
PIPBench: A Profile-Inclusive Framework for Personalized Image Generation Evaluation
Yuhang Wu, Shuxiang Zhang, Wee Hian Ching +2
Recent text-to-image models such as DALLE-3 excel at following diverse prompts yet remain blind to individual aesthetic preferences. We study personalized image generation, where m…