10 papers
HelloWorld: Enabling Socially Interactive Characters in Video World Models
Liangyang Ouyang, Ruicong Liu, Xuangeng Chu +2
Despite the remarkable recent progress of video world models, social interaction between users and the characters within these worlds remains unsupported. To fill this gap, we pres…
Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?
Caixin Kang, Tianyu Yan, Sitong Gong +8
Multimodal Large Language Models (MLLMs) are increasingly deployed in human-facing roles where personality perception is critical, yet existing benchmarks evaluate this capability…
SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting
Ruicong Liu, Yifei Huang, Liangyang Ouyang +2
Real-time 3D hand forecasting is a critical component for fluid human-computer interaction in applications like AR and assistive robotics. However, existing methods are ill-suited…
SocialDirector: Training-Free Social Interaction Control for Multi-Person Video Generation
Liangyang Ouyang, Ruicong Liu, Caixin Kang +2
Video generation has advanced rapidly, producing photorealistic videos from text or image prompts. Meanwhile, film production and social robotics increasingly demand multi-person v…
Multi-speaker Attention Alignment for Multimodal Social Interaction
Liangyang Ouyang, Yifei Huang, Mingfang Zhang +3
Understanding social interaction in video requires reasoning over a dynamic interplay of verbal and non-verbal cues: who is speaking, to whom, and with what gaze or gestures. While…
Can MLLMs Read the Room? A Multimodal Benchmark for Assessing Deception in Multi-Party Social Interactions
Caixin Kang, Yifei Huang, Liangyang Ouyang +3
Despite their advanced reasoning capabilities, state-of-the-art Multimodal Large Language Models (MLLMs) demonstrably lack a core component of human intelligence: the ability to `r…