collaborators

6 papers

cs.CV2025

Multi-speaker Attention Alignment for Multimodal Social Interaction

Liangyang Ouyang, Yifei Huang, Mingfang Zhang +3

Understanding social interaction in video requires reasoning over a dynamic interplay of verbal and non-verbal cues: who is speaking, to whom, and with what gaze or gestures. While…

cs.CV2025

Can MLLMs Read the Room? A Multimodal Benchmark for Assessing Deception in Multi-Party Social Interactions

Caixin Kang, Yifei Huang, Liangyang Ouyang +3

Despite their advanced reasoning capabilities, state-of-the-art Multimodal Large Language Models (MLLMs) demonstrably lack a core component of human intelligence: the ability to `r…

cs.CV2025

Can MLLMs Read the Room? A Multimodal Benchmark for Verifying Truthfulness in Multi-Party Social Interactions

Caixin Kang, Yifei Huang, Liangyang Ouyang +2

As AI systems become increasingly integrated into human lives, endowing them with robust social intelligence has emerged as a critical frontier. A key aspect of this intelligence i…

cs.CV2025

LORE: Latent Optimization for Precise Semantic Control in Rectified Flow-based Image Editing

Liangyang Ouyang, Jiafeng Mao

Text-driven image editing enables users to flexibly modify visual content through natural language instructions, and is widely applied to tasks such as semantic object replacement,…

cs.CV2025

Leadership Assessment in Pediatric Intensive Care Unit Team Training

Liangyang Ouyang, Yuki Sakai, Ryosuke Furuta +3

This paper addresses the task of assessing PICU team's leadership skills by developing an automated analysis framework based on egocentric vision. We identify key behavioral cues,…

cs.CV2025

Egocentric Action-aware Inertial Localization in Point Clouds with Vision-Language Guidance

Mingfang Zhang, Ryo Yonetani, Yifei Huang +3

This paper presents a novel inertial localization framework named Egocentric Action-aware Inertial Localization (EAIL), which leverages egocentric action cues from head-mounted IMU…