4 papers
GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions
Junho Kim, Xu Cao, Houze Yang +6
Understanding social interactions requires reasoning over subtle non-verbal cues, yet current multimodal large language models (MLLMs) often fail to identify who interacts with who…
MuCo: Multi-turn Contrastive Learning for Multimodal Embedding Model
Geonmo Gu, Byeongho Heo, Jaemyung Yu +7
Universal Multimodal embedding models built on Multimodal Large Language Models (MLLMs) have traditionally employed contrastive learning, which aligns representations of query-targ…
Unified Text-Image-to-Video Generation: A Training-Free Approach to Flexible Visual Conditioning
Bolin Lai, Sangmin Lee, Xu Cao +2
Text-image-to-video (TI2V) generation is a critical problem for controllable video generation using both semantic and visual conditions. Most existing methods typically add visual…
GaussianMotion: End-to-End Learning of Animatable Gaussian Avatars with Pose Guidance from Text
Gyumin Shim, Sangmin Lee, Jaegul Choo
In this paper, we introduce GaussianMotion, a novel human rendering model that generates fully animatable scenes aligned with textual descriptions using Gaussian Splatting. Althoug…