7 papers
ReFree: Towards Realistic Co-Speech Video Generation via Reward-Free RL and Multilevel Speech Guidance
Salaheldin Mohamed, M. Hamza Mughal, Rishabh Dabral +1
Speech-driven talking character animation seeks to generate life-like portrait videos that convey natural conversation behavior, aligning facial motion with spoken audio. Although…
CoMoGen: COntrollable MOtion Dynamics and Interactions with Mask-Guided Video GENeration
Adil Meric, Lin Geng Foo, Mert Kiray +3
We present CoMoGen, a controllable video generation framework that generates realistic interactive dynamics from a single binary mask sequence conditioned on an input image. CoMoGe…
VHOI: Controllable Video Generation of Human-Object Interactions from Sparse Trajectories via Motion Densification
Wanyue Zhang, Lin Geng Foo, Thabo Beeler +2
Synthesizing realistic human-object interactions (HOI) in video is challenging due to the complex, instance-specific interaction dynamics of both humans and objects. Incorporating…
MIBURI: Towards Expressive Interactive Gesture Synthesis
M. Hamza Mughal, Rishabh Dabral, Vera Demberg +1
Embodied Conversational Agents (ECAs) aim to emulate human face-to-face interaction through speech, gestures, and facial expressions. Current large language model (LLM)-based conve…
Physical Simulator In-the-Loop Video Generation
Lin Geng Foo, Mark He Huang, Alexandros Lattas +3
Recent advances in diffusion-based video generation have achieved remarkable visual realism but still struggle to obey basic physical laws such as gravity, inertia, and collision.…
OnlineSplatter: Pose-Free Online 3D Reconstruction for Free-Moving Objects
Mark He Huang, Lin Geng Foo, Christian Theobalt +2
Free-moving object reconstruction from monocular video remains challenging, particularly without reliable pose or depth cues and under arbitrary object motion. We introduce OnlineS…