4 papers
GIRAF: Towards Generalizable Human Interactions with Articulated Objects
Xiaohan Zhang, Sebastian Starke, Alexander Winkler +3
Synthesizing realistic full-body human interactions with articulated objects is a fundamental challenge for embodied AI and graphics, with applications in robotics training and vir…
AeSlides: Incentivizing Aesthetic Layout in LLM-Based Slide Generation via Verifiable Rewards
Yiming Pan, Chengwei Hu, Xuancheng Huang +6
Large language models (LLMs) have demonstrated strong potential in agentic tasks, particularly in slide generation. However, slide generation poses a fundamental challenge: the gen…
MoLingo: Motion-Language Alignment for Text-to-Human Motion Generation
Yannan He, Garvita Tiwari, Xiaohan Zhang +4
We introduce MoLingo, a text-to-motion (T2M) model that generates realistic, lifelike human motion by denoising in a continuous latent space. Recent works perform latent space diff…
SCENIC: Scene-aware Semantic Navigation with Instruction-guided Control
Xiaohan Zhang, Sebastian Starke, Vladimir Guzov +3
Synthesizing natural human motion that adapts to complex environments while allowing creative control remains a fundamental challenge in motion synthesis. Existing models often fal…