7 papers
Attacking the Trusted Imagination: Oracle-Level Integrity Attacks on Imagine-then-Act World Models
Linghan Chen, Kaiyan Ji, Minyu Guo
Many recent vision-language-action (VLA) policies adopt an imagine-then-act design. A world-action model (WAM) first imagines a short future as a latent trajectory z~, on which the…
DiscoForcing: A Unified Framework for Real-Time Audio-Driven Character Control with Diffusion Forcing
Kaiyang Ji, Bingsheng Qian, Binghuan Wu +3
We study real-time audio-responsive character control as a deployment-faithful problem: strictly causal, bounded-latency streaming that must generate coherent full-body motion at i…
DreamPolicy: A Unified World-model Policy for Scalable Humanoid Locomotion
Yahao Fan, Tianxiang Gui, Kaiyang Ji +8
Achieving versatile humanoid locomotion with a single policy presents a critical scalability challenge. Prevailing methods often rely on distilling multiple terrain-specific teache…
Commanding Humanoid by Free-form Language: A Large Language Action Model with Unified Motion Vocabulary
Zhirui Liu, Kaiyang Ji, Ke Yang +4
Enabling humanoid robots to follow free-form natural language commands is a critical step toward seamless human-robot interaction and general-purpose embodied AI. However, existing…
Human-Object Interaction via Automatically Designed VLM-Guided Motion Policy
Zekai Deng, Ye Shi, Kaiyang Ji +3
Human-object interaction (HOI) synthesis is crucial for applications in animation, simulation, and robotics. However, existing approaches either rely on expensive motion capture da…
Towards Immersive Human-X Interaction: A Real-Time Framework for Physically Plausible Motion Synthesis
Kaiyang Ji, Ye Shi, Zichen Jin +5
Real-time synthesis of physically plausible human interactions remains a critical challenge for immersive VR/AR systems and humanoid robotics. While existing methods demonstrate pr…