12 papers
Vesta: A Generalist Embodied Reasoning Model
Johan Bjorck, Zhiqi Li, Yunze Man +29
Robots operating in open-world environments must seamlessly integrate localization, spatial reasoning, navigation, and long-horizon planning. While specialist models excel at indiv…
NextMotionQA: Benchmarking and Judging Human Motion Understanding with Vision-Language Models
Yong Cao, Chuqiao Li, Xianghui Xie +2
Reliable evaluation of human motion understanding is fundamental to advancing embodied AI, robotics, and animation. However, existing benchmarks suffer from coarse semantic granula…
CARI4D: Category Agnostic 4D Reconstruction of Human-Object Interaction
Xianghui Xie, Bowen Wen, Yan Chang +5
Accurate capture of human-object interaction from ubiquitous sensors like RGB cameras is important for applications in human understanding, gaming, and robot learning. However, inf…
ActionPlan: Future-Aware Streaming Motion Synthesis via Frame-Level Action Planning
Eric Nazarenus, Chuqiao Li, Yannan He +3
We present ActionPlan, a unified motion diffusion framework that bridges real-time streaming with high-quality offline generation within a single model. The core idea is to introdu…
Hoi3DGen: Generating High-Quality Human-Object-Interactions in 3D
Agniv Sharma, Xianghui Xie, Tom Fischer +2
Modeling and generating 3D human-object interactions from text is crucial for applications in AR, XR, and gaming. Existing approaches often rely on score distillation from text-to-…
FrankenMotion: Part-level Human Motion Generation and Composition
Chuqiao Li, Xianghui Xie, Yong Cao +2
Human motion generation from text prompts has made remarkable progress in recent years. However, existing methods primarily rely on either sequence-level or action-level descriptio…