2 citations · 2 across the 32 of their papers we have counts for
30 papers · 1 filter
ReMoMask-2: Latent Retrieval-Augmented Masked Motion Generation
Yiran Wang, Zeyu Zhang, Ling Shao +1
Text-to-motion (T2M) generation maps natural language to human joint movements, aiding gaming, VR, and robotics. Retrieval-Augmented Text-to-Motion (RAG-T2M) improves generation on…
MobileVLA-R1 2.0: RL-Enhanced Reasoning for Mobile Robot Control
Ting Huang, Yue Huang, Zeyu Zhang +2
Grounding natural-language instructions into reliable and executable actions remains a fundamental challenge for vision-language-action (VLA) systems on mobile robots, due to the p…
ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning
Ting Huang, Zhenyu Zhang, Wenyuan Huang +2
Video spatial reasoning is essential for navigation-oriented perception and long-video question answering, where models must infer spatial relations across long horizons under chan…
GeneralVLA-2: Geometry-Aware Reconstruction and Governed Memory for Robot Planning
Haoyu Wang, Guoqing Ma, Zeyu Zhang +3
Generalist vision-language-action systems need object-centric 3D evidence and reusable manipulation experience to plan reliable robot trajectories. GeneralVLA provides a hierarchic…
MotionVLA: Vision-Language-Action Model for Humanoid Motion
Nonghai Zhang, Siyu Zhai, Yanjun Li +5
Generating realistic humanoid motion from scene images and text involves both low-frequency pose semantics and high-frequency physical dynamics. However, many existing methods toke…
PresentAgent-2: Towards Generalist Multimodal Presentation Agents
Wei Wu, Ziyang Xu, Zeyu Zhang +2
Presentation generation is moving beyond static slide creation toward end-to-end presentation video generation with research grounding, multimodal media, and interactive delivery.…