8 papers
DriveGen3D: Boosting Feed-Forward Driving Scene Generation with Efficient Video Diffusion
Weijie Wang, Jiagang Zhu, Zeyu Zhang +14
We present DriveGen3D, a novel framework for generating high-quality and highly controllable dynamic 3D driving scenes that addresses critical limitations in existing methodologies…
GigaBrain-0: A World Model-Powered Vision-Language-Action Model
GigaBrain Team, Angen Ye, Boyuan Wang +24
Training Vision-Language-Action (VLA) models for generalist robots typically requires large-scale real-world robot data, which is expensive and time-consuming to collect. The ineff…
SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal Overhead
Chaojun Ni, Cheng Chen, Xiaofeng Wang +12
Vision-Language-Action (VLA) models built on pretrained Vision-Language Models (VLMs) show strong potential but are limited in practicality due to their large parameter counts. To…
GigaWorld-0: World Models as Data Engine to Empower Embodied AI
GigaWorld Team, Angen Ye, Boyuan Wang +22
World models are emerging as a foundational paradigm for scalable, data-efficient embodied AI. In this work, we present GigaWorld-0, a unified world model framework designed explic…
Motion-R1: Enhancing Motion Generation with Decomposed Chain-of-Thought and RL Binding
Runqi Ouyang, Haoyun Li, Zhenyuan Zhang +6
Text-to-Motion generation has become a fundamental task in human-machine interaction, enabling the synthesis of realistic human motions from natural language descriptions. Although…
MimicDreamer: Aligning Human and Robot Demonstrations for Scalable VLA Training
Haoyun Li, Ivan Zhang, Runqi Ouyang +12
Vision Language Action (VLA) models derive their generalization capability from diverse training data, yet collecting embodied robot interaction data remains prohibitively expensiv…