6 papers
DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation
Haozhe Xie, Beichen Wen, Jiarui Zheng +4
Manipulating dynamic objects remains an open challenge for Vision-Language-Action (VLA) models, which, despite strong generalization in static manipulation, struggle in dynamic sce…
4DNeX: Feed-Forward 4D Generative Modeling Made Easy
Zhaoxi Chen, Tianqi Liu, Long Zhuo +6
We present 4DNeX, the first feed-forward framework for generating 4D (i.e., dynamic 3D) scene representations from a single image. In contrast to existing methods that rely on comp…
EgoTwin: Dreaming Body and View in First Person
Jingqiao Xiu, Fangzhou Hong, Yicong Li +5
While exocentric video synthesis has achieved great progress, egocentric video generation remains largely underexplored, which requires modeling first-person view content along wit…
Compositional Generative Model of Unbounded 4D Cities
Haozhe Xie, Zhaoxi Chen, Fangzhou Hong +1
3D scene generation has garnered growing attention in recent years and has made significant progress. Generating 4D cities is more challenging than 3D scenes due to the presence of…
GeneMAN: Generalizable Single-Image 3D Human Reconstruction from Multi-Source Human Data
Wentao Wang, Hang Ye, Fangzhou Hong +5
Given a single in-the-wild human photo, it remains a challenging task to reconstruct a high-fidelity 3D human model. Existing methods face difficulties including a) the varying bod…
EgoLM: Multi-Modal Language Model of Egocentric Motions
Fangzhou Hong, Vladimir Guzov, Hyo Jin Kim +4
As the prevalence of wearable devices, learning egocentric motions becomes essential to develop contextual AI. In this work, we present EgoLM, a versatile framework that tracks and…