17 papers
Dynamic Execution Commitment of Vision-Language-Action Models
Feng Chen, Xianghui Wang, Yuxuan Chen +4
Vision-Language-Action (VLA) models predominantly adopt action chunking, i.e., predicting and committing to a short horizon of consecutive low-level actions in a single forward pas…
GeneralVLA-2: Geometry-Aware Reconstruction and Governed Memory for Robot Planning
Haoyu Wang, Guoqing Ma, Zeyu Zhang +3
Generalist vision-language-action systems need object-centric 3D evidence and reusable manipulation experience to plan reliable robot trajectories. GeneralVLA provides a hierarchic…
SpatialAvatar-0: High-Quality 4D Head Avatar with Multi-Stage Reconstruction
Yiran Wang, Zeyu Zhang, Yuanming Li +2
High-quality 4D head avatars from one or a few source portraits are central to telepresence, AR/VR, and digital-human interaction. 3D Gaussian Splatting (3DGS) has emerged as the d…
MotionVLA: Vision-Language-Action Model for Humanoid Motion
Nonghai Zhang, Siyu Zhai, Yanjun Li +5
Generating realistic humanoid motion from scene images and text involves both low-frequency pose semantics and high-frequency physical dynamics. However, many existing methods toke…
DragMesh-2: Physically Plausible Dexterous Hand-Object Interaction with Articulated Objects
Tianshan Zhang, Yijia Duan, Yanjun Li +2
Dexterous interaction with articulated objects is important for household, assistive, and humanoid manipulation, where multi-finger hands can provide compliant contact patterns bey…
PresentAgent-2: Towards Generalist Multimodal Presentation Agents
Wei Wu, Ziyang Xu, Zeyu Zhang +2
Presentation generation is moving beyond static slide creation toward end-to-end presentation video generation with research grounding, multimodal media, and interactive delivery.…