7 papers
LIRA: Local Cross-Layer Information Routing for Vision-Language-Action Decoding
Zhewei Zhang, Puyue Wang, Guanren Qiao +10
Vision-Language-Action (VLA) models transform representations from pretrained vision-language models (VLMs) into robot actions, yet the interface that routes intermediate VLM featu…
XTransfer: Modality-Agnostic Few-Shot Model Transfer for Human Sensing at the Edge
Yu Zhang, Xi Zhang, Hualin Zhou +6
Deep learning for human sensing on edge systems presents significant potential for smart applications. However, its training and development are hindered by the limited availabilit…
ST-Gen4D: Embedding 4D Spatiotemporal Cognition into World Model for 4D Generation
Haonan Wang, Hanyu Zhou, Tao Gu +1
Generative models have achieved success in producing apparently coherent 2D videos, but remain challenging in the physical world due to lack of 4D spatiotemporal scale. Typically,…
ST-: Structured SpatioTemporal VLA for Robotic Manipulation
Chuanhao Ma, Hanyu Zhou, Shihan Peng +3
Vision-language-action (VLA) models have achieved great success on general robotic tasks, but still face challenges in fine-grained spatiotemporal manipulation. Typically, existing…
HoRD: Robust Humanoid Control via History-Conditioned Reinforcement Learning and Online Distillation
Puyue Wang, Jiawei Hu, Yan Gao +7
Humanoid robots can suffer significant performance drops under small changes in dynamics, task specifications, or environment setup. We propose HoRD, a two-stage learning framework…
Cog2Gen3D: Sculpturing 3D Semantic-Geometric Cognition for 3D Generation
Haonan Wang, Hanyu Zhou, Haoyue Liu +2
Generative models have achieved success in producing semantically plausible 2D images, but it remains challenging in 3D generation due to the absence of spatial geometry constraint…