collaborators

7 papers

cs.RO2026

LIRA: Local Cross-Layer Information Routing for Vision-Language-Action Decoding

Zhewei Zhang, Puyue Wang, Guanren Qiao +10

Vision-Language-Action (VLA) models transform representations from pretrained vision-language models (VLMs) into robot actions, yet the interface that routes intermediate VLM featu…

cs.CV2026

XTransfer: Modality-Agnostic Few-Shot Model Transfer for Human Sensing at the Edge

Yu Zhang, Xi Zhang, Hualin Zhou +6

Deep learning for human sensing on edge systems presents significant potential for smart applications. However, its training and development are hindered by the limited availabilit…

cs.CV2026

ST-Gen4D: Embedding 4D Spatiotemporal Cognition into World Model for 4D Generation

Haonan Wang, Hanyu Zhou, Tao Gu +1

Generative models have achieved success in producing apparently coherent 2D videos, but remain challenging in the physical world due to lack of 4D spatiotemporal scale. Typically,…

cs.RO2026

ST-: Structured SpatioTemporal VLA for Robotic Manipulation

Chuanhao Ma, Hanyu Zhou, Shihan Peng +3

Vision-language-action (VLA) models have achieved great success on general robotic tasks, but still face challenges in fine-grained spatiotemporal manipulation. Typically, existing…

cs.RO2026

HoRD: Robust Humanoid Control via History-Conditioned Reinforcement Learning and Online Distillation

Puyue Wang, Jiawei Hu, Yan Gao +7

Humanoid robots can suffer significant performance drops under small changes in dynamics, task specifications, or environment setup. We propose HoRD, a two-stage learning framework…

cs.CV2026

Cog2Gen3D: Sculpturing 3D Semantic-Geometric Cognition for 3D Generation

Haonan Wang, Hanyu Zhou, Haoyue Liu +2

Generative models have achieved success in producing semantically plausible 2D images, but it remains challenging in 3D generation due to the absence of spatial geometry constraint…