4 papers
ST-: Structured SpatioTemporal VLA for Robotic Manipulation
Chuanhao Ma, Hanyu Zhou, Shihan Peng +3
Vision-language-action (VLA) models have achieved great success on general robotic tasks, but still face challenges in fine-grained spatiotemporal manipulation. Typically, existing…
HoRD: Robust Humanoid Control via History-Conditioned Reinforcement Learning and Online Distillation
Puyue Wang, Jiawei Hu, Yan Gao +7
Humanoid robots can suffer significant performance drops under small changes in dynamics, task specifications, or environment setup. We propose HoRD, a two-stage learning framework…
Cog2Gen3D: Sculpturing 3D Semantic-Geometric Cognition for 3D Generation
Haonan Wang, Hanyu Zhou, Haoyue Liu +2
Generative models have achieved success in producing semantically plausible 2D images, but it remains challenging in 3D generation due to the absence of spatial geometry constraint…
LQA: A Lightweight Quantized-Adaptive Framework for Vision-Language Models on the Edge
Xin Wang, Hong Jia, Hualin Zhou +4
Deploying Vision-Language Models (VLMs) on edge devices is challenged by resource constraints and performance degradation under distribution shifts. While test-time adaptation (TTA…