2 papers
cs.LG2026
HBVLA: Pushing 1-Bit Post-Training Quantization for Vision-Language-Action Models
Xin Yan, Zhenglin Wan, Feiyang Ye +4
Vision-Language-Action (VLA) models enable instruction-following embodied control, but their large compute and memory footprints hinder deployment on resource-constrained robots an…
cs.CV2025
RapVerse: Coherent Vocals and Whole-Body Motions Generations from Text
Jiaben Chen, Xin Yan, Yihang Chen +7
In this work, we introduce a challenging task for simultaneously generating 3D holistic body motions and singing vocals directly from textual lyrics inputs, advancing beyond existi…