17 papers
OmniVTLA: Vision-Tactile-Language-Action Models with Semantic-Aligned Tactile Sensing
Zhengxue Cheng, Yiqian Zhang, Anni Tang +5
Recent vision-language-action (VLA) models build upon vision-language foundations, and have achieved promising results and exhibit the possibility of task generalization in robot m…
GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training
Yuelin Hu, Zhenbo Yu, Zhengxue Cheng +2
Hybrid post-training usually combines supervised fine-tuning and reinforcement learning, but fixed mixing schedules cannot adapt when the relative noise of the two signals changes…
Full-4D: Generating Full-Scope 4D Scenes from a Single-View Video
Tingxi Chen, Ke Hao, Yabo Chen +6
Generating 4D scenes from a single-view video is inherently ill-posed: a single viewpoint lacks the information needed to recover a complete, dynamic scene with full coverage. Exis…
Hidden Failure Modes of Gradient Modification under Adam in Continual Learning, and Adaptive Decoupled Moment Routing as a Repair
Yuelin Hu, Zhenbo Yu, Zhengxue Cheng +2
Many continual-learning methods modify gradients upstream (e.g., projection, penalty rescaling, replay mixing) while treating Adam as a neutral backend. We show this composition ha…
TaCo: A Benchmark for Lossless and Lossy Codecs of Heterogeneous Tactile Data
Zhengxue Cheng, Yan Zhao, Keyu Wang +2
Tactile sensing is crucial for embodied intelligence, providing fine-grained perception and control in complex environments. However, efficient tactile data compression, which is e…
WebExpert: domain-aware web agents with critic-guided expert experience for high-precision search
Yuelin Hu, Zhengxue Cheng, Ronghua Wu +5
Specialized web tasks in finance, biomedicine, and pharmaceuticals remain challenging due to missing domain priors: queries drift, evidence is noisy, and reasoning is brittle. We p…