13 papers
ChainVLA: Chaining Vision-Language-Action Queries through a Unified Execution State for Long-Horizon Manipulation
Yuzhi Huang, Weijue Bu, Ziyi Xiong +4
Humans perform long-horizon manipulation by retaining knowledge of what earlier actions have established while continuously adapting the motion underway. By contrast, action-chunke…
RoboStream: Weaving Spatio-Temporal Reasoning with Memory in Vision-Language Models for Robotics
Yuzhi Huang, Jie Wu, Weijue Bu +9
Enabling reliable long-horizon robotic manipulation is a crucial step toward open-world embodied intelligence. However, VLM-based planners treat each step as an isolated observatio…
What Drives Test-Time Adaptation for CLIP? A Controlled Empirical Study from an Update Perspective
Jiazhen Huang, Xiao Chen, Zhiming Liu +3
Vision-Language Models (VLMs) such as CLIP have become a standard backbone for open-vocabulary recognition, yet their zero-shot predictions remain vulnerable to distribution shifts…
Test-Time Distillation for Continual Model Adaptation
Xiao Chen, Jiazhen Huang, Zhiming Liu +4
Deep neural networks often suffer performance degradation upon deployment due to distribution shifts. Continual Test-Time Adaptation (CTTA) aims to address this issue in an unsuper…
Semantic-Space Exploration and Exploitation in RLVR for LLM Reasoning
Fanding Huang, Guanbo Huang, Xiao Fan +7
Reinforcement Learning with Verifiable Rewards (RLVR) for LLM reasoning is often framed as balancing exploration and exploitation in action space, typically operationalized with to…
Neural Collapse in Test-Time Adaptation
Xiao Chen, Zhongjing Du, Jiazhen Huang +4
Test-Time Adaptation (TTA) enhances model robustness to out-of-distribution (OOD) data by updating the model online during inference, yet existing methods lack theoretical insights…