5 papers
ChainVLA: Chaining Vision-Language-Action Queries through a Unified Execution State for Long-Horizon Manipulation
Yuzhi Huang, Weijue Bu, Ziyi Xiong +4
Humans perform long-horizon manipulation by retaining knowledge of what earlier actions have established while continuously adapting the motion underway. By contrast, action-chunke…
RoboStream: Weaving Spatio-Temporal Reasoning with Memory in Vision-Language Models for Robotics
Yuzhi Huang, Jie Wu, Weijue Bu +9
Enabling reliable long-horizon robotic manipulation is a crucial step toward open-world embodied intelligence. However, VLM-based planners treat each step as an isolated observatio…
Conscious Gaze: Adaptive Attention Mechanisms for Hallucination Mitigation in Vision-Language Models
Weijue Bu, Guan Yuan, Guixian Zhang
Large Vision-Language Models (VLMs) often exhibit text inertia, where attention drifts from visual evidence toward linguistic priors, resulting in object hallucinations. Existing d…
Globally aware optimization with resurgence
Wei Bu
Modern optimization faces a fundamental challenge: local gradient-based methods provide no global information about the objective function landscape, often leading to suboptima…
Fokker-Planck to Callan-Symanzik: evolution of weight matrices under training
Wei Bu, Uri Kol, Ziming Liu
The dynamical evolution of a neural network during training has been an incredibly fascinating subject of study. First principal derivation of generic evolution of variables in sta…