6 papers
GeneralVLA: Generalizable Vision-Language-Action Models with Knowledge-Guided Trajectory Planning
Guoqing Ma, Siheng Wang, Zeyu Zhang +2
Large foundation models have shown strong open-world generalization to complex problems in vision and language, but similar levels of generalization have yet to be achieved in robo…
Clustering-Based Weight Orthogonalization for Stabilizing Deep Reinforcement Learning
Guoqing Ma, Yuhan Zhang, Yuming Dai +3
Reinforcement learning (RL) has made significant advancements, achieving superhuman performance in various tasks. However, RL agents often operate under the assumption of environme…
Multi-dimensional Neural Decoding with Orthogonal Representations for Brain-Computer Interfaces
Kaixi Tian, Shengjia Zhao, Yuhan Zhang +1
Current brain-computer interfaces primarily decode single motor variables, limiting their ability to support natural, high-bandwidth neural control that requires simultaneous extra…
Visual Large Language Models Exhibit Human-Level Cognitive Flexibility in the Wisconsin Card Sorting Test
Guangfu Hao, Frederic Alexandre, Shan Yu
Cognitive flexibility has been extensively studied in human cognition but remains relatively unexplored in the context of Visual Large Language Models (VLLMs). This study assesses…
Flexible Tool Selection through Low-dimensional Attribute Alignment of Vision and Language
Guangfu Hao, Haojie Wen, Liangxuan Guo +3
Flexible tool selection reflects a complex cognitive ability that distinguishes humans from other species, yet computational models that capture this ability remain underdeveloped.…
Efficient Reinforcement Learning Through Adaptively Pretrained Visual Encoder
Yuhan Zhang, Guoqing Ma, Guangfu Hao +3
While Reinforcement Learning (RL) agents can successfully learn to handle complex tasks, effectively generalizing acquired skills to unfamiliar settings remains a challenge. One of…