6 papers
Black-Box Continual Learning for Vision-Language Models
Yuting Li, Weihang Fang, Haoyuan Gao +4
The rapid deployment of Vision-Language Models (VLMs) in dynamic environments necessitates the ability to learn continuously without forgetting. However, traditional continual lear…
IDER: IDempotent Experience Replay for Reliable Continual Learning
Zhanwang Liu, Yuting Li, Haoyuan Gao +4
Catastrophic forgetting, the tendency of neural networks to forget previously learned knowledge when learning new tasks, has been a major challenge in continual learning (CL). To t…
First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-Training
Lai Wei, Yuting Li, Chen Wang +4
Improving Multi-modal Large Language Models (MLLMs) in the post-training stage typically relies on supervised fine-tuning (SFT) or reinforcement learning (RL), which require expens…
Revisiting Visual Understanding in Multimodal Reasoning through a Lens of Image Perturbation
Yuting Li, Lai Wei, Kaipeng Zheng +6
Despite the rapid progress of multimodal large language models (MLLMs), they have largely overlooked the importance of visual processing. In a simple yet revealing experiment, we i…
Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start
Lai Wei, Yuting Li, Kaipeng Zheng +5
Recent advancements in large language models (LLMs) have demonstrated impressive chain-of-thought reasoning capabilities, with reinforcement learning (RL) playing a crucial role in…
Agentic Robot: A Brain-Inspired Framework for Vision-Language-Action Models in Embodied Agents
Zhejian Yang, Yongchao Chen, Xueyang Zhou +8
Long-horizon robotic manipulation poses significant challenges for autonomous systems, requiring extended reasoning, precise execution, and robust error recovery across complex seq…