9 papers
Black-Box Continual Learning for Vision-Language Models
Yuting Li, Weihang Fang, Haoyuan Gao +4
The rapid deployment of Vision-Language Models (VLMs) in dynamic environments necessitates the ability to learn continuously without forgetting. However, traditional continual lear…
PANTHER: Generative Pretraining Beyond Language for Sequential User Behavior Modeling
Guilin Li, Yun Zhang, Xiuyuan Chen +6
Large language models (LLMs) have shown that generative pretraining can distill vast world knowledge into compact token representations. While LLMs encapsulate extensive world know…
Enhanced Continual Learning of Vision-Language Models with Model Fusion
Haoyuan Gao, Zicong Zhang, Yuqi Wei +6
Vision-Language Models (VLMs) represent a significant breakthrough in artificial intelligence by integrating visual and textual modalities to achieve impressive zero-shot capabilit…
IDER: IDempotent Experience Replay for Reliable Continual Learning
Zhanwang Liu, Yuting Li, Haoyuan Gao +4
Catastrophic forgetting, the tendency of neural networks to forget previously learned knowledge when learning new tasks, has been a major challenge in continual learning (CL). To t…
Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception
Lai Wei, Liangbo He, Jun Lan +9
Multimodal Large Language Models (MLLMs) excel at broad visual understanding but still struggle with fine-grained perception, where decisive evidence is small and easily overwhelme…
First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-Training
Lai Wei, Yuting Li, Chen Wang +4
Improving Multi-modal Large Language Models (MLLMs) in the post-training stage typically relies on supervised fine-tuning (SFT) or reinforcement learning (RL), which require expens…