8 papers · 1 filter
Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
Jingjing Jiang, Chongjie Si, Jun Luo +2
This paper presents a pioneering exploration of reinforcement learning (RL) via group relative policy optimization for unified multimodal large language models (ULMs), aimed at sim…
More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models
Xurui Song, Shuo Huai, JingJing Jiang +2
Vision-Language Model (VLM) driving agents promise explainable end-to-end autonomy by first producing natural-language reasoning and then predicting trajectory planning. However, w…
Rehearsal-free and Task-free Online Continual Learning With Contrastive Prompt
Aopeng Wang, Ke Deng, Yongli Ren +1
The main challenge of continual learning is \textit{catastrophic forgetting}. Because of processing data in one pass, online continual learning (OCL) is one of the most difficult c…
Reinitializing weights vs units for maintaining plasticity in neural networks
J. Fernando Hernandez-Garcia, Shibhansh Dohare, Jun Luo +1
Loss of plasticity is a phenomenon in which a neural network loses its ability to learn when trained for an extended time on non-stationary data. It is a crucial problem to overcom…
Corvid: Improving Multimodal Large Language Models Towards Chain-of-Thought Reasoning
Jingjing Jiang, Chao Ma, Xurui Song +2
Recent advancements in multimodal large language models (MLLMs) have demonstrated exceptional performance in multimodal perception and understanding. However, leading open-source M…
A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations
Tian Lan, Yang-Hao Zhou, Zi-Ao Ma +8
Recent advances in deep learning have significantly enhanced generative AI capabilities across text, images, and audio. However, automatically evaluating the quality of these gener…