8 papers
Miles: Metric Learning with Expandable Subspace for Pre-Trained Model-Based Class-Incremental Learning
Kai Jiang, Zisong Lin, Hongyuan Zhang +2
Class Incremental Learning (CIL) aims to learn new concepts consistently from a data stream without forgetting. Unlike typical CIL methods which need to learn a model from scratch,…
Latent Noise Mask for Reducing Visual Redundancy in Multimodal Large Language Models
Kai Jiang, Ruishu Zhu, Siqi Huang +2
Multimodal large language models (MLLMs) often fail in fine-grained visual reasoning, as question-relevant visual cues are diluted by dense and redundant image tokens. Recent multi…
ViewMask-1-to-3: Multi-View Consistent Image Generation via Multimodal Discrete Diffusion Models
Ruishu Zhu, Zhihao Huang, Jiacheng Sun +3
Motivated by discrete diffusion's success in language-vision modeling, we explore its potential for multi-view generation, a task dominated by continuous approaches. We introduce V…
Dynamic Mixture of Progressive Parameter-Efficient Expert Library for Lifelong Robot Learning
Yuheng Lei, Sitong Mao, Shunbo Zhou +3
A generalist agent must continuously learn and adapt throughout its lifetime, achieving efficient forward transfer while minimizing catastrophic forgetting. Previous work within th…
Crab: A Scalable and Unified Audio-Visual Scene Understanding Model with Explicit Cooperation
Dongnuan Cai, Henghui Du, Chang Zhou +5
Developing Audio-Visual Large Language Models (AV-LLMs) for unified scene understanding is pivotal in multimodal intelligence. While instruction tuning enables pre-trained models w…
Object-AVEdit: An Object-level Audio-Visual Editing Model
Youquan Fu, Ruiyang Si, Hongfa Wang +6
There is a high demand for audio-visual editing in video post-production and the film making field. While numerous models have explored audio and video editing, they struggle with…