From the 2 of 5 linked papers with an AI index.
5 papers
ReToken: One Token to Improve Vision-Language Models for Visual Retrieval
Yao Xiao, Reuben Tan, Zhen Zhu +3
ReToken introduces a single learnable embedding that acts as a retrieval token to select a sparse set of relevant visual tokens from a cached representation, improving vision-langu…
AlphaWiSE: Adaptive Weight Interpolation for Continual Multimodal Representation Learning
Sarthak Jain, Qiran Hu, Zhen Zhu +1
The paper introduces AlphaWiSE, a post‑hoc weight‑space interpolation technique that combines two frozen checkpoints with learned scalar coefficients to improve continual learning…
Norm Anchors Make Model Edits Last
Mingda Liu, Zhenghan Zhu, Ze'an Miao +1
Sequential Locate-and-Edit (L&E) model editing can fail abruptly after many edits. We identify and formalize this failure as a positive norm-feedback loop, in which solved value ve…
How to Teach Large Multimodal Models New Skills
Zhen Zhu, Yiming Gong, Yao Xiao +2
How can we teach large multimodal models (LMMs) new skills without erasing prior abilities? We study sequential fine-tuning on five target skills while monitoring general ability o…
TextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models
Yao Xiao, Qiqian Fu, Heyi Tao +3
Image-text models excel at image-level tasks but struggle with detailed visual understanding. While these models provide strong visual-language alignment, segmentation models like…