6 papers
DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models
Zhuoming Liu, Jinhong Lin, Kwan Man Cheng +3
Many modern vision-language models (VLMs) build on autoregressive decoding of discrete tokens. While text-based output interfaces enable scalable pretraining and strong zero-shot g…
MEval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasks
Jie Huang, Ruixun Liu, Sirui Sun +4
As multi-modal models advance towards long-form video understanding, memory emerges as a critical capability. Despite substantial efforts in developing video datasets and benchmark…
DUALVISION: RGB-Infrared Multimodal Large Language Models for Robust Visual Reasoning
Abrar Majeedi, Zhiyuan Ruan, Ziyi Zhao +3
Multimodal large language models (MLLMs) have achieved impressive performance on visual perception and reasoning tasks with RGB imagery, yet they remain fragile under common degrad…
Audio-Enhanced Vision-Language Modeling with Latent Space Broadening for High Quality Data Expansion
Yu Sun, Yin Li, Ruixiao Sun +7
Transformer-based multimodal models are widely used in industrial-scale recommendation, search, and advertising systems for content understanding and relevance ranking. Enhancing l…
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning
Yiwu Zhong, Zhuoming Liu, Yin Li +1
Large language models (LLMs) have enabled the creation of multi-modal LLMs that exhibit strong comprehension of visual data such as images and videos. However, these models usually…
PAVE: Patching and Adapting Video Large Language Models
Zhuoming Liu, Yiquan Li, Khoi Duc Nguyen +2
Pre-trained video large language models (Video LLMs) exhibit remarkable reasoning capabilities, yet adapting these models to new tasks involving additional modalities or data types…