7 papers
X2SAM: Any Segmentation in Images and Videos
Hao Wang, Limeng Qiao, Chi Zhang +4
Multimodal Large Language Models (MLLMs) have demonstrated strong image-level visual understanding and reasoning, yet their pixel-level perception across both images and videos rem…
UniComp: Rethinking Video Compression Through Informational Uniqueness
Chao Yuan, Shimin Chen, Minliang Lin +3
Distinct from attention-based compression methods, this paper presents an information uniqueness driven video compression framework, termed UniComp, which aims to maximize the info…
STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning
Jie Qin, Jiancheng Huang, Limeng Qiao +1
Multimodal large language models (MLLMs) play a pivotal role in advancing the quest for general artificial intelligence. However, achieving unified target for multimodal understand…
VITRIX-CLIPIN: Enhancing Fine-Grained Visual Understanding in CLIP via Instruction Editing Data and Long Captions
Ziteng Wang, Siqi Yang, Limeng Qiao +1
Despite the success of Vision-Language Models (VLMs) like CLIP in aligning vision and language, their proficiency in detailed, fine-grained visual comprehension remains a key chall…
RIV: Recursive Introspection Mask Diffusion Vision Language Model
YuQian Li, Limeng Qiao, Lin Ma
Mask Diffusion-based Vision Language Models (MDVLMs) have achieved remarkable progress in multimodal understanding tasks. However, these models are unable to correct errors in gene…
Scalable Training for Vector-Quantized Networks with 100% Codebook Utilization
Yifan Chang, Jie Qin, Limeng Qiao +4
Vector quantization (VQ) is a key component in discrete tokenizers for image generation, but its training is often unstable due to straight-through estimation bias, one-step-behind…