5 papers
STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning
Jie Qin, Jiancheng Huang, Limeng Qiao +1
Multimodal large language models (MLLMs) play a pivotal role in advancing the quest for general artificial intelligence. However, achieving unified target for multimodal understand…
RIV: Recursive Introspection Mask Diffusion Vision Language Model
YuQian Li, Limeng Qiao, Lin Ma
Mask Diffusion-based Vision Language Models (MDVLMs) have achieved remarkable progress in multimodal understanding tasks. However, these models are unable to correct errors in gene…
Scalable Training for Vector-Quantized Networks with 100% Codebook Utilization
Yifan Chang, Jie Qin, Limeng Qiao +4
Vector quantization (VQ) is a key component in discrete tokenizers for image generation, but its training is often unstable due to straight-through estimation bias, one-step-behind…
Beyond the Visible: Benchmarking Occlusion Perception in Multimodal Large Language Models
Zhaochen Liu, Kaiwen Gao, Shuyi Liang +4
Occlusion perception, a critical foundation for human-level spatial understanding, embodies the challenge of integrating visual recognition and reasoning. Though multimodal large l…
VITRIX-CLIPIN: Enhancing Fine-Grained Visual Understanding in CLIP via Instruction Editing Data and Long Captions
Ziteng Wang, Siqi Yang, Limeng Qiao +1
Despite the success of Vision-Language Models (VLMs) like CLIP in aligning vision and language, their proficiency in detailed, fine-grained visual comprehension remains a key chall…