3 papers
cs.CV2026
On the Limits of Token Reduction for Efficient Unified Vision Language Training
Siyi Chen, Weiming Zhuang, Jingtao Li +1
Unified vision-language models (VLMs) integrate visual understanding and visual generation within a single autoregressive backbone, but their joint training is computationally expe…
cs.CV2026
VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations
Maitreya Patel, Jingtao Li, Weiming Zhuang +2
We introduce an efficient, resolution-agnostic autoregressive (AR) image synthesis approach that generalizes to arbitrary resolutions and aspect ratios, narrowing the gap to diffus…
cs.LG2026
IRIS: Implicit Reward-Guided Internal Sifting for Mitigating Multimodal Hallucination
Yuanshuai Li, Yuping Yan, Jirui Han +3
Hallucination remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While Direct Preference Optimization (DPO) is a key alignment framework, existing approa…