3 papers
cs.CV2026
The Perceptual Bandwidth Bottleneck in Vision-Language Models: Active Visual Reasoning via Sequential Experimental Design
Anjie Liu, Ziqin Gong, Yan Song +6
Visual perception in modern Vision-Language Models (VLMs) is constrained by a perceptual bandwidth bottleneck: a broad field of view preserves global context but sacrifices the fin…
cs.CV2026
Chain-of-Models Pre-Training: Rethinking Training Acceleration of Vision Foundation Models
Jiawei Fan, Shigeng Wang, Chao Li +2
In this paper, we present Chain-of-Models Pre-Training (CoM-PT), a novel performance-lossless training acceleration method for vision foundation models (VFMs). This approach fundam…
cs.CV2025
JoyType: A Robust Design for Multilingual Visual Text Creation
Chao Li, Chen Jiang, Xiaolong Liu +2
Generating images with accurately represented text, especially in non-Latin languages, poses a significant challenge for diffusion models. Existing approaches, such as the integrat…