3 papers
cs.CV2026
Foveated Reasoning: Stateful, Action-based Visual Focusing for Vision-Language Models
Juhong Min, Lazar Valkov, Vitali Petsiuk +2
Vision-language models benefit from high-resolution images, but the increase in visual-token count incurs high compute overhead. Humans resolve this tension via foveation: a coarse…
cs.CV2026
GoldiCLIP: The Goldilocks Approach for Balancing Explicit Supervision for Language-Image Pretraining
Deen Dayal Mohan, Hossein Souri, Vitali Petsiuk +4
Until recently, the success of large-scale vision-language models (VLMs) has primarily relied on billion-sample datasets, posing a significant barrier to progress. Latest works hav…
cs.CV2026
ARGENT: Adaptive Hierarchical Image-Text Representations
Chuong Huynh, Hossein Souri, Abhinav Kumar +3
Large-scale Vision-Language Models (VLMs) such as CLIP learn powerful semantic representations but operate in Euclidean space, which fails to capture the inherent hierarchical stru…