Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
VSAS-Bench: Real-Time Evaluation of Visual Streaming Assistant Models
Pavan Kumar Anasosalu Vasu, Cem Koc, Fartash Faghri +6
Streaming vision-language models (VLMs) continuously generate responses given an instruction prompt and an online stream of input frames. This is a core mechanism for real-time vis…
cs.CV2025
MobileCLIP2: Improving Multi-Modal Reinforced Training
Fartash Faghri, Pavan Kumar Anasosalu Vasu, Cem Koc +4
Foundation image-text models such as CLIP with zero-shot capabilities enable a wide array of applications. MobileCLIP is a recent family of image-text models at 3-15ms latency and…
cs.CV2025
FastVLM: Efficient Vision Encoding for Vision Language Models
Pavan Kumar Anasosalu Vasu, Fartash Faghri, Chun-Liang Li +8
Scaling the input image resolution is essential for enhancing the performance of Vision Language Models (VLMs), particularly in text-rich image understanding tasks. However, popula…