3 citations · 3 across the 4 of their papers we have counts for
4 papers · 1 filter
MobileCLIP2: Improving Multi-Modal Reinforced Training
Fartash Faghri, Pavan Kumar Anasosalu Vasu, Cem Koc +4
Foundation image-text models such as CLIP with zero-shot capabilities enable a wide array of applications. MobileCLIP is a recent family of image-text models at 3-15ms latency and…
FocalLens: Instruction Tuning Enables Zero-Shot Conditional Image Representations
Cheng-Yu Hsieh, Pavan Kumar Anasosalu Vasu, Fartash Faghri +5
Visual understanding is inherently contextual -- what we focus on in an image depends on the task at hand. For instance, given an image of a person holding a bouquet of flowers, we…
FastVLM: Efficient Vision Encoding for Vision Language Models
Pavan Kumar Anasosalu Vasu, Fartash Faghri, Chun-Liang Li +8
Scaling the input image resolution is essential for enhancing the performance of Vision Language Models (VLMs), particularly in text-rich image understanding tasks. However, popula…
Instance-Level Task Parameters: A Robust Multi-task Weighting Framework
Pavan Kumar Anasosalu Vasu, Shreyas Saxena, Oncel Tuzel
Recent works have shown that deep neural networks benefit from multi-task learning by learning a shared representation across several related tasks. However, performance of such sy…