1 paper · 1 filter
Max Hartman, Vidhata Jayaraman, Moulik Choraria +2
Vision-language models achieve incredible performance across a wide range of tasks, but their large size makes inference costly. Recent work has shown that multimodal processing co…