1 paper · 1 filter
Yasmine Omri, Connor Ding, Tsachy Weissman +1
Modern vision language pipelines are driven by RGB vision encoders trained on massive image text corpora. While these pipelines have enabled impressive zero-shot capabilities and s…