25 citations · 53 across the 17 of their papers we have counts for
1 paper · 2 filters
Gregor Geigle, Chen Cecilia Liu, Jonas Pfeiffer +1
Current multimodal models, aimed at solving Vision and Language (V+L) tasks, predominantly repurpose Vision Encoders (VE) as feature extractors. While many VEs -- of different arch…