1 paper · 1 filter
Fabian Paischer, Markus Hofmarcher, Sepp Hochreiter +1
Recently, vision-language models like CLIP have advanced the state of the art in a variety of multi-modal tasks including image captioning and caption evaluation. Many approaches l…