1 paper
Mitja Nikolaus, Emmanuelle Salin, Stephane Ayache +2
Recent advances in vision-and-language modeling have seen the development of Transformer architectures that achieve remarkable performance on multimodal reasoning tasks. Yet, the e…