1 paper
Letitia Parcalabescu, Anette Frank
Vision and language model (VLM) decoders are currently the best-performing architectures on multimodal tasks. Next to answers, they are able to produce natural language explanation…