most citedRetrieval-augmented Image Captioning

4 citations · 5 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CL2024

Sequential Compositional Generalization in Multimodal Models

Semih Yagcioglu, Osman Batur İnce, Aykut Erdem +3

The rise of large-scale multimodal models has paved the pathway for groundbreaking advances in generative modeling and reasoning, unlocking transformative applications in a variety…

cs.CL2023

PHD: Pixel-Based Language Modeling of Historical Documents

Nadav Borenstein, Phillip Rust, Desmond Elliott +1

The digitisation of historical documents has provided historians with unprecedented research opportunities. Yet, the conventional approach to analysing historical documents involve…

cs.CL2023

Text Rendering Strategies for Pixel Language Models

Jonas F. Lotz, Elizabeth Salesky, Phillip Rust +1

Pixel-based language models process text rendered as images, which allows them to handle any script, making them a promising approach to open vocabulary language modelling. However…

cs.CV2023

Evaluating Bias and Fairness in Gender-Neutral Pretrained Vision-and-Language Models

Laura Cabello, Emanuele Bugliarello, Stephanie Brandl +1

Pretrained machine learning models are known to perpetuate and even amplify existing biases in data, which can result in unfair outcomes that ultimately impact user experience. The…

cs.CL20231 cited

LMCap: Few-shot Multilingual Image Captioning by Retrieval Augmented Language Model Prompting

Rita Ramos, Bruno Martins, Desmond Elliott

Multilingual image captioning has recently been tackled by training with large-scale machine translated data, which is an expensive, noisy, and time-consuming process. Without requ…

cs.CV20234 cited

Retrieval-augmented Image Captioning

Rita Ramos, Desmond Elliott, Bruno Martins

Inspired by retrieval-augmented language generation and pretrained Vision and Language (V&L) encoders, we present a new approach to image captioning that generates sentences given…