17 citations · 20 across the 10 of their papers we have counts for
21 papers
Multilingual Multimodal Learning with Machine Translated Text
Chen Qiu, Dan Oneata, Emanuele Bugliarello +2
Most vision-and-language pretraining research focuses on English tasks. However, the creation of multilingual multimodal evaluation datasets (e.g. Multi30K, xGQA, XVNLI, and MaRVL)…
An Exploration of Hierarchical Attention Transformers for Efficient Long Document Classification
Ilias Chalkidis, Xiang Dai, Manos Fergadiotis +2
Non-hierarchical sparse attention Transformer-based models, such as Longformer and Big Bird, are popular approaches to working with long documents. There are clear benefits to thes…
Visually Grounded Reasoning across Languages and Cultures
Fangyu Liu, Emanuele Bugliarello, Edoardo Maria Ponti +3
The design of widespread vision-and-language datasets and pre-trained encoders directly adopts, or draws inspiration from, the concepts and images of ImageNet. While one can hardly…
MDAPT: Multilingual Domain Adaptive Pretraining in a Single Model
Rasmus Kær Jørgensen, Mareike Hartmann, Xiang Dai +1
Domain adaptive pretraining, i.e. the continued unsupervised pretraining of a language model on domain-specific text, improves the modelling of text for downstream tasks within the…
Vision-and-Language or Vision-for-Language? On Cross-Modal Influence in Multimodal Transformers
Stella Frank, Emanuele Bugliarello, Desmond Elliott
Pretrained vision-and-language BERTs aim to learn representations that combine information from both modalities. We propose a diagnostic method based on cross-modal input ablation…
The Role of Syntactic Planning in Compositional Image Captioning
Emanuele Bugliarello, Desmond Elliott
Image captioning has focused on generalizing to images drawn from the same distribution as the training set, and not to the more challenging problem of generalizing to different di…