activity
20162022
most citedAn Exploration of Hierarchical Attention Transformers for Efficient Long Document Classification

17 citations · 20 across the 10 of their papers we have counts for

collaborators

21 papers

cs.CL2022

Multilingual Multimodal Learning with Machine Translated Text

Chen Qiu, Dan Oneata, Emanuele Bugliarello +2

Most vision-and-language pretraining research focuses on English tasks. However, the creation of multilingual multimodal evaluation datasets (e.g. Multi30K, xGQA, XVNLI, and MaRVL)…

cs.CL202217 cited

An Exploration of Hierarchical Attention Transformers for Efficient Long Document Classification

Ilias Chalkidis, Xiang Dai, Manos Fergadiotis +2

Non-hierarchical sparse attention Transformer-based models, such as Longformer and Big Bird, are popular approaches to working with long documents. There are clear benefits to thes…

cs.CL20211 cited

Visually Grounded Reasoning across Languages and Cultures

Fangyu Liu, Emanuele Bugliarello, Edoardo Maria Ponti +3

The design of widespread vision-and-language datasets and pre-trained encoders directly adopts, or draws inspiration from, the concepts and images of ImageNet. While one can hardly…

cs.CL2021

MDAPT: Multilingual Domain Adaptive Pretraining in a Single Model

Rasmus Kær Jørgensen, Mareike Hartmann, Xiang Dai +1

Domain adaptive pretraining, i.e. the continued unsupervised pretraining of a language model on domain-specific text, improves the modelling of text for downstream tasks within the…

cs.CL2021

Vision-and-Language or Vision-for-Language? On Cross-Modal Influence in Multimodal Transformers

Stella Frank, Emanuele Bugliarello, Desmond Elliott

Pretrained vision-and-language BERTs aim to learn representations that combine information from both modalities. We propose a diagnostic method based on cross-modal input ablation…

cs.CL2021

The Role of Syntactic Planning in Compositional Image Captioning

Emanuele Bugliarello, Desmond Elliott

Image captioning has focused on generalizing to images drawn from the same distribution as the training set, and not to the more challenging problem of generalizing to different di…