activity
20182022
most citedWinoground: Probing Vision and Language Models for Visio-Linguistic Compositionality

14 citations · 16 across the 3 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL20222 cited

Dynatask: A Framework for Creating Dynamic AI Benchmark Tasks

Tristan Thrush, Kushal Tirumala, Anmol Gupta +7

We introduce Dynatask: an open source system for setting up custom NLP tasks that aims to greatly lower the technical knowledge and effort required for hosting and evaluating state…

cs.CL2021

Contrasting Human- and Machine-Generated Word-Level Adversarial Examples for Text Classification

Maximilian Mozes, Max Bartolo, Pontus Stenetorp +2

Research shows that natural language processing models are generally considered to be vulnerable to adversarial attacks; but recent work has drawn attention to the issue of validat…

cs.CL2021

Dynabench: Rethinking Benchmarking in NLP

Douwe Kiela, Max Bartolo, Yixin Nie +16

We introduce Dynabench, an open-source platform for dynamic dataset creation and model benchmarking. Dynabench runs in a web browser and supports human-and-model-in-the-loop datase…

cs.CL2020

Undersensitivity in Neural Reading Comprehension

Johannes Welbl, Pasquale Minervini, Max Bartolo +2

Current reading comprehension models generalise well to in-distribution test sets, yet perform poorly on adversarially selected inputs. Most prior work on adversarial inputs studie…

cs.CL2020

Beat the AI: Investigating Adversarial Human Annotation for Reading Comprehension

Max Bartolo, Alastair Roberts, Johannes Welbl +2

Innovations in annotation methodology have been a catalyst for Reading Comprehension (RC) datasets and models. One recent trend to challenge current RC models is to involve a model…

cs.CL2018

Interpretation of Natural Language Rules in Conversational Machine Reading

Marzieh Saeidi, Max Bartolo, Patrick Lewis +5

Most work in machine reading focuses on question answering problems where the answer is directly expressed in the text to read. However, many real-world question answering problems…