16 citations · 39 across the 7 of their papers we have counts for
10 papers
Evaluate & Evaluation on the Hub: Better Best Practices for Data and Model Measurements
Leandro von Werra, Lewis Tunstall, Abhishek Thakur +16
Evaluation is a key part of machine learning (ML), yet there is a lack of support and tooling to enable its informed and systematic practice. We introduce Evaluate and Evaluation o…
Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality
Tristan Thrush, Ryan Jiang, Max Bartolo +4
We present a novel task and dataset for evaluating the ability of vision and language models to conduct visio-linguistic compositional reasoning, which we call Winoground. Given tw…
Dynatask: A Framework for Creating Dynamic AI Benchmark Tasks
Tristan Thrush, Kushal Tirumala, Anmol Gupta +7
We introduce Dynatask: an open source system for setting up custom NLP tasks that aims to greatly lower the technical knowledge and effort required for hosting and evaluating state…
Dynaboard: An Evaluation-As-A-Service Platform for Holistic Next-Generation Benchmarking
Zhiyi Ma, Kawin Ethayarajh, Tristan Thrush +6
We introduce Dynaboard, an evaluation-as-a-service framework for hosting benchmarks and conducting holistic model comparison, integrated with the Dynabench platform. Our platform e…
Dynabench: Rethinking Benchmarking in NLP
Douwe Kiela, Max Bartolo, Yixin Nie +16
We introduce Dynabench, an open-source platform for dynamic dataset creation and model benchmarking. Dynabench runs in a web browser and supports human-and-model-in-the-loop datase…
Rover Relocalization for Mars Sample Return by Virtual Template Synthesis and Matching
Tu-Hoa Pham, William Seto, Shreyansh Daftry +11
We consider the problem of rover relocalization in the context of the notional Mars Sample Return campaign. In this campaign, a rover (R1) needs to be capable of autonomously navig…