16 citations · 48 across the 5 of their papers we have counts for
6 papers
FinanceBench: A New Benchmark for Financial Question Answering
Pranab Islam, Anand Kannappan, Douwe Kiela +3
FinanceBench is a first-of-its-kind test suite for evaluating the performance of LLMs on open book financial question answering (QA). It comprises 10,231 questions about publicly t…
Evaluate & Evaluation on the Hub: Better Best Practices for Data and Model Measurements
Leandro von Werra, Lewis Tunstall, Abhishek Thakur +16
Evaluation is a key part of machine learning (ML), yet there is a lack of support and tooling to enable its informed and systematic practice. We introduce Evaluate and Evaluation o…
Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality
Tristan Thrush, Ryan Jiang, Max Bartolo +4
We present a novel task and dataset for evaluating the ability of vision and language models to conduct visio-linguistic compositional reasoning, which we call Winoground. Given tw…
Dynatask: A Framework for Creating Dynamic AI Benchmark Tasks
Tristan Thrush, Kushal Tirumala, Anmol Gupta +7
We introduce Dynatask: an open source system for setting up custom NLP tasks that aims to greatly lower the technical knowledge and effort required for hosting and evaluating state…
What's Hidden in a One-layer Randomly Weighted Transformer?
Sheng Shen, Zhewei Yao, Douwe Kiela +2
We demonstrate that, hidden within one-layer randomly weighted neural networks, there exist subnetworks that can achieve impressive performance, without ever modifying the weight i…
Dynaboard: An Evaluation-As-A-Service Platform for Holistic Next-Generation Benchmarking
Zhiyi Ma, Kawin Ethayarajh, Tristan Thrush +6
We introduce Dynaboard, an evaluation-as-a-service framework for hosting benchmarks and conducting holistic model comparison, integrated with the Dynabench platform. Our platform e…