11 citations · 22 across the 10 of their papers we have counts for
Showing 2023Show all
2 papers · 1 filter
cs.CL2023★ 11 cited
FinanceBench: A New Benchmark for Financial Question Answering
Pranab Islam, Anand Kannappan, Douwe Kiela +3
FinanceBench is a first-of-its-kind test suite for evaluating the performance of LLMs on open book financial question answering (QA). It comprises 10,231 questions about publicly t…
cs.CL2023★ 5 cited
SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models
Bertie Vidgen, Nino Scherrer, Hannah Rose Kirk +4
The past year has seen rapid acceleration in the development of large language models (LLMs). However, without proper steering and safeguards, LLMs will readily follow malicious in…