activity
20162023
most citedEvaluating GPT-4 and ChatGPT on Japanese Medical Licensing Examinations

51 citations · 160 across the 22 of their papers we have counts for

collaborators
Showing 2019 · cs.CLShow all

5 papers · 2 filters

cs.CL2019

The Universal Decompositional Semantics Dataset and Decomp Toolkit

Aaron Steven White, Elias Stengel-Eskin, Siddharth Vashishtha +9

We present the Universal Decompositional Semantics (UDS) dataset (v1.0), which is bundled with the Decomp toolkit (v0.1). UDS1.0 unifies five high-quality, decompositional semantic…

cs.CL2019

WIQA: A dataset for "What if..." reasoning over procedural text

Niket Tandon, Bhavana Dalvi Mishra, Keisuke Sakaguchi +2

We introduce WIQA, the first large-scale dataset of "What if..." questions over procedural text. WIQA contains three parts: a collection of paragraphs each describing a process, e.…

cs.CL2019

Uncertain Natural Language Inference

Tongfei Chen, Zhengping Jiang, Adam Poliak +2

We introduce Uncertain Natural Language Inference (UNLI), a refinement of Natural Language Inference (NLI) that shifts away from categorical labels, targeting instead the direct pr…

cs.CL2019

Abductive Commonsense Reasoning

Chandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya +6

Abductive reasoning is inference to the most plausible explanation. For example, if Jenny finds her house in a mess when she returns from work, and remembers that she left a window…

cs.CL2019

WinoGrande: An Adversarial Winograd Schema Challenge at Scale

Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula +1

The Winograd Schema Challenge (WSC) (Levesque, Davis, and Morgenstern 2011), a benchmark for commonsense reasoning, is a set of 273 expert-crafted pronoun resolution problems origi…