activity
20162024
most citedLLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts

43 citations · 66 across the 13 of their papers we have counts for

collaborators
Showing 2022Show all

6 papers · 1 filter

cs.CL2022★ 1 cited

Ontologically Faithful Generation of Non-Player Character Dialogues

Nathaniel Weir, Ryan Thomas, Randolph D'Amore +3

We introduce a language generation task grounded in a popular video game environment. KNUDGE (KNowledge Constrained User-NPC Dialogue GEneration) requires models to produce trees o…

cs.CV2022★ 4 cited

Super-CLEVR: A Virtual Benchmark to Diagnose Domain Robustness in Visual Reasoning

Zhuowan Li, Xingrui Wang, Elias Stengel-Eskin +4

Visual Question Answering (VQA) models often perform poorly on out-of-distribution data and struggle on domain generalization. Due to the multi-modal nature of this task, multiple…

cs.CL2022

Why Did the Chicken Cross the Road? Rephrasing and Analyzing Ambiguous Questions in VQA

Elias Stengel-Eskin, Jimena Guallar-Blasco, Yi Zhou +1

Natural language is ambiguous. Resolving ambiguous questions is key to successfully answering them. Focusing on questions about images, we create a dataset of ambiguous examples. W…

cs.CL2022

Calibrated Interpretation: Confidence Estimation in Semantic Parsing

Elias Stengel-Eskin, Benjamin Van Durme

Sequence generation models are increasingly being used to translate natural language into programs, i.e. to perform executable semantic parsing. The fact that semantic parsing aims…

cs.CL2022

Iterative Document-level Information Extraction via Imitation Learning

Yunmo Chen, William Gantt, Weiwei Gu +3

We present a novel iterative extraction model, IterX, for extracting complex relations, or templates (i.e., N-tuples representing a mapping from named slots to spans of text) withi…

cs.CL2022

Zero-shot Cross-lingual Transfer is Under-specified Optimization

Shijie Wu, Benjamin Van Durme, Mark Dredze

Pretrained multilingual encoders enable zero-shot cross-lingual transfer, but often produce unreliable models that exhibit high performance variance on the target language. We post…