activity
20192025
most citedSuper-CLEVR: A Virtual Benchmark to Diagnose Domain Robustness in Visual Reasoning

4 citations · 7 across the 10 of their papers we have counts for

collaborators

16 papers

cs.CL2025

Calibrating Verbalized Confidence with Self-Generated Distractors

Victor Wang, Elias Stengel-Eskin

Calibrated confidence estimates are necessary for large language model (LLM) outputs to be trusted by human users. While LLMs can express their confidence in human-interpretable wa…

cs.CL2023★ 1 cited

Rephrase, Augment, Reason: Visual Grounding of Questions for Vision-Language Models

Archiki Prasad, Elias Stengel-Eskin, Mohit Bansal

An increasing number of vision-language tasks can be handled with little to no training, i.e., in a zero and few-shot manner, by marrying large language models (LLMs) to vision enc…

cs.CL2023★ 2 cited

Zero and Few-shot Semantic Parsing with Ambiguous Inputs

Elias Stengel-Eskin, Kyle Rawlins, Benjamin Van Durme

Despite the frequent challenges posed by ambiguity when representing meaning via natural language, it is often ignored or deliberately removed in tasks mapping language to formally…

cs.CL2023

Did You Mean...? Confidence-based Trade-offs in Semantic Parsing

Elias Stengel-Eskin, Benjamin Van Durme

We illustrate how a calibrated model can help balance common trade-offs in task-oriented parsing. In a simulated annotator-in-the-loop experiment, we show that well-calibrated conf…

cs.CV2022★ 4 cited

Super-CLEVR: A Virtual Benchmark to Diagnose Domain Robustness in Visual Reasoning

Zhuowan Li, Xingrui Wang, Elias Stengel-Eskin +4

Visual Question Answering (VQA) models often perform poorly on out-of-distribution data and struggle on domain generalization. Due to the multi-modal nature of this task, multiple…

cs.CL2022

Why Did the Chicken Cross the Road? Rephrasing and Analyzing Ambiguous Questions in VQA

Elias Stengel-Eskin, Jimena Guallar-Blasco, Yi Zhou +1

Natural language is ambiguous. Resolving ambiguous questions is key to successfully answering them. Focusing on questions about images, we create a dataset of ambiguous examples. W…