160 citations · 293 across the 51 of their papers we have counts for
4 papers · 2 filters
Unifying Conformal Language Tasks with In-Context Ensembles
Xiao Shi Huang, Chen-Yuan Lin, Bruce Kuwahara +2
Many NLP tasks, such as summarization and extractive question answering, reduce to retrieving relevant content from documents under two constraints: coverage, retaining enough pert…
LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes
Michael Solodko, Steven Gong, Guangwei Yu +3
While modern question answering (QA) systems excel on clean, schema-aligned corpora, real-world knowledge is rarely so neatly packaged. Answering questions over enterprise and scie…
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator
Zhenwei Tang, Zhaoyan Liu, Rasa Hosseinzadeh +3
As interactive LLM-based applications are created and refined, model developers need to evaluate the quality of generated text along many possible axes. For simpler systems, human…
Chartographer: Counterfactual Chart Generation for Evaluating Vision-Language Models
Yifan Jiang, Dae Yon Hwang, Jesse C. Cresswell +1
Chart question-answering (QA) benchmarks aim to pose questions that require visual reasoning to correctly answer, but vision-language models (VLMs) can often reach solutions throug…