36 citations · 66 across the 32 of their papers we have counts for
7 papers · 1 filter
Unifying Conformal Language Tasks with In-Context Ensembles
Xiao Shi Huang, Chen-Yuan Lin, Bruce Kuwahara +2
Many NLP tasks, such as summarization and extractive question answering, reduce to retrieving relevant content from documents under two constraints: coverage, retaining enough pert…
LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes
Michael Solodko, Steven Gong, Guangwei Yu +3
While modern question answering (QA) systems excel on clean, schema-aligned corpora, real-world knowledge is rarely so neatly packaged. Answering questions over enterprise and scie…
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator
Zhenwei Tang, Zhaoyan Liu, Rasa Hosseinzadeh +3
As interactive LLM-based applications are created and refined, model developers need to evaluate the quality of generated text along many possible axes. For simpler systems, human…
Chartographer: Counterfactual Chart Generation for Evaluating Vision-Language Models
Yifan Jiang, Dae Yon Hwang, Jesse C. Cresswell +1
Chart question-answering (QA) benchmarks aim to pose questions that require visual reasoning to correctly answer, but vision-language models (VLMs) can often reach solutions throug…
Classifying and Addressing the Diversity of Errors in Retrieval-Augmented Generation Systems
Kin Kwan Leung, Mouloud Belbahri, Yi Sui +4
Retrieval-augmented generation (RAG) is a prevalent approach for building LLM-based question-answering systems that can take advantage of external knowledge databases. Due to the c…
Document Summarization with Conformal Importance Guarantees
Bruce Kuwahara, Chen-Yuan Lin, Xiao Shi Huang +5
Automatic summarization systems have advanced rapidly with large language models (LLMs), yet they still lack reliable guarantees on inclusion of critical content in high-stakes dom…