activity
20152026
most citedCaloMan: Fast generation of calorimeter showers with density estimation on learned manifolds

36 citations · 66 across the 32 of their papers we have counts for

collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2026

Unifying Conformal Language Tasks with In-Context Ensembles

Xiao Shi Huang, Chen-Yuan Lin, Bruce Kuwahara +2

Many NLP tasks, such as summarization and extractive question answering, reduce to retrieving relevant content from documents under two constraints: coverage, retaining enough pert…

cs.CL2026

LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes

Michael Solodko, Steven Gong, Guangwei Yu +3

While modern question answering (QA) systems excel on clean, schema-aligned corpora, real-world knowledge is rarely so neatly packaged. Answering questions over enterprise and scie…

cs.CL2026

RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator

Zhenwei Tang, Zhaoyan Liu, Rasa Hosseinzadeh +3

As interactive LLM-based applications are created and refined, model developers need to evaluate the quality of generated text along many possible axes. For simpler systems, human…

cs.CL2026

Chartographer: Counterfactual Chart Generation for Evaluating Vision-Language Models

Yifan Jiang, Dae Yon Hwang, Jesse C. Cresswell +1

Chart question-answering (QA) benchmarks aim to pose questions that require visual reasoning to correctly answer, but vision-language models (VLMs) can often reach solutions throug…

cs.CL2025

Classifying and Addressing the Diversity of Errors in Retrieval-Augmented Generation Systems

Kin Kwan Leung, Mouloud Belbahri, Yi Sui +4

Retrieval-augmented generation (RAG) is a prevalent approach for building LLM-based question-answering systems that can take advantage of external knowledge databases. Due to the c…

cs.CL2025

Document Summarization with Conformal Importance Guarantees

Bruce Kuwahara, Chen-Yuan Lin, Xiao Shi Huang +5

Automatic summarization systems have advanced rapidly with large language models (LLMs), yet they still lack reliable guarantees on inclusion of critical content in high-stakes dom…