2 papers
cs.CL2026
Calibrating Verbalized Confidence with Self-Generated Distractors
Victor Wang, Elias Stengel-Eskin
Calibrated confidence estimates are necessary for large language model (LLM) outputs to be trusted by human users. While LLMs can express their confidence in human-interpretable wa…
cs.CL2025
Improving LLM-as-a-Judge Inference with the Judgment Distribution
Victor Wang, Michael J. Q. Zhang, Eunsol Choi
Using language models to scalably approximate human preferences on text quality (LLM-as-a-judge) has become a standard practice applicable to many tasks. A judgment is often extrac…