3 papers
cs.LG2026
Unsupervised Confidence Calibration for Reasoning LLMs from a Single Generation
Thomas Zollo, Jimmy Wang, Richard Zemel
Reasoning language models can solve increasingly complex tasks, but struggle to produce the calibrated confidence estimates necessary for reliable deployment. Existing calibration…
cs.CL2025
Adaptive Elicitation of Latent Information Using Natural Language
Jimmy Wang, Thomas Zollo, Richard Zemel +1
Eliciting information to reduce uncertainty about a latent entity is a critical task in many application domains, e.g., assessing individual student learning outcomes, diagnosing u…
cs.LG2024
Optimization-Driven Adaptive Experimentation
Ethan Che, Daniel R. Jiang, Hongseok Namkoong +1
Real-world experiments involve batched & delayed feedback, non-stationarity, multiple objectives & constraints, and (often some) personalization. Tailoring adaptive methods to addr…