5 papers
Unsupervised Confidence Calibration for Reasoning LLMs from a Single Generation
Thomas Zollo, Jimmy Wang, Richard Zemel
Reasoning language models can solve increasingly complex tasks, but struggle to produce the calibrated confidence estimates necessary for reliable deployment. Existing calibration…
Tell Me What To Learn: Generalizing Neural Memory to be Controllable in Natural Language
Max S. Bennett, Thomas P. Zollo, Richard Zemel
Modern machine learning models are deployed in diverse, non-stationary environments where they must continually adapt to new tasks and evolving knowledge. Continual fine-tuning and…
Test-Time Warmup for Multimodal Large Language Models
Nikita Rajaneesh, Thomas Zollo, Richard Zemel
Multimodal Large Language Models (MLLMs) hold great promise for advanced reasoning at the intersection of text and images, yet they have not fully realized this potential. MLLMs ty…
QuEst: Enhancing Estimates of Quantile-Based Distributional Measures Using Model Predictions
Zhun Deng, Thomas P Zollo, Benjamin Eyre +3
As machine learning models grow increasingly competent, their predictions can supplement scarce or expensive data in various important domains. In support of this paradigm, algorit…
Adaptive Elicitation of Latent Information Using Natural Language
Jimmy Wang, Thomas Zollo, Richard Zemel +1
Eliciting information to reduce uncertainty about a latent entity is a critical task in many application domains, e.g., assessing individual student learning outcomes, diagnosing u…