4 papers
Trained on Tokens, Calibrated on Concepts: The Emergence of Semantic Calibration in LLMs
Preetum Nakkiran, Arwen Bradley, Adam GoliÅski +3
Large Language Models (LLMs) often lack meaningful confidence estimates for their outputs. While base LLMs are known to exhibit next-token calibration, it remains unclear whether t…
Revisiting Uncertainty Quantification Evaluation in Language Models: Spurious Interactions with Response Length Bias Results
Andrea Santilli, Adam Golinski, Michael Kirchhof +5
Uncertainty Quantification (UQ) in Language Models (LMs) is key to improving their safety and reliability. Evaluations often use metrics like AUROC to assess how well UQ methods (e…
Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs
Yinong Oliver Wang, Nivedha Sivakumar, Falaah Arif Khan +6
The recent rapid adoption of large language models (LLMs) highlights the critical need for benchmarking their fairness. Conventional fairness metrics, which focus on discrete accur…
Considerations for Distribution Shift Robustness of Diagnostic Models in Healthcare
Arno Blaas, Adam GoliÅski, Andrew Miller +3
We consider robustness to distribution shifts in the context of diagnostic models in healthcare, where the prediction target , e.g., the presence of a disease, is causally upstr…