14 papers
Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration
Ruochen Jin, Zhanliang Wang, Zongyu Dai +2
Preference alignment often makes large language models (LLMs) overconfident and poorly calibrated. Traditional post-hoc temperature scaling is inherently domain-dependent: a temper…
Stochastic Regret Guarantees for Online Zeroth- and First-Order Bilevel Optimization
Parvin Nazari, Bojian Hou, Davoud Ataee Tarzanagh +2
Online bilevel optimization (OBO) is a powerful framework for machine learning problems where both outer and inner objectives evolve over time, requiring dynamic updates. Current O…
A Semantic-Sampling Framework for Evaluating Calibration in Open-Ended Question Answering
Zhanliang Wang, Jiancong Xiao, Ruochen Jin +3
Calibration measures whether a model's predicted confidence aligns with its empirical accuracy, and is central to the reliable deployment of large language models (LLMs) in high-st…
Tabular LLMs for Interpretable Few-Shot Alzheimer's Disease Prediction with Multimodal Biomedical Data
Sophie Kearney, Shu Yang, Zixuan Wen +8
Accurate diagnosis of Alzheimer's disease (AD) requires handling tabular biomarker data, yet such data are often small and incomplete, where deep learning models frequently fail to…
Restoring Calibration for Aligned Large Language Models: A Calibration-Aware Fine-Tuning Approach
Jiancong Xiao, Bojian Hou, Zhanliang Wang +4
One of the key technologies for the success of Large Language Models (LLMs) is preference alignment. However, a notable side effect of preference alignment is poor calibration: whi…
Enabling Few-Shot Alzheimer's Disease Diagnosis on Biomarker Data with Tabular LLMs
Sophie Kearney, Shu Yang, Zixuan Wen +6
Early and accurate diagnosis of Alzheimer's disease (AD), a complex neurodegenerative disorder, requires analysis of heterogeneous biomarkers (e.g., neuroimaging, genetic risk fact…