6 papers
JADAI: Jointly Amortizing Adaptive Design and Bayesian Inference
Niels Bracher, Lars Kühmichel, Desi R. Ivanova +3
We consider problems of parameter estimation where design variables can be actively optimized to maximize information gain. To this end, we introduce JADAI, a framework that jointl…
Improving Semantic Uncertainty Quantification in Language Model Question-Answering via Token-Level Temperature Scaling
Tom A. Lamb, Desi R. Ivanova, Philip H. S. Torr +1
Calibration is central to reliable semantic uncertainty quantification, yet prior work has largely focused on discrimination, neglecting calibration. As calibration and discriminat…
Step-DAD: Semi-Amortized Policy-Based Bayesian Experimental Design
Marcel Hedman, Desi R. Ivanova, Cong Guan +1
We develop a semi-amortized, policy-based, approach to Bayesian experimental design (BED) called Stepwise Deep Adaptive Design (Step-DAD). Like existing, fully amortized, policy-ba…
Is merging worth it? Securely evaluating the information gain for causal dataset acquisition
Jake Fawkes, Lucile Ter-Minassian, Desi Ivanova +2
Merging datasets across institutions is a lengthy and costly procedure, especially when it involves private information. Data hosts may therefore want to prospectively gauge which…
Extending Epistemic Uncertainty Beyond Parameters Would Assist in Designing Reliable LLMs
T. Duy Nguyen-Hien, Desi R. Ivanova, Yee Whye Teh +1
Although large language models (LLMs) are highly interactive and extendable, current approaches to ensure reliability in deployments remain mostly limited to rejecting outputs with…
Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints
Sam Bowyer, Laurence Aitchison, Desi R. Ivanova
Rigorous statistical evaluations of large language models (LLMs), including valid error bars and significance testing, are essential for meaningful and reliable performance assessm…