4 papers
NuclearQAv2: A Structured Benchmark for Evaluating Domain-Science Competence in Large Language Models
Henry Shaowu Yuchi, Michal Kucer, Benjamin H. Sims +2
Large language models (LLMs) have demonstrated strong performance across a wide range of tasks, but ensuring their reliability in highly technical domains remains a significant cha…
Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences
Sweta Karlekar, Carolina Zheng, Magnus Saebo +5
Many applications seek to optimize LLM outputs at test time by iteratively proposing, scoring, and refining candidates over a discrete output space. Existing methods use a calibrat…
The Trust Calibration Maturity Model for Characterizing and Communicating Trustworthiness of AI Systems
Scott T Steinmetz, Asmeret Naugle, Paul Schutte +6
Recent proliferation of powerful AI systems has created a strong need for capabilities that help users to calibrate trust in those systems. As AI systems grow in scale, information…
Posterior Mean Matching: Generative Modeling through Online Bayesian Inference
Sebastian Salazar, Michal Kucer, Yixin Wang +2
This paper introduces posterior mean matching (PMM), a new method for generative modeling that is grounded in Bayesian inference. PMM uses conjugate pairs of distributions to model…