2 papers
cs.CL2026
NuclearQAv2: A Structured Benchmark for Evaluating Domain-Science Competence in Large Language Models
Henry Shaowu Yuchi, Michal Kucer, Benjamin H. Sims +2
Large language models (LLMs) have demonstrated strong performance across a wide range of tasks, but ensuring their reliability in highly technical domains remains a significant cha…
stat.ME2026
Discrepancy Modeling with Intermediate Variables: A New Framework for Robust Gaussian Process Calibration
Henry Shaowu Yuchi, Michael Grosskopf, Aman Sharma +4
Gaussian processes are widely used for surrogate modeling in computer experiments, which often produce numerous intermediate variables that are not explicitly used in standard cali…