3 papers
cs.CL2026
A Robust Evaluation of Probe Robustness: Lessons for Reliable OOD Uncertainty Quantification
Joe Stacey, Hadas Orgad, Kentaro Inui +2
Recent work has shown that the hidden states of large language models contain signals useful for uncertainty estimation, motivating a growing interest in efficient probe-based appr…
cs.CL2025
AgentCoMa: A Compositional Benchmark Mixing Commonsense and Mathematical Reasoning in Real-World Scenarios
Lisa Alazraki, Lihu Chen, Ana Brassard +3
Large Language Models (LLMs) have achieved high accuracy on complex commonsense and mathematical problems that involve the composition of multiple reasoning steps. However, current…
cs.CL2025
Improving the OOD Performance of Closed-Source LLMs on NLI Through Strategic Data Selection
Joe Stacey, Lisa Alazraki, Aran Ubhi +3
We investigate the robustness of fine-tuned Large Language Models (LLMs) for the task of Natural Language Inference (NLI), finding that the in-distribution gains from fine-tuning c…