4 papers
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377
To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…
FormInv: A Measurement Protocol for Semantic Invariance in Mathematical Reasoning Benchmarks
Nishal Thomas, Noel Thomas
A paraphrase-quality audit of MathCheck (ICLR 2025) detected 4 semantically incorrect paraphrases in 129 groups (3.1%); removing them drops GPT-4o from rank 2 to rank 4 and elevate…
ChaosBench-Logic v2: Evaluating LLM Logical Reasoning over Dynamical Systems at Scale
Noel Thomas
Standard accuracy on binary reasoning benchmarks hides critical failure modes: prior collapse, inconsistency under paraphrase, and inability to reason about parameter-dependent dyn…
Regime-Conditioned Evaluation in Multi-Context Bayesian Optimization
Noel Thomas
Published transfer-BO comparisons often estimate an average treatment effect of acquisition choice over hidden regime variables, while practitioners need the conditional effect for…