Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
FormInv: A Measurement Protocol for Semantic Invariance in Mathematical Reasoning Benchmarks
Nishal Thomas, Noel Thomas
A paraphrase-quality audit of MathCheck (ICLR 2025) detected 4 semantically incorrect paraphrases in 129 groups (3.1%); removing them drops GPT-4o from rank 2 to rank 4 and elevate…
cs.LG2026
ChaosBench-Logic v2: Evaluating LLM Logical Reasoning over Dynamical Systems at Scale
Noel Thomas
Standard accuracy on binary reasoning benchmarks hides critical failure modes: prior collapse, inconsistency under paraphrase, and inability to reason about parameter-dependent dyn…
cs.LG2026
Regime-Conditioned Evaluation in Multi-Context Bayesian Optimization
Noel Thomas
Published transfer-BO comparisons often estimate an average treatment effect of acquisition choice over hidden regime variables, while practitioners need the conditional effect for…