3 papers
cs.LG2026
Online Reasoning Calibration: Test-Time Training Enables Generalizable Conformal LLM Reasoning
Cai Zhou, Zekai Wang, Menghua Wu +6
While test-time scaling has enabled large language models to solve highly difficult tasks, state-of-the-art results come at exorbitant compute costs. These inefficiencies can be at…
cs.GT2025
Instance-Adaptive Hypothesis Tests with Heterogeneous Agents
Flora C. Shi, Martin J. Wainwright, Stephen Bates
We study hypothesis testing over a heterogeneous population of strategic agents with private information. Any single test applied uniformly across the population yields statistical…
stat.ME2024
Sharp Results for Hypothesis Testing with Risk-Sensitive Agents
Flora C. Shi, Stephen Bates, Martin J. Wainwright
Statistical protocols are often used for decision-making involving multiple parties, each with their own incentives, private information, and ability to influence the distributiona…