37 citations · 117 across the 38 of their papers we have counts for
34 papers · 1 filter
A positive resolution of the gap-entropy conjecture
P. M. Aronow, Nathan Kallus, Patrick Lopatto
We prove the gap-entropy conjecture for fixed-confidence best-arm identification with independent unit-variance Gaussian arms, means in , and a unique optimal arm. For each…
Causal Inference on Networks under Misspecified Exposure Mappings: A Partial Identification Framework
Maresa Schröder, Miruna Oprescu, Stefan Feuerriegel +1
Estimating treatment effects in networks is challenging, as each potential outcome depends on the treatments of all other nodes in the network. To overcome this difficulty, existin…
Exploration in the Limit
Brian M. Cho, Nathan Kallus
In fixed-confidence best arm identification (BAI), the objective is to quickly identify the optimal option while controlling the probability of error below a desired threshold. Des…
Entropy After </Think> for reasoning model early exiting
Xi Wang, James McInerney, Lequn Wang +1
Reasoning LLMs show improved performance with longer chains of thought. However, recent work has highlighted their tendency to overthink, continuing to revise answers even after re…
Optimization of Epsilon-Greedy Exploration
Ethan Che, Hakan Ceylan, James McInerney +1
Modern recommendation systems rely on exploration to learn user preferences for new items, typically implementing uniform exploration policies (e.g., epsilon-greedy) due to their s…
Value-Guided Search for Efficient Chain-of-Thought Reasoning
Kaiwen Wang, Jin Peng Zhou, Jonathan Chang +4
In this paper, we propose a simple and efficient method for value model training on long-context reasoning traces. Compared to existing process reward models (PRMs), our method doe…