9 papers · 1 filter
Knowing When to Quit: A Principled Framework for Dynamic Abstention in LLM Reasoning
Hen Davidov, Nachshon Cohen, Oren Kalinsky +4
LLMs utilizing chain-of-thought reasoning often waste substantial compute by producing long, incorrect responses. Abstention can mitigate this by withholding outputs unlikely to be…
Sharp Risk Bounds for Early-Stopping in Gaussian Linear Regression
Tobias Wegel, Gil Kur, Patrick Rebeschini
We study early-stopped mirror descent (ESMD) for high-dimensional Gaussian linear regression over arbitrary convex bodies and design matrices, where the task is to minimize the in-…
On-Average Stability of Multipass Preconditioned SGD and Effective Dimension
Simon Vary, Tyler Farghly, Ilja Kuzborskij +1
We study trade-offs between the population risk curvature, geometry of the noise, and preconditioning on the generalisation ability of the multipass Preconditioned Stochastic Gradi…
Meta-Learning Objectives for Preference Optimization
Carlo Alfano, Silvia Sapora, Jakob Nicolaus Foerster +2
Evaluating preference optimization (PO) algorithms on LLM alignment is a challenging task that presents prohibitive costs, noise, and several variables like model size and hyper-pa…
On the necessity of adaptive regularisation:Optimal anytime online learning on -balls
Emmeran Johnson, David MartÃnez-Rubio, Ciara Pike-Burke +1
We study online convex optimization on -balls in for . While always sub-linear, the optimal regret exhibits a shift between the high-dimensional setti…
Stochastic Shortest Path with Sparse Adversarial Costs
Emmeran Johnson, Alberto Rumi, Ciara Pike-Burke +1
We study the adversarial Stochastic Shortest Path (SSP) problem with sparse costs under full-information feedback. In the known transition setting, existing bounds based on Online…