10 papers
Toward Simultaneously Optimal Regret in U-Calibration
Rafael Frongillo, Haipeng Luo, Nishant A. Mehta +1
U-calibration studies online forecasting algorithms whose predictions can be consumed by any unknown downstream agent, guaranteeing sublinear regret simultaneously for all proper l…
Swap Regret Minimization Through Response-Based Approachability
Ioannis Anagnostides, Gabriele Farina, Maxwell Fishelson +2
We consider the problem of minimizing different notions of swap regret in online optimization. These forms of regret are tightly connected to correlated equilibrium concepts in gam…
Generalized Distributional Alignment Games for Unbiased Answer-Level Fine-Tuning
Mehryar Mohri, Jon Schneider, Yutao Zhong
The Distributional Alignment Game framework provides a powerful variational perspective on Answer-Level Fine-Tuning (ALFT). However, standard algorithms for these games rely on est…
Distributional Alignment Games for Answer-Level Fine-Tuning
Mehryar Mohri, Jon Schneider, Yifan Wu
We focus on the problem of \emph{Answer-Level Fine-Tuning} (ALFT), where the goal is to optimize a language model based on the correctness or properties of its final answers, rathe…
Next-Token Prediction and Regret Minimization
Mehryar Mohri, Clayton Sanford, Jon Schneider +2
We consider the question of how to employ next-token prediction algorithms in adversarial online decision-making environments. Specifically, if we train a next-token prediction mod…
Efficient Opportunistic Approachability
Teodor Vanislavov Marinov, Mehryar Mohri, Princewill Okoroafor +2
We study the problem of opportunistic approachability: a generalization of Blackwell approachability where the learner would like to obtain stronger guarantees (i.e., approach a sm…