3 papers
cs.LG2025
Generalized Linear Bandits: Almost Optimal Regret with One-Pass Update
Yu-Jie Zhang, Sheng-An Xu, Peng Zhao +1
We study the generalized linear bandit (GLB) problem, a contextual multi-armed bandit framework that extends the classical linear model by incorporating a non-linear link function,…
cs.LG2025
Non-stationary Online Learning for Curved Losses: Improved Dynamic Regret via Mixability
Yu-Jie Zhang, Peng Zhao, Masashi Sugiyama
Non-stationary online learning has drawn much attention in recent years. Despite considerable progress, dynamic regret minimization has primarily focused on convex functions, leavi…
cs.LG2025
On Symmetric Losses for Robust Policy Optimization with Noisy Preferences
Soichiro Nishimori, Yu-Jie Zhang, Thanawat Lodkaew +1
Optimizing policies based on human preferences is key to aligning language models with human intent. This work focuses on reward modeling, a core component in reinforcement learnin…