5 papers
Generalized Linear Bandits: Almost Optimal Regret with One-Pass Update
Yu-Jie Zhang, Sheng-An Xu, Peng Zhao +1
We study the generalized linear bandit (GLB) problem, a contextual multi-armed bandit framework that extends the classical linear model by incorporating a non-linear link function,…
Recursive Reward Aggregation
Yuting Tang, Yivan Zhang, Johannes Ackermann +3
In reinforcement learning (RL), aligning agent behavior with specific objectives typically requires careful design of the reward function, which can be challenging when the desired…
Non-stationary Online Learning for Curved Losses: Improved Dynamic Regret via Mixability
Yu-Jie Zhang, Peng Zhao, Masashi Sugiyama
Non-stationary online learning has drawn much attention in recent years. Despite considerable progress, dynamic regret minimization has primarily focused on convex functions, leavi…
On Symmetric Losses for Robust Policy Optimization with Noisy Preferences
Soichiro Nishimori, Yu-Jie Zhang, Thanawat Lodkaew +1
Optimizing policies based on human preferences is key to aligning language models with human intent. This work focuses on reward modeling, a core component in reinforcement learnin…
Enriching Disentanglement: From Logical Definitions to Quantitative Metrics
Yivan Zhang, Masashi Sugiyama
Disentangling the explanatory factors in complex data is a promising approach for generalizable and data-efficient representation learning. While a variety of quantitative metrics…