2 papers
cs.LG2026
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective
Haichuan Wang, Tao Lin, Lingkai Kong +3
Existing alignment methods directly use the reward model learned from user preference data to optimize an LLM policy, subject to KL regularization with respect to the base policy.…
cs.GT2025
Information Design with Unknown Prior
Ce Li, Tao Lin
Information designers, such as online platforms, often do not know the beliefs of their receivers. We design learning algorithms so that the information designer can learn the rece…