3 papers
cs.LG2026
Inverse Contextual Bandits without Rewards: Learning from a Non-Stationary Learner via Suffix Imitation
Yuqi Kong, Xiao Zhang, Weiran Shen
We study the Inverse Contextual Bandit (ICB) problem, in which a learner seeks to optimize a policy while an observer, who cannot access the learner's rewards and only observes act…
cs.IR2026
Enhancing Long-Term Welfare in Recommender Systems: An Information Revelation Approach
Xu Zhao, Xiaopeng Ye, Chen Xu +2
Improving the long-term user welfare (e.g., sustained user engagement) has become a central objective of recommender systems (RS). In real-world platforms, the creation behaviors o…
cs.LG2025
IBCB: Efficient Inverse Batched Contextual Bandit for Behavioral Evolution History
Yi Xu, Weiran Shen, Xiao Zhang +1
Traditional imitation learning focuses on modeling the behavioral mechanisms of experts, which requires a large amount of interaction history generated by some fixed expert. Howeve…