3 papers
cs.LG2026
Selective Ensemble Based on Preference-Directed Multi-Objective Bandits
Lanjihong Ma, Zhen-Yu Zhang, Masashi Sugiyama +1
Selective ensemble for modern machine learning systems requires choosing promising model candidates under limited evaluation budgets, while downstream tasks often specify only part…
cs.LG2026
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback
Zhen-Yu Zhang, Yuting Tang, Jiandong Zhang +2
Online reinforcement learning from human feedback (RLHF) has emerged as a promising paradigm for aligning large language models (LLMs) by continuously collecting new preference fee…
cs.LG2019
An Unbiased Risk Estimator for Learning with Augmented Classes
Yu-Jie Zhang, Peng Zhao, Zhi-Hua Zhou
This paper studies the problem of learning with augmented classes (LAC), where augmented classes unobserved in the training data might emerge in the testing phase. Previous studies…