4 papers
Selective Ensemble Based on Preference-Directed Multi-Objective Bandits
Lanjihong Ma, Zhen-Yu Zhang, Masashi Sugiyama +1
Selective ensemble for modern machine learning systems requires choosing promising model candidates under limited evaluation budgets, while downstream tasks often specify only part…
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback
Zhen-Yu Zhang, Yuting Tang, Jiandong Zhang +2
Online reinforcement learning from human feedback (RLHF) has emerged as a promising paradigm for aligning large language models (LLMs) by continuously collecting new preference fee…
In-context Demonstration Matters: On Prompt Optimization for Pseudo-Supervision Refinement
Zhen-Yu Zhang, Jiandong Zhang, Huaxiu Yao +2
Large language models (LLMs) have achieved great success across diverse tasks, and fine-tuning is sometimes needed to further enhance generation quality. Most existing methods rely…
Generating Chain-of-Thoughts with a Pairwise-Comparison Approach to Searching for the Most Promising Intermediate Thought
Zhen-Yu Zhang, Siwei Han, Huaxiu Yao +2
To improve the ability of the large language model (LLMs) to tackle complex reasoning problems, chain-of-thoughts (CoT) methods were proposed to guide LLMs to reason step-by-step,…