4 papers
Efficient Multi-objective Prompt Optimization via Pure-exploration Bandits
Donghao Li, Chengshuai Shi, Weijuan Ou +2
Prompt engineering has become central to eliciting the capabilities of large language models (LLMs). At its core lies prompt selection -- efficiently identifying the most effective…
Breaking the Computational Barrier: Provably Efficient Actor-Critic for Low-Rank MDPs
Ruiquan Huang, Donghao Li, Yingbin Liang +1
Reinforcement learning (RL) is a fundamental framework for sequential decision-making, in which an agent learns an optimal policy through interactions with an unknown environment.…
Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis
Ruiquan Huang, Donghao Li, Chengshuai Shi +2
This paper investigates a hybrid learning framework for reinforcement learning (RL) in which the agent can leverage both an offline dataset and online interactions to learn the opt…
A Shared Low-Rank Adaptation Approach to Personalized RLHF
Renpu Liu, Peng Wang, Donghao Li +2
Reinforcement Learning from Human Feedback (RLHF) has emerged as a pivotal technique for aligning artificial intelligence systems with human values, achieving remarkable success in…