3 citations · 6 across the 6 of their papers we have counts for
6 papers
Toward Optimal LLM Alignments Using Two-Player Games
Rui Zheng, Hongyi Guo, Zhihan Liu +10
The standard Reinforcement Learning from Human Feedback (RLHF) framework primarily focuses on optimizing the performance of large language models using pre-collected prompts. Howev…
Improving Reinforcement Learning from Human Feedback Using Contrastive Rewards
Wei Shen, Xiaoying Zhang, Yuanshun Yao +3
Reinforcement learning from human feedback (RLHF) is the mainstream paradigm used to align large language models (LLMs) with human preferences. Yet existing RLHF heavily relies on…
Learning to Watermark LLM-generated Text via Reinforcement Learning
Xiaojun Xu, Yuanshun Yao, Yang Liu
We study how to watermark LLM outputs, i.e. embedding algorithmically detectable signals into LLM-generated text to track misuse. Unlike the current mainstream methods that work wi…
Human-Instruction-Free LLM Self-Alignment with Limited Samples
Hongyi Guo, Yuanshun Yao, Wei Shen +4
Aligning large language models (LLMs) with human values is a vital task for LLM practitioners. Current alignment techniques have several limitations: (1) requiring a large amount o…
Fair Classifiers that Abstain without Harm
Tongxin Yin, Jean-François Ton, Ruocheng Guo +3
In critical applications, it is vital for classifiers to defer decision-making to humans. We propose a post-hoc method that makes existing classifiers selectively abstain from pred…
DPAUC: Differentially Private AUC Computation in Federated Learning
Jiankai Sun, Xin Yang, Yuanshun Yao +3
Federated learning (FL) has gained significant attention recently as a privacy-enhancing tool to jointly train a machine learning model by multiple participants. The prior work on…