activity
20222024
most citedLearning to Watermark LLM-generated Text via Reinforcement Learning

3 citations · 6 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CL2024

Toward Optimal LLM Alignments Using Two-Player Games

Rui Zheng, Hongyi Guo, Zhihan Liu +10

The standard Reinforcement Learning from Human Feedback (RLHF) framework primarily focuses on optimizing the performance of large language models using pre-collected prompts. Howev…

cs.CL2024

Improving Reinforcement Learning from Human Feedback Using Contrastive Rewards

Wei Shen, Xiaoying Zhang, Yuanshun Yao +3

Reinforcement learning from human feedback (RLHF) is the mainstream paradigm used to align large language models (LLMs) with human preferences. Yet existing RLHF heavily relies on…

cs.LG20243 cited

Learning to Watermark LLM-generated Text via Reinforcement Learning

Xiaojun Xu, Yuanshun Yao, Yang Liu

We study how to watermark LLM outputs, i.e. embedding algorithmically detectable signals into LLM-generated text to track misuse. Unlike the current mainstream methods that work wi…

cs.CL20242 cited

Human-Instruction-Free LLM Self-Alignment with Limited Samples

Hongyi Guo, Yuanshun Yao, Wei Shen +4

Aligning large language models (LLMs) with human values is a vital task for LLM practitioners. Current alignment techniques have several limitations: (1) requiring a large amount o…

cs.LG2023

Fair Classifiers that Abstain without Harm

Tongxin Yin, Jean-François Ton, Ruocheng Guo +3

In critical applications, it is vital for classifiers to defer decision-making to humans. We propose a post-hoc method that makes existing classifiers selectively abstain from pred…

cs.LG20221 cited

DPAUC: Differentially Private AUC Computation in Federated Learning

Jiankai Sun, Xin Yang, Yuanshun Yao +3

Federated learning (FL) has gained significant attention recently as a privacy-enhancing tool to jointly train a machine learning model by multiple participants. The prior work on…