most citedLearning to Watermark LLM-generated Text via Reinforcement Learning

3 citations · 5 across the 2 of their papers we have counts for

collaborators

6 papers

cs.AI2024

Robust Multi-bit Text Watermark with LLM-based Paraphrasers

Xiaojun Xu, Jinghan Jia, Yuanshun Yao +2

We propose an imperceptible multi-bit text watermark embedded by paraphrasing with LLMs. We fine-tune a pair of LLM paraphrasers that are designed to behave differently so that the…

cs.LG20243 cited

Learning to Watermark LLM-generated Text via Reinforcement Learning

Xiaojun Xu, Yuanshun Yao, Yang Liu

We study how to watermark LLM outputs, i.e. embedding algorithmically detectable signals into LLM-generated text to track misuse. Unlike the current mainstream methods that work wi…

cs.LG2024

Fairness Without Harm: An Influence-Guided Active Sampling Approach

Jinlong Pang, Jialu Wang, Zhaowei Zhu +3

The pursuit of fairness in machine learning (ML), ensuring that the models do not exhibit biases toward protected demographic groups, typically results in a compromise scenario. Th…

cs.CL2024

Measuring and Reducing LLM Hallucination without Gold-Standard Answers

Jiaheng Wei, Yuanshun Yao, Jean-Francois Ton +3

LLM hallucination, i.e. generating factually incorrect yet seemingly convincing answers, is currently a major threat to the trustworthiness and reliability of LLMs. The first step…

cs.CL20242 cited

Human-Instruction-Free LLM Self-Alignment with Limited Samples

Hongyi Guo, Yuanshun Yao, Wei Shen +4

Aligning large language models (LLMs) with human values is a vital task for LLM practitioners. Current alignment techniques have several limitations: (1) requiring a large amount o…

cs.CL2023

Large Language Model Unlearning

Yuanshun Yao, Xiaojun Xu, Yang Liu

We study how to perform unlearning, i.e. forgetting undesirable misbehaviors, on large language models (LLMs). We show at least three scenarios of aligning LLMs with human preferen…