3 citations · 5 across the 2 of their papers we have counts for
6 papers
Robust Multi-bit Text Watermark with LLM-based Paraphrasers
Xiaojun Xu, Jinghan Jia, Yuanshun Yao +2
We propose an imperceptible multi-bit text watermark embedded by paraphrasing with LLMs. We fine-tune a pair of LLM paraphrasers that are designed to behave differently so that the…
Learning to Watermark LLM-generated Text via Reinforcement Learning
Xiaojun Xu, Yuanshun Yao, Yang Liu
We study how to watermark LLM outputs, i.e. embedding algorithmically detectable signals into LLM-generated text to track misuse. Unlike the current mainstream methods that work wi…
Fairness Without Harm: An Influence-Guided Active Sampling Approach
Jinlong Pang, Jialu Wang, Zhaowei Zhu +3
The pursuit of fairness in machine learning (ML), ensuring that the models do not exhibit biases toward protected demographic groups, typically results in a compromise scenario. Th…
Measuring and Reducing LLM Hallucination without Gold-Standard Answers
Jiaheng Wei, Yuanshun Yao, Jean-Francois Ton +3
LLM hallucination, i.e. generating factually incorrect yet seemingly convincing answers, is currently a major threat to the trustworthiness and reliability of LLMs. The first step…
Human-Instruction-Free LLM Self-Alignment with Limited Samples
Hongyi Guo, Yuanshun Yao, Wei Shen +4
Aligning large language models (LLMs) with human values is a vital task for LLM practitioners. Current alignment techniques have several limitations: (1) requiring a large amount o…
Large Language Model Unlearning
Yuanshun Yao, Xiaojun Xu, Yang Liu
We study how to perform unlearning, i.e. forgetting undesirable misbehaviors, on large language models (LLMs). We show at least three scenarios of aligning LLMs with human preferen…