1 citations · 1 across the 3 of their papers we have counts for
5 papers
Users as Annotators: LLM Preference Learning from Comparison Mode
Zhongze Cai, Xiaocheng Li
Pairwise preference data have played an important role in the alignment of large language models (LLMs). Each sample of such data consists of a prompt, two different responses to t…
Incentivizing High-Quality Human Annotations with Golden Questions
Shang Liu, Zhongze Cai, Hanzhao Wang +2
Human-annotated data plays a vital role in training large language models (LLMs), such as supervised fine-tuning and human preference alignment. However, it is not guaranteed that…
Towards Better Statistical Understanding of Watermarking LLMs
Zhongze Cai, Shang Liu, Hanzhao Wang +2
In this paper, we study the problem of watermarking large language models (LLMs). We consider the trade-off between model distortion and detection ability and formulate it as a con…
What Matters in Data for DPO?
Yu Pan, Zhongze Cai, Guanting Chen +2
Direct Preference Optimization (DPO) has emerged as a simple and effective approach for aligning large language models (LLMs) with human preferences, bypassing the need for a learn…
Out-of-distribution Robust Optimization
Zhongze Cai, Hansheng Jiang, Xiaocheng Li
In this paper, we consider the contextual robust optimization problem under an out-of-distribution setting. The contextual robust optimization problem considers a risk-sensitive ob…