3 papers
cs.GT2025
Incentivizing High-Quality Human Annotations with Golden Questions
Shang Liu, Zhongze Cai, Hanzhao Wang +2
Human-annotated data plays a vital role in training large language models (LLMs), such as supervised fine-tuning and human preference alignment. However, it is not guaranteed that…
cs.LG2025
How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators
Shang Liu, Hanzhao Wang, Zhongyao Ma +1
Human-annotated preference data play an important role in aligning large language models (LLMs). In this paper, we study two connected questions: how to monitor the quality of huma…
math.OC2024
Out-of-distribution Robust Optimization
Zhongze Cai, Hansheng Jiang, Xiaocheng Li
In this paper, we consider the contextual robust optimization problem under an out-of-distribution setting. The contextual robust optimization problem considers a risk-sensitive ob…