103 citations · 335 across the 29 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2023
BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset
Jiaming Ji, Mickel Liu, Juntao Dai +6
In this paper, we introduce the BeaverTails dataset, aimed at fostering research on safety alignment in large language models (LLMs). This dataset uniquely separates annotations of…
cs.CL2023
Heterogeneous Value Alignment Evaluation for Large Language Models
Zhaowei Zhang, Ceyao Zhang, Nian Liu +5
The emergent capabilities of Large Language Models (LLMs) have made it crucial to align their values with those of humans. However, current methodologies typically attempt to assig…