15 citations · 21 across the 5 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2023★ 1 cited
HelpSteer: Multi-attribute Helpfulness Dataset for SteerLM
Zhilin Wang, Yi Dong, Jiaqi Zeng +8
Existing open-source helpfulness preference datasets do not specify what makes some responses more helpful and others less so. Models trained on these datasets can incidentally lea…
cs.CL2023★ 3 cited
SteerLM: Attribute Conditioned SFT as an (User-Steerable) Alternative to RLHF
Yi Dong, Zhilin Wang, Makesh Narsimhan Sreedhar +2
Model alignment with human preferences is an essential step in making Large Language Models (LLMs) helpful and consistent with human values. It typically consists of supervised fin…