19 citations · 20 across the 2 of their papers we have counts for
2 papers
cs.CL2024★ 19 cited
RedPajama: an Open Dataset for Training Large Language Models
Maurice Weber, Daniel Fu, Quentin Anthony +16
Large language models are increasingly becoming a cornerstone technology in artificial intelligence, the sciences, and society as a whole, yet the optimal strategies for dataset co…
cs.CL2023★ 1 cited
HelpSteer: Multi-attribute Helpfulness Dataset for SteerLM
Zhilin Wang, Yi Dong, Jiaqi Zeng +8
Existing open-source helpfulness preference datasets do not specify what makes some responses more helpful and others less so. Models trained on these datasets can incidentally lea…