1 paper
Honggen Zhang, Xufeng Zhao, Igor Molybog +1
Aligning large language models (LLMs) to human preferences is a crucial step in building helpful and safe AI tools, which usually involve training on supervised datasets. Popular a…