1 paper · 1 filter
Lijun Zhang, Lin Li, Yajie Qi +4
When fine-tuning pre-trained Large Language Models (LLMs) to align with human values and intentions, maximizing the estimated reward can lead to superior performance, but it also i…