1 paper
Tien-Phat Nguyen, Truong Nguyen, Thin Nguyen +3
Aligning language models for both helpfulness and safety typically requires complex pipelines-separate reward and cost models, online reinforcement learning, and primal-dual update…