1 paper
Hritik Bansal, John Dang, Aditya Grover
Aligning large language models (LLMs) with human values and intents critically involves the use of human or AI feedback. While dense feedback annotations are expensive to acquire a…