1 paper
Bilgehan Sel, Vaishakh Keshava, Phillip Wallis +3
Addressing the critical need for robust safety in Large Language Models (LLMs), particularly against adversarial attacks and in-distribution errors, we introduce Reinforcement Lear…