1 paper · 1 filter
Bat-Sheva Einbinder, Hen Davidov, Yee Whye Teh +2
While large language models (LLMs) are trained to align with human values, their generations may still violate safety constraints. A growing line of work addresses this problem by…