40 citations · 117 across the 21 of their papers we have counts for
1 paper · 1 filter
Aryan Keluskar, Amrita Bhattacharjee, Huan Liu
Safety alignment in LLMs aims to align models with human values, but which values take precedence when they conflict? We investigate this question in the context of tool-calling LL…