1 paper
Rheeya Uppaal, Seungwoo Lyu, Selina Sung +1
Safe completion requires models to provide useful assistance without enabling harm, but this behavior is difficult to evaluate with isolated prompts. We introduce OpenSafeIntent, a…