1 paper · 1 filter
Md Asiful Islam, Mihai Surdeanu
We propose a lightweight explainable guardrail (LEG) method to detect unsafe prompts. LEG uses a multi-task learning architecture to jointly learn a prompt classifier and an explan…