2 papers
cs.AI2026
A Dual-Hypothesis Reasoning Framework for LLM Guardrails
Md Asiful Islam, Mihai Surdeanu
We propose ARBITER, a novel LLM guardrail framework that introduces two key ideas: (i) dual-hypothesis reasoning, a reasoning method for LLM guardrails that explicitly considers bo…
cs.CL2026
A Lightweight Explainable Guardrail for Prompt Safety
Md Asiful Islam, Mihai Surdeanu
We propose a lightweight explainable guardrail (LEG) method to detect unsafe prompts. LEG uses a multi-task learning architecture to jointly learn a prompt classifier and an explan…