Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Safeguarding LLM Agents from Misalignment through Provenance Analysis
Yining She, Yiliang Liang, Eunsuk Kang
As LLM agents gain increasing access to powerful tools, ensuring that their actions align with the user's intent becomes critical. When an agent's proposed action deviates from tha…
cs.CL2025
RAG Makes Guardrails Unsafe? Investigating Robustness of Guardrails under RAG-style Contexts
Yining She, Daniel W. Peterson, Marianne Menglin Liu +4
With the increasing adoption of large language models (LLMs), ensuring the safety of LLM systems has become a pressing concern. External LLM-based guardrail models have emerged as…