4 papers
Don't Make Models Guess Security and Safety: Symbolic Guardrails for Domain-Specific AI Agents
Yining Hong, Yining She, Eunsuk Kang +2
There is increasing interest in integrating AI agents that invoke tools into domain-specific commercial software, where unintended tool calls can cause serious security and safety…
Safeguarding LLM Agents from Misalignment through Provenance Analysis
Yining She, Yiliang Liang, Eunsuk Kang
As LLM agents gain increasing access to powerful tools, ensuring that their actions align with the user's intent becomes critical. When an agent's proposed action deviates from tha…
RAG Makes Guardrails Unsafe? Investigating Robustness of Guardrails under RAG-style Contexts
Yining She, Daniel W. Peterson, Marianne Menglin Liu +4
With the increasing adoption of large language models (LLMs), ensuring the safety of LLM systems has become a pressing concern. External LLM-based guardrail models have emerged as…
FairSense: Long-Term Fairness Analysis of ML-Enabled Systems
Yining She, Sumon Biswas, Christian Kästner +1
Algorithmic fairness of machine learning (ML) models has raised significant concern in the recent years. Many testing, verification, and bias mitigation techniques have been propos…