6 papers
Don't Make Models Guess Security and Safety: Symbolic Guardrails for Domain-Specific AI Agents
Yining Hong, Yining She, Eunsuk Kang +2
There is increasing interest in integrating AI agents that invoke tools into domain-specific commercial software, where unintended tool calls can cause serious security and safety…
FASR: Automated Identification of Unsafe Control Actions in STPA
Ian Dardik, Yining She, Sam Procter +3
The System-Theoretic Process Analysis (STPA) is a well-established hazard analysis technique that has been applied to a wide range of safety-critical systems. Despite its popularit…
Safeguarding LLM Agents from Misalignment through Provenance Analysis
Yining She, Yiliang Liang, Eunsuk Kang
As LLM agents gain increasing access to powerful tools, ensuring that their actions align with the user's intent becomes critical. When an agent's proposed action deviates from tha…
Towards Verifiably Safe Tool Use for LLM Agents
Aarya Doshi, Yining Hong, Congying Xu +3
Large language model (LLM)-based AI agents extend LLM capabilities by enabling access to tools such as data sources, APIs, search engines, code sandboxes, and even other agents. Wh…
RAG Makes Guardrails Unsafe? Investigating Robustness of Guardrails under RAG-style Contexts
Yining She, Daniel W. Peterson, Marianne Menglin Liu +4
With the increasing adoption of large language models (LLMs), ensuring the safety of LLM systems has become a pressing concern. External LLM-based guardrail models have emerged as…
FairSense: Long-Term Fairness Analysis of ML-Enabled Systems
Yining She, Sumon Biswas, Christian Kästner +1
Algorithmic fairness of machine learning (ML) models has raised significant concern in the recent years. Many testing, verification, and bias mitigation techniques have been propos…