3 papers
cs.CY2026
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
Jing-Jing Li, Joel Mire, Eve Fleisig +4
Current AI safety frameworks, which often treat harmfulness as binary, lack the flexibility to handle borderline cases where humans meaningfully disagree. To build more pluralistic…
cs.CR2026
STAC: When Innocent Tools Form Dangerous Chains for LLM Agents
Jing-Jing Li, Jianfeng He, Chao Shang +6
As LLMs advance into autonomous agents with tool-use capabilities, they introduce security challenges that extend beyond traditional content-based LLM safety concerns. This paper i…
cs.CL2025
SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
Jing-Jing Li, Valentina Pyatkin, Max Kleiman-Weiner +7
The ideal AI safety moderation system would be both structurally interpretable (so its decisions can be reliably explained) and steerable (to align to safety standards and reflect…