8 papers
SoK: Intent-Oriented Systematization of Multi-Turn LLM Jailbreaks
Siyuan Li, Aodu Wulianghai, Zehao Liu +8
Large Language Models (LLMs) are increasingly deployed in interactive settings, where user intent commonly unfolds through multi-turn dialogue. Multi-turn jailbreaks exploit this p…
Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks
Siyuan Li, Zehao Liu, Haoyu Li +5
As LLMs become increasingly integrated into complex applications, their vulnerability to adversarial attacks has raised significant concerns. However, existing defenses remain reac…
Lightweight Stylistic Consistency Profiling: Robust Detection of LLM-Generated Textual Content for Multimedia Moderation
Siyuan Li, Aodu Wulianghai, Xi Lin +6
The increasing prevalence of Large Language Models (LLMs) in content creation has made distinguishing human-written textual content from LLM-generated counterparts a critical task…
CoE: Collaborative Entropy for Uncertainty Quantification in Agentic Multi-LLM Systems
Kangkang Sun, Jun Wu, Jianhua Li +3
Uncertainty estimation in multi-LLM systems remains largely single-model-centric: existing methods quantify uncertainty within each model but do not adequately capture semantic dis…
HoneyTrap: Deceiving Large Language Model Attackers to Honeypot Traps with Resilient Multi-Agent Defense
Siyuan Li, Xi Lin, Jun Wu +5
Jailbreak attacks pose significant threats to large language models (LLMs), enabling attackers to bypass safeguards. However, existing reactive defense approaches struggle to keep…
BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems
Andy K. Zhang, Joey Ji, Celeste Menders +31
AI agents have the potential to significantly alter the cybersecurity landscape. Here, we introduce the first framework to capture offensive and defensive cyber-capabilities in evo…