6 papers
DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model
Wenhao Lin, Chenyu Yu, Xingwei Lin +6
As large language model (LLM) agents increasingly invoke external tools and interact with real-world systems, unsafe actions may cause irreversible consequences on external states,…
SoK: Intent-Oriented Systematization of Multi-Turn LLM Jailbreaks
Siyuan Li, Aodu Wulianghai, Zehao Liu +8
Large Language Models (LLMs) are increasingly deployed in interactive settings, where user intent commonly unfolds through multi-turn dialogue. Multi-turn jailbreaks exploit this p…
Lightweight Stylistic Consistency Profiling: Robust Detection of LLM-Generated Textual Content for Multimedia Moderation
Siyuan Li, Aodu Wulianghai, Xi Lin +6
The increasing prevalence of Large Language Models (LLMs) in content creation has made distinguishing human-written textual content from LLM-generated counterparts a critical task…
HoneyTrap: Deceiving Large Language Model Attackers to Honeypot Traps with Resilient Multi-Agent Defense
Siyuan Li, Xi Lin, Jun Wu +5
Jailbreak attacks pose significant threats to large language models (LLMs), enabling attackers to bypass safeguards. However, existing reactive defense approaches struggle to keep…
StyleDecipher: Robust and Explainable Detection of LLM-Generated Texts with Stylistic Analysis
Siyuan Li, Aodu Wulianghai, Xi Lin +4
With the increasing integration of large language models (LLMs) into open-domain writing, detecting machine-generated text has become a critical task for ensuring content authentic…
Trustworthy AI-Generative Content for Intelligent Network Service: Robustness, Security, and Fairness
Siyuan Li, Xi Lin, Yaju Liu +2
AI-generated content (AIGC) models, represented by large language models (LLM), have revolutionized content creation. High-speed next-generation communication technology is an idea…