3 papers
cs.AI2026
DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model
Wenhao Lin, Chenyu Yu, Xingwei Lin +6
As large language model (LLM) agents increasingly invoke external tools and interact with real-world systems, unsafe actions may cause irreversible consequences on external states,…
cs.CR2026
Lorica: A Synergistic Fine-Tuning Framework for Advancing Personalized Adversarial Robustness
Tianyu Qi, Lei Xue, Yufeng Zhan +1
The growing use of large pre-trained models in edge computing has made model inference on mobile clients both feasible and popular. Yet these devices remain vulnerable to adversari…
cs.CR2026
ICON: Intent-Context Coupling for Efficient Multi-Turn Jailbreak Attack
Xingwei Lin, Wenhao Lin, Sicong Cao +4
Multi-turn jailbreak attacks have emerged as a critical threat to Large Language Models (LLMs), bypassing safety mechanisms by progressively constructing adversarial contexts from…