Showing cs.CRShow all
3 papers · 1 filter
cs.CR2026
MAGE: Safeguarding LLM Agents against Long-Horizon Threats via Shadow Memory
Yuhui Wang, Tanqiu Jiang, Jiacheng Liang +2
As large language model (LLM)-powered agents are increasingly deployed to perform complex, real-world tasks, they face a growing class of attacks that exploit extended user-agent-e…
cs.CR2025
SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimization
Xuankun Rong, Wenke Huang, Tingfeng Wang +3
Multimodal large language models (MLLMs) have demonstrated impressive reasoning and instruction-following capabilities, yet their expanded modality space introduces new composition…
cs.CR2025
RAPID: Retrieval Augmented Training of Differentially Private Diffusion Models
Tanqiu Jiang, Changjiang Li, Fenglong Ma +1
Differentially private diffusion models (DPDMs) harness the remarkable generative capabilities of diffusion models while enforcing differential privacy (DP) for sensitive data. How…