20 papers
When Claws Remember but Do Not Tell: Stealthy Memory Injection in Persistent Personal Agents
Yechao Zhang, Shiqian Zhao, Jiawen Zhang +5
Persistent personal agents combine long-term memory with access to users' external environments, enabling personalized foreground assistance and proactive background execution. Thi…
MaskForge: Structure-Aware Adaptive Attacks for Jailbreaking Diffusion Large Language Models
Yingzi Ma, Zhengyue Zhao, Xiaogeng Liu +3
Diffusion large language models (dLLMs) generate text by iteratively denoising partially masked sequences under bidirectional context, exposing a safety surface distinct from autor…
SafeGen-Bench: Benchmarking Safety in Image-Conditioned Text-to-Video Generation
Yingzi Ma, Xiaogeng Liu, Yawen Zheng +1
With the rapid advancements in text-to-image diffusion models, generative video models (T2V models) like Sora can now produce short synthetic videos from a text prompt or an initia…
When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning
Xiaogeng Liu, Xinyan Wang, Yingzi Ma +2
On-policy self-distillation (OPSD) trains a student on its own rollouts using a privileged teacher, but its standard objective weights all generated tokens equally, implicitly trea…
ROM: Real-time Overthinking Mitigation via Streaming Detection and Intervention
Xinyan Wang, Xiaogeng Liu, Ming Pei +1
Large Reasoning Models (LRMs) often reach a correct solution before their long Chain-of-Thought trace ends, yet continue with redundant verification, repeated attempts, or unnecess…
DecodingTrust-Agent Platform (DTap): A Controllable and Interactive Red-Teaming Platform for AI Agents
Zhaorun Chen, Xun Liu, Haibo Tong +14
AI agents are increasingly deployed across diverse domains to automate complex workflows through long-horizon and high-stakes action executions. Due to their high capability and fl…