10 papers
Vulnerable Agent Identification in Large-Scale Multi-Agent Reinforcement Learning
Simin Li, Zihao Mao, Zheng Yuwei +12
Partial agent failure becomes inevitable when systems scale up, making it crucial to identify the subset of agents whose failure causes worst-case system performance degradations.…
GuardAD: Safeguarding Autonomous Driving MLLMs via Markovian Safety Logic
Tianyuan Zhang, Peng Yue, Zihao Peng +8
Multimodal large language models (MLLMs) are increasingly integrated into autonomous driving (AD) systems; however, they remain vulnerable to diverse safety threats, particularly i…
AgentVisor: Defending LLM Agents Against Prompt Injection via Semantic Virtualization
Zonghao Ying, Haozheng Wang, Jiangfan Liu +5
Large Language Model (LLM) agents are increasingly used to automate complex workflows, but integrating untrusted external data with privileged execution exposes them to severe secu…
Reading Between the Pixels: An Inscriptive Jailbreak Attack on Text-to-Image Models
Zonghao Ying, Haowen Dai, Lianyu Hu +5
Modern text-to-image (T2I) models can now render legible, paragraph-length text, enabling a fundamentally new class of misuse. We identify and formalize the inscriptive jailbreak,…
Evolving Deception: When Agents Evolve, Deception Wins
Zonghao Ying, Haowen Dai, Tianyuan Zhang +6
Self-evolving agents offer a promising path toward scalable autonomy. However, in this work, we show that in competitive environments, self-evolution can instead give rise to a ser…
Does LLM Alignment Really Need Diversity? An Empirical Study of Adapting RLVR Methods for Moral Reasoning
Zhaowei Zhang, Xiaohan Liu, Xuekai Zhu +6
Reinforcement learning with verifiable rewards (RLVR) has achieved remarkable success in logical reasoning tasks, yet whether large language model (LLM) alignment requires fundamen…