4 papers
Stealthy World Model Manipulation via Data Poisoning
Yibin Hu, Xiaolin Sun, Zizhan Zheng
Model-based learning agents use learned world models to predict future states, plan actions, and adapt to new environments. However, the process of updating world models from colle…
Insider Attacks in Multi-Agent LLM Consensus Systems
Xiaolin Sun, Zixuan Liu, Yibin Hu +1
Large language models (LLMs) are increasingly deployed in multi-agent systems where agents communicate in natural language to solve tasks jointly. A key capability in such systems…
Robust Optimization for Mitigating Reward Hacking with Correlated Proxies
Zixuan Liu, Xiaolin Sun, Zizhan Zheng
Designing robust reinforcement learning (RL) agents in the presence of imperfect reward signals remains a core challenge. In practice, agents are often trained with proxy rewards t…
Diffusion Guided Adversarial State Perturbations in Reinforcement Learning
Xiaolin Sun, Feidi Liu, Zhengming Ding +1
Reinforcement learning (RL) systems, while achieving remarkable success across various domains, are vulnerable to adversarial attacks. This is especially a concern in vision-based…