8 papers
Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges
Xianglin Yang, Bryan Hooi, Gelei Deng +2
The known stylistic biases in LLM judges, such as a preference for verbosity or specific sentence structures, present an underexplored security vulnerability. In this work, we intr…
Inverting the Shield: Systematically Generating Safety Tests from Policy Specifications
Xiaoyue Lu, Xianglin Yang, Haijun Liu +4
The widespread integration of Large Language Models (LLMs) necessitates rigorous and systematic safety evaluation. Existing paradigms either rely on constructed benchmarks to asses…
Zombie Agents: Persistent Control of Self-Evolving LLM Agents via Self-Reinforcing Injections
Xianglin Yang, Yufei He, Shuo Ji +2
Self-evolving LLM agents update their internal state across sessions, often by writing and reusing long-term memory. This design improves performance on long-horizon tasks but crea…
LLM-enabled Applications Require System-Level Threat Monitoring
Yedi Zhang, Haoyu Wang, Xianglin Yang +2
LLM-enabled applications are rapidly reshaping the software ecosystem by using large language models as core reasoning components for complex task execution. This paradigm shift, h…
Enhancing Model Defense Against Jailbreaks with Proactive Safety Reasoning
Xianglin Yang, Gelei Deng, Jieming Shi +2
Large language models (LLMs) are vital for a wide range of applications yet remain susceptible to jailbreak threats, which could lead to the generation of inappropriate responses.…
AutoContext: Instance-Level Context Learning for LLM Agents
Kuntai Cai, Juncheng Liu, Xianglin Yang +3
Current LLM agents typically lack instance-level context, which comprises concrete facts such as environment structure, system configurations, and local mechanics. Consequently, ex…