4 papers
Zombie Agents: Persistent Control of Self-Evolving LLM Agents via Self-Reinforcing Injections
Xianglin Yang, Yufei He, Shuo Ji +2
Self-evolving LLM agents update their internal state across sessions, often by writing and reusing long-term memory. This design improves performance on long-horizon tasks but crea…
Enhancing Model Defense Against Jailbreaks with Proactive Safety Reasoning
Xianglin Yang, Gelei Deng, Jieming Shi +2
Large language models (LLMs) are vital for a wide range of applications yet remain susceptible to jailbreak threats, which could lead to the generation of inappropriate responses.…
TSRE: Channel-Aware Typical Set Refinement for Out-of-Distribution Detection
Weijun Gao, Rundong He, Jinyang Dong +1
Out-of-Distribution (OOD) detection is a critical capability for ensuring the safe deployment of machine learning models in open-world environments, where unexpected or anomalous i…
Neural Surveillance: Live-Update Visualization of Latent Training Dynamics
Xianglin Yang, Jin Song Dong
Monitoring the inner state of deep neural networks is essential for auditing the learning process and enabling timely interventions. While conventional metrics like validation loss…