2 papers
cs.LG2025
EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit Identification
Lin Zhang, Wenshuo Dong, Zhuoran Zhang +5
Understanding the internal mechanisms of transformer-based language models remains challenging. Mechanistic interpretability based on circuit discovery aims to reverse engineer neu…
cs.LG2024
Provably Efficient Action-Manipulation Attack Against Continuous Reinforcement Learning
Zhi Luo, Xiyuan Yang, Pan Zhou +1
Manipulating the interaction trajectories between the intelligent agent and the environment can control the agent's training and behavior, exposing the potential vulnerabilities of…