3 papers
cs.LG2026
IBCircuit: Towards Holistic Circuit Discovery with Information Bottleneck
Tian Bian, Yifan Niu, Chaohao Yuan +7
Circuit discovery has recently attracted attention as a potential research direction to explain the non-trivial behaviors of language models. It aims to find the computational subg…
cs.LG2026
Mitigating the Safety Alignment Tax with Null-Space Constrained Policy Optimization
Yifan Niu, Han Xiao, Dongyi Liu +2
As Large Language Models (LLMs) are increasingly deployed in real-world applications, it is important to ensure their behaviors align with human values, societal norms, and ethical…
cs.LG2025
Attacking and Securing Community Detection: A Game-Theoretic Framework
Yifan Niu, Aochuan Chen, Tingyang Xu +1
It has been demonstrated that adversarial graphs, i.e., graphs with imperceptible perturbations, can cause deep graph models to fail on classification tasks. In this work, we exten…