Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Certified Causal Defense with Generalizable Robustness
Yiran Qiao, Yu Yin, Chen Chen +1
While machine learning models have proven effective across various scenarios, it is widely acknowledged that many models are vulnerable to adversarial attacks. Recently, there have…
cs.LG2025
SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
Zirui He, Haiyan Zhao, Yiran Qiao +4
The ability of large language models (LLMs) to follow instructions is crucial for their practical applications, yet the underlying mechanisms remain poorly understood. This paper p…