From the 1 of 7 linked papers with an AI index.
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Beyond Steering Vector: Flow-based Activation Steering for Inference-Time Intervention
Zehao Jin, Ruixuan Deng, Junran Wang +2
Activation steering has emerged as a promising alternative for controlling language-model behavior at inference time by modifying intermediate representations while keeping model p…
cs.CL2026
Sparse Feature Coactivation Reveals Causal Semantic Modules in Large Language Models
Ruixuan Deng, Xiaoyang Hu, Miles Gilberti +5
We identify semantically coherent, context-consistent network components in large language models (LLMs) using coactivation of sparse autoencoder (SAE) features collected from just…