2 papers
cs.CR2026
Compiling Activation Steering into Weights via Null-Space Constraints for Stealthy Backdoors
Rui Yin, Tianxu Han, Naen Xu +8
Safety-aligned large language models (LLMs) are increasingly deployed in real-world pipelines, yet this deployment also enlarges the supply-chain attack surface: adversaries can di…
cs.CV2025
Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation
Zhenglin Hua, Jinghan He, Zijun Yao +4
Large vision-language models (LVLMs) have achieved remarkable performance on multimodal tasks. However, they still suffer from hallucinations, generating text inconsistent with vis…