1 citations · 1 across the 4 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
AgentLens: Interpretable Safety Steering via Mechanistic Subspaces for Multi-Turn Coding Agent
Weidi Luo, Qiming Zhang, Yihao Quan +5
Coding agents based on large language models (LLMs) demonstrate remarkable autonomous capabilities, but they also introduce significant safety and misuse risks during multi-turn in…
cs.AI2026
Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs
Qitao Tan, Xiaoying Song, Arman Akbari +7
Current safety alignment of foundation models largely follows a \emph{one-size-fits-all} paradigm, applying the same refusal policy across users and contexts. As a result, models m…