4 papers · 1 filter
Forgotten in Weights, Recovered by Tools: Agentic Tool Unlearning for LLM Agents
Baicheng Chen, Zheyuan Liu, Jingyu Zhang +4
Large language models (LLMs) are increasingly deployed as tool-augmented agents, where responses can depend on tool calls and external observations rather than model parameters alo…
Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions
Diancheng Kang, Zheyuan Liu, Ningshan Ma +3
Activation steering controls language model behavior by adding directions to internal representations at inference time, but standard residual-stream steering can fail in stateful…
The Chameleon's Limit: Investigating Persona Collapse and Homogenization in Large Language Models
Yunze Xiao, Vivienne J. Zhang, Chenghao Yang +3
Applications based on large language models (LLMs), such as multi-agent simulations, require population diversity among agents. We identify a pervasive failure mode we term \emph{P…
How Private are Language Models in Abstractive Summarization?
Anthony Hughes, Ning Ma, Nikolaos Aletras
In sensitive domains such as medical and legal, protecting sensitive information is critical, with protective laws strictly prohibiting the disclosure of personal data. This poses…