Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
Exploring the Personality Traits of LLMs through Latent Features Steering
Shu Yang, Shenzhe Zhu, Liang Liu +3
Large language models (LLMs) have significantly advanced dialogue systems and role-playing agents through their ability to generate human-like text. While prior studies have shown…
cs.CL2024
Dissecting Fine-Tuning Unlearning in Large Language Models
Yihuai Hong, Yuelin Zou, Lijie Hu +3
Fine-tuning-based unlearning methods prevail for preventing targeted harmful, sensitive, or copyrighted information within large language models while preserving overall capabiliti…
cs.CL2024
Private Language Models via Truncated Laplacian Mechanism
Tianhao Huang, Tao Yang, Ivan Habernal +2
Deep learning models for NLP tasks are prone to variants of privacy attacks. To prevent privacy leakage, researchers have investigated word-level perturbations, relying on the form…