5 papers
Beyond Forgetting: Machine Unlearning Elicits Controllable Side Behaviors and Capabilities
Tien Dang, The-Hai Nguyen, Dinh Mai Phuong +5
We consider Representation Misdirection (RM), a class of large language model (LLM) unlearning methods that achieve forgetting by redirecting the forget-representations, that is, l…
Improving LLM Unlearning Robustness via Random Perturbations
Dang Huu-Tien, Hoang Thanh-Tung, Anh Bui +3
Here, we show that current LLM unlearning methods inherently reduce models' robustness, causing them to misbehave even when a single non-adversarial forget-token is present in the…
Improving Chain-of-Thought for Logical Reasoning via Attention-Aware Intervention
Nguyen Minh Phuong, Dang Huu Tien, Naoya Inoue
Modern logical reasoning with LLMs primarily relies on employing complex interactive frameworks that decompose the reasoning process into subtasks solved through carefully designed…
Detecting and Rectifying Noisy Labels: A Similarity-based Approach
Dang Huu-Tien, Minh-Phuong Nguyen, Naoya Inoue
Label noise in datasets could significantly damage the performance and robustness of deep neural networks (DNNs) trained on these datasets. As the size of modern DNNs grows, there…
On Effects of Steering Latent Representation for Large Language Model Unlearning
Dang Huu-Tien, Trung-Tin Pham, Hoang Thanh-Tung +1
Representation Misdirection for Unlearning (RMU), which steers model representation in the intermediate layer to a target random representation, is an effective method for large la…