2 papers
cs.LG2026
Adaptive Helpfulness-Harmlessness Alignment with Preference Vectors
Ren-Wei Liang, Chin-Ting Hsu, Chan-Hung Yu +6
Ensuring that large language models (LLMs) are both helpful and harmless is a critical challenge, as overly strict constraints can lead to excessive refusals, while permissive mode…
cs.CL2025
SHA256 at SemEval-2025 Task 4: Selective Amnesia -- Constrained Unlearning for Large Language Models via Knowledge Isolation
Saransh Agrawal, Kuan-Hao Huang
Large language models (LLMs) frequently memorize sensitive information during training, posing risks when deploying publicly accessible models. Current machine unlearning methods s…