3 papers
cs.LG2025
Reference-Specific Unlearning Metrics Can Hide the Truth: A Reality Check
Sungjun Cho, Dasol Hwang, Frederic Sala +3
Current unlearning metrics for generative models evaluate success based on reference responses or classifier outputs rather than assessing the core objective: whether the unlearned…
cs.LG2025
Why Alignment Must Precede Distillation: A Minimal Working Explanation
Sungmin Cha, Kyunghyun Cho
For efficiency, preference alignment is often performed on compact, knowledge-distilled (KD) models. We argue this common practice introduces a significant limitation by overlookin…
cs.LG2025
Why Knowledge Distillation Works in Generative Models: A Minimal Working Explanation
Sungmin Cha, Kyunghyun Cho
Knowledge distillation (KD) is a core component in the training and deployment of modern generative models, particularly large language models (LLMs). While its empirical benefits…