4 citations · 4 across the 2 of their papers we have counts for
4 papers
Distill on a Diet: Efficient Knowledge Distillation via Learnable Data Pruning
Yifan Wu, Yiqi Wang, Xichen Ye +5
Knowledge Distillation (KD) is widely used to obtain compact models for efficient inference in resource-constrained environments. Yet the computational overhead of the distillation…
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
Xiaoyu Xu, Xiang Yue, Yang Liu +5
Unlearning in large language models (LLMs) aims to remove specified data, but its efficacy is typically assessed with task-level metrics like accuracy and perplexity. We show that…
Data Engineering for Scaling Language Models to 128K Context
Yao Fu, Rameswar Panda, Xinyao Niu +4
We study the continual pretraining recipe for scaling language models' context lengths to 128K, with a focus on data engineering. We hypothesize that long context modeling, in part…
Machine Unlearning of Pre-trained Large Language Models
Jin Yao, Eli Chien, Minxin Du +4
This study investigates the concept of the `right to be forgotten' within the context of large language models (LLMs). We explore machine unlearning as a pivotal solution, with a f…