3 citations · 4 across the 2 of their papers we have counts for
2 papers
cs.LG2024★ 3 cited
Get more for less: Principled Data Selection for Warming Up Fine-Tuning in LLMs
Feiyang Kang, Hoang Anh Just, Yifan Sun +5
This work focuses on leveraging and selecting from vast, unlabeled, open data to pre-fine-tune a pre-trained language model. The goal is to minimize the need for costly domain-spec…
cs.CR2022★ 1 cited
How to Sift Out a Clean Data Subset in the Presence of Data Poisoning?
Yi Zeng, Minzhou Pan, Himanshu Jahagirdar +3
Given the volume of data needed to train modern machine learning models, external suppliers are increasingly used. However, incorporating external data poses data poisoning risks,…