2 citations · 4 across the 3 of their papers we have counts for
3 papers
cs.CV2023★ 2 cited
Can pre-trained models assist in dataset distillation?
Yao Lu, Xuguang Chen, Yuchen Zhang +7
Dataset Distillation (DD) is a prominent technique that encapsulates knowledge from a large-scale original dataset into a small synthetic dataset for efficient training. Meanwhile,…
cs.LG2023★ 2 cited
A Theory on Adam Instability in Large-Scale Machine Learning
Igor Molybog, Peter Albert, Moya Chen +14
We present a theory for the previously unexplained divergent behavior noticed in the training of large language models. We argue that the phenomenon is an artifact of the dominant…
cs.CL2021
The Influence of Data Pre-processing and Post-processing on Long Document Summarization
Xinwei Du, Kailun Dong, Yuchen Zhang +2
Long document summarization is an important and hard task in the field of natural language processing. A good performance of the long document summarization reveals the model has a…