3 papers
cs.CL2025
Approximating Language Model Training Data from Weights
John X. Morris, Junjie Oscar Yin, Woojeong Kim +2
Modern language models often have open weights but closed training data. We formalize the problem of data approximation from model weights and propose several baselines and metrics…
cs.LG2025
Compute-Constrained Data Selection
Junjie Oscar Yin, Alexander M. Rush
Data selection can reduce the amount of training data needed to finetune LLMs; however, the efficacy of data selection scales directly with its compute. Motivated by the practical…
cs.LG2024
NetFlowGen: Leveraging Generative Pre-training for Network Traffic Dynamics
Jiawei Zhou, Woojeong Kim, Zhiying Xu +2
Understanding the traffic dynamics in networks is a core capability for automated systems to monitor and analyze networking behaviors, reducing expensive human efforts and economic…