5 papers
A Non-compact Positivity-Preserving Scheme for Parabolic PDE via Conditional Expectation
Haoran Xu, Jie Ren, Xingye Yue
We propose a novel non-compact, positivity-preserving scheme for linear non-divergence form parabolic equations. Based on the Feynman-Kac formula, the solution is expressed as a co…
FlashDP: Private Training Large Language Models with Efficient DP-SGD
Liangyu Wang, Junxiao Wang, Jie Ren +3
As large language models (LLMs) increasingly underpin technological advancements, the privacy of their training data emerges as a critical concern. Differential Privacy (DP) serves…
On the Cone Effect in the Learning Dynamics
Zhanpeng Zhou, Yongyi Yang, Jie Ren +2
Understanding the learning dynamics of neural networks is a central topic in the deep learning community. In this paper, we take an empirical perspective to study the learning dyna…
ZO2: Scalable Zeroth-Order Fine-Tuning for Extremely Large Language Models with Limited GPU Memory
Liangyu Wang, Jie Ren, Hang Xu +4
Fine-tuning large pre-trained LLMs generally demands extensive GPU memory. Traditional first-order optimizers like SGD encounter substantial difficulties due to increased memory re…
Accelerating Mixed-Precision Out-of-Core Cholesky Factorization with Static Task Scheduling
Jie Ren, Hatem Ltaief, Sameh Abdulah +1
This paper explores the performance optimization of out-of-core (OOC) Cholesky factorization on shared-memory systems equipped with multiple GPUs. We employ fine-grained computatio…