2 papers
cs.LG2025
DP-AdamW: Investigating Decoupled Weight Decay and Bias Correction in Private Deep Learning
Jay Chooi, Kevin Cong, Russell Li +1
As deep learning methods increasingly utilize sensitive data on a widespread scale, differential privacy (DP) offers formal guarantees to protect against information leakage during…
cs.LG2025
Optimal Inference Schedules for Masked Diffusion Models
Sitan Chen, Kevin Cong, Jerry Li
A major bottleneck of standard auto-regressive large language models is that their inference process is inherently sequential, resulting in very long and costly inference times. To…