7 papers
MoRE: Mixture of Reused Experts
Eric S. Qiu, Utku Umur Acikalin, Justin Lovelace +4
Mixture-of-Experts (MoE) architectures decouple model capacity from computational cost, yet incur high memory footprints as parameters grow linearly with the number of experts. Rec…
Optimize Your Sampling: Tuned Diffusion Sampling with Bayesian Optimization
Travis Zhang, Christian Belardi, Justin Lovelace +4
Sampling from a diffusion model typically requires many forward passes through a large neural network, making generation computationally expensive. While much work has focused on e…
Prescriptive Scaling Laws for Data Constrained Training
Justin Lovelace, Christian Belardi, Srivatsa Kundurthy +2
Training compute is increasingly outpacing the availability of high-quality data. This shifts the central challenge from optimal compute allocation to extracting maximum value from…
Adaptive Moments are Surprisingly Effective for Plug-and-Play Diffusion Sampling
Christian Belardi, Justin Lovelace, Kilian Q. Weinberger +1
Guided diffusion sampling relies on approximating often intractable likelihood scores, which introduces significant noise into the sampling dynamics. We propose using adaptive mome…
Improving Multislice Electron Ptychography with a Generative Prior
Christian K. Belardi, Chia-Hao Lee, Yingheng Wang +4
Multislice electron ptychography (MEP) is an inverse imaging technique that computationally reconstructs the highest-resolution images of atomic crystal structures from diffraction…
Pre-training Limited Memory Language Models with Internal and External Knowledge
Linxi Zhao, Sofian Zalouk, Christian K. Belardi +7
Neural language models are black-boxes--both linguistic patterns and factual knowledge are distributed across billions of opaque parameters. This entangled encoding makes it diffic…