Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs
Di He, Songjun Tu, Keyu Wang +2
Learning rate configuration is a fundamental aspect of modern deep learning. The prevailing practice of applying a uniform learning rate across all layers overlooks the structural…
cs.LG2026
Quotient-Space Diffusion Models
Yixian Xu, Yusong Wang, Shengjie Luo +4
Diffusion-based generative models have reformed generative AI, and also enabled new capabilities in the science domain, e.g., fast generation of 3D structures of molecules. In such…
cs.LG2026
Lossless Anti-Distillation Sampling
Zibo Diao, Jingchu Gai, Xinyue Ai +3
Frontier commercial generative models face a growing threat from distillation, whereby a distiller harvests generated responses and trains a competing model of its own at drastical…