2 papers
cs.CL2026
Predicting Deep Neural Network Training Outcomes from Early Training Telemetry
Ranjita Naik, Anh D. Nguyen, Pankaj Kumar Singh
Large hyperparameter sweeps for deep neural networks spend substantial compute on configurations that are effectively doomed from the first few epochs. We study whether a single tr…
cs.LG2025
Layer-wise Quantization for Quantized Optimistic Dual Averaging
Anh Duc Nguyen, Ilia Markov, Frank Zhengqing Wu +4
Modern deep neural networks exhibit heterogeneity across numerous layers of various types such as residuals, multi-head attention, etc., due to varying structures (dimensions, acti…