Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Parcae: Scaling Laws For Stable Looped Language Models
Hayden Prairie, Zachary Novack, Taylor Berg-Kirkpatrick +1
Traditional fixed-depth architectures scale quality by increasing training FLOPs, typically through increased parameterization, at the expense of a higher memory footprint, or data…
cs.LG2021
Double Descent Optimization Pattern and Aliasing: Caveats of Noisy Labels
Florian Dubost, Erin Hong, Max Pike +5
Optimization plays a key role in the training of deep neural networks. Deciding when to stop training can have a substantial impact on the performance of the network during inferen…