4 citations · 4 across the 1 of their papers we have counts for
3 papers · 1 filter
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
Jonas Geiping, Sean McLeish, Neel Jain +6
We study a novel language model architecture that is capable of scaling test-time computation by implicitly reasoning in latent space. Our model works by iterating a recurrent bloc…
A Winning Hand: Compressing Deep Networks Can Improve Out-Of-Distribution Robustness
James Diffenderfer, Brian R. Bartoldson, Shreya Chaganti +2
Successful adoption of deep learning (DL) in the wild requires models to be: (1) compact, (2) accurate, and (3) robust to distributional shifts. Unfortunately, efforts towards simu…
The Generalization-Stability Tradeoff In Neural Network Pruning
Brian R. Bartoldson, Ari S. Morcos, Adrian Barbu +1
Pruning neural network parameters is often viewed as a means to compress models, but pruning has also been motivated by the desire to prevent overfitting. This motivation is partic…