3 papers
cs.LG2026
Double Preconditioning (DoPr): Optimization for Test-Time Performance, not Validation Loss
Thomas T. Zhang, Alok Shah, Yifei Zhang +3
Many modern applications of deep learning involve training a neural network via a one-step prediction loss (e.g., regression, cross-entropy), but deploy the network by rollin…
cs.LG2025
Action Chunking and Exploratory Data Collection Yield Exponential Improvements in Behavior Cloning for Continuous Control
Thomas T. Zhang, Daniel Pfrommer, Chaoyi Pan +2
This paper presents a theoretical analysis of two of the most impactful interventions in modern learning from demonstration in robotics and continuous control: the practice of acti…
cs.LG2025
On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning
Thomas T. Zhang, Behrad Moniri, Ansh Nagwekar +4
Layer-wise preconditioning methods are a family of memory-efficient optimization algorithms that introduce preconditioners per axis of each layer's weight tensors. These methods ha…