2 papers
cs.LG2026
Double Preconditioning (DoPr): Optimization for Test-Time Performance, not Validation Loss
Thomas T. Zhang, Alok Shah, Yifei Zhang +3
Many modern applications of deep learning involve training a neural network via a one-step prediction loss (e.g., regression, cross-entropy), but deploy the network by rollin…
cs.CL2025
Language Modeling with Learned Meta-Tokens
Alok N. Shah, Khush Gupta, Keshav Ramji +1
While modern Transformer-based language models (LMs) have achieved major success in multi-task generalization, they often struggle to capture long-range dependencies within their c…