3 papers
cs.SE2026
Speculative Refinement: A Hybrid Autoregressive Diffusion Decoding Strategy and Its Behavior Across Benchmarks
Aditi Gupta, Neel Mishra, Kushagra Trivedi +1
How should we evaluate generation systems that combine autoregressive (AR) and diffusion decoding? We study this question through Speculative Refinement (SpecRef), a training-free…
cs.LG2026
SignMuon: Communication-Efficient Distributed Muon Optimization
Neel Mishra, Kushagara Trivedi, Pawan Kumar
Distributed training of large neural networks is bottlenecked by full-precision gradient communication and by coordinatewise optimizers that ignore the matrix structure of weight t…
cs.LG2024
A Gauss-Newton Approach for Min-Max Optimization in Generative Adversarial Networks
Neel Mishra, Bamdev Mishra, Pratik Jawanpuria +1
A novel first-order method is proposed for training generative adversarial networks (GANs). It modifies the Gauss-Newton method to approximate the min-max Hessian and uses the Sher…