2 papers
cs.SE2026
Speculative Refinement: A Hybrid Autoregressive Diffusion Decoding Strategy and Its Behavior Across Benchmarks
Aditi Gupta, Neel Mishra, Kushagra Trivedi +1
How should we evaluate generation systems that combine autoregressive (AR) and diffusion decoding? We study this question through Speculative Refinement (SpecRef), a training-free…
cs.LG2025
Hierarchical Sparse Plus Low Rank Compression of LLM
Pawan Kumar, Aditi Gupta
Modern large language models (LLMs) place extraordinary pressure on memory and compute budgets, making principled compression indispensable for both deployment and continued traini…