Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Stochastic Sparse Attention for Memory-Bound Inference
Kyle Lee, Corentin Delacour, Kevin Callahan-Coray +5
Autoregressive decoding becomes bandwidth-limited at long contexts, as generating each token requires reading all key and value vectors from KV cache. We present Stochastic A…
cs.LG2025
IsingFormer: Augmenting Parallel Tempering With Learned Proposals
Saleh Bunaiyan, Corentin Delacour, Shuvro Chowdhury +2
Markov Chain Monte Carlo (MCMC) underlies both statistical physics and combinatorial optimization, but mixes slowly near critical points and in rough landscapes. Parallel Tempering…