3 papers
cs.AI2026
xPress: Parallel Refinement for Diffusion Drafters in Speculative Decoding
Zheng Wang, Davis Wertheimer, Yu Chin Fabian Lim +4
Block-diffusion drafters like dFlash generate an entire block of draft tokens in a single forward pass, drastically reducing the overhead of multiple-token drafting in speculative…
cs.LG2026
FlashSinkhorn: IO-Aware Entropic Optimal Transport on GPU
Felix X. -F. Ye, Xingjie Li, An Yu +3
Entropic optimal transport (EOT) via Sinkhorn iterations is widely used in modern machine learning, yet GPU solvers remain inefficient at scale. Tensorized implementations suffer q…
cs.LG2026
Frayed RoPE and Long Inputs: A Geometric Perspective
Davis Wertheimer, Aozhong Zhang, Derrick Liu +2
Rotary Positional Embedding (RoPE) is a widely adopted technique for encoding position in language models, which, while effective, causes performance breakdown when input length ex…