Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
xPress: Parallel Refinement for Diffusion Drafters in Speculative Decoding
Zheng Wang, Davis Wertheimer, Yu Chin Fabian Lim +4
Block-diffusion drafters like dFlash generate an entire block of draft tokens in a single forward pass, drastically reducing the overhead of multiple-token drafting in speculative…
cs.AI2025
Using Span Queries to Optimize for Cache and Attention Locality
Paul Castro, Nick Mitchell, Nathan Ordonez +3
Clients are evolving beyond chat completion, and now include a variety of innovative inference-time scaling and deep reasoning techniques. At the same time, inference servers remai…