5 papers
xPress: Parallel Refinement for Diffusion Drafters in Speculative Decoding
Zheng Wang, Davis Wertheimer, Yu Chin Fabian Lim +4
Block-diffusion drafters like dFlash generate an entire block of draft tokens in a single forward pass, drastically reducing the overhead of multiple-token drafting in speculative…
CaloChallenge 2022: A Community Challenge for Fast Calorimeter Simulation
Claudius Krause, Michele Faucci Giannelli, Gregor Kasieczka +66
We present the results of the "Fast Calorimeter Simulation Challenge 2022" - the CaloChallenge. We study state-of-the-art generative models on four calorimeter shower datasets of i…
Using Span Queries to Optimize for Cache and Attention Locality
Paul Castro, Nick Mitchell, Nathan Ordonez +3
Clients are evolving beyond chat completion, and now include a variety of innovative inference-time scaling and deep reasoning techniques. At the same time, inference servers remai…
Transforming the Hybrid Cloud for Emerging AI Workloads
Deming Chen, Alaa Youssef, Ruchi Pendse +42
This white paper, developed through close collaboration between IBM Research and UIUC researchers within the IIDAI Institute, envisions transforming hybrid cloud systems to meet th…
HadaCore: Tensor Core Accelerated Hadamard Transform Kernel
Krish Agarwal, Rishi Astra, Adnan Hoque +4
We present HadaCore, a modified Fast Walsh-Hadamard Transform (FWHT) algorithm optimized for the Tensor Cores present in modern GPU hardware. HadaCore follows the recursive structu…