7 papers
xPress: Parallel Refinement for Diffusion Drafters in Speculative Decoding
Zheng Wang, Davis Wertheimer, Yu Chin Fabian Lim +4
Block-diffusion drafters like dFlash generate an entire block of draft tokens in a single forward pass, drastically reducing the overhead of multiple-token drafting in speculative…
Transforming the Hybrid Cloud for Emerging AI Workloads
Deming Chen, Alaa Youssef, Ruchi Pendse +42
This white paper, developed through close collaboration between IBM Research and UIUC researchers within the IIDAI Institute, envisions transforming hybrid cloud systems to meet th…
HadaCore: Tensor Core Accelerated Hadamard Transform Kernel
Krish Agarwal, Rishi Astra, Adnan Hoque +4
We present HadaCore, a modified Fast Walsh-Hadamard Transform (FWHT) algorithm optimized for the Tensor Cores present in modern GPU hardware. HadaCore follows the recursive structu…
Next generation Co-Packaged Optics Technology to Train & Run Generative AI Models in Data Centers and Other Computing Applications
John Knickerbocker, Jean Benoit Heroux, Griselda Bonilla +15
We report on the successful design and fabrication of optical modules using a 50 micron pitch polymer waveguide interface, integrated for low loss, high density optical data transf…
Flexible and Effective Mixing of Large Language Models into a Mixture of Domain Experts
Rhui Dih Lee, Laura Wynter, Raghu Kiran Ganti
We present a toolkit for creating low-cost Mixture-of-Domain-Experts (MOE) from trained models. The toolkit can be used for creating a mixture from models or from adapters. We perf…
Enhancing Training Efficiency Using Packing with Flash Attention
Achintya Kundu, Rhui Dih Lee, Laura Wynter +2
Padding is often used in tuning LLM models by adding special tokens to shorter training examples to match the length of the longest sequence in each batch. While this ensures unifo…