Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
HiFloat4 Format for Language Model Inference
Yuanyong Luo, Jing Huang, Yu Cheng +19
This paper introduces HiFloat4 (HiF4), a block floating-point data format tailored for deep learning. Each HiF4 unit packs 64 4-bit elements with 32 bits of shared scaling metadata…
cs.LG2024
Learning-to-Cache: Accelerating Diffusion Transformer via Layer Caching
Xinyin Ma, Gongfan Fang, Michael Bi Mi +1
Diffusion Transformers have recently demonstrated unprecedented generative capabilities for various tasks. The encouraging results, however, come with the cost of slow inference, s…