5 papers
TWEO: Transformers Without Extreme Outliers Enables FP8 Training And Quantization For Dummies
Guang Liang, Jie Shao, Ningyuan Tang +2
Native FP8 support in modern hardware is essential for training large Transformers, but is severely hindered by extreme activation outliers. Existing solutions either rely on compl…
Images Speak Louder Than Scores: Failure Mode Escape for Enhancing Generative Quality
Jie Shao, Ke Zhu, Minghao Fu +2
Diffusion models have achieved remarkable progress in class-to-image generation. However, we observe that despite impressive FID scores, state-of-the-art models often generate dist…
Quantization without Tears
Minghao Fu, Hao Yu, Jie Shao +3
Deep neural networks, while achieving remarkable success across diverse tasks, demand significant resources, including computation, GPU memory, bandwidth, storage, and energy. Netw…
Who Reasons in the Large Language Models?
Jie Shao, Jianxin Wu
Despite the impressive performance of large language models (LLMs), the process of endowing them with new capabilities--such as mathematical reasoning--remains largely empirical an…
Diffusion Product Quantization
Jie Shao, Hanxiao Zhang, Jianxin Wu
In this work, we explore the quantization of diffusion models in extreme compression regimes to reduce model size while maintaining performance. We begin by investigating classical…