2 papers
cs.CV2025
SQ-DM: Accelerating Diffusion Models with Aggressive Quantization and Temporal Sparsity
Zichen Fan, Steve Dai, Rangharajan Venkatesan +2
Diffusion models have gained significant popularity in image generation tasks. However, generating high-quality content remains notably slow because it requires running model infer…
cs.AR2024
ConSmax: Hardware-Friendly Alternative Softmax with Learnable Parameters
Shiwei Liu, Guanchen Tao, Yifei Zou +7
The self-attention mechanism distinguishes transformer-based large language models (LLMs) apart from convolutional and recurrent neural networks. Despite the performance improvemen…