Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Quantized Keys Steal Attention: Bias Correction for KV-Cache Compression in Video Diffusion
Tuna Tuncer, Felix Becker, Thomas Pfeil
Chunk-wise autoregressive video diffusion models rely on a KV cache of previously generated chunks to avoid redundant computation, but this cache quickly becomes a memory bottlenec…
cs.LG2024
Interactions Across Blocks in Post-Training Quantization of Large Language Models
Khasmamad Shabanovi, Lukas Wiest, Vladimir Golkov +2
Post-training quantization is widely employed to reduce the computational demands of neural networks. Typically, individual substructures, such as layers or blocks of layers, are q…