1 paper
Yitao Jiang, Yaoqing Yang, Luyang Zhao +2
Long-context Transformer inference increasingly relies on KV-cache compression or quantization. Prior rotation and transform-coding results suggest that the channel basis of each k…