From the 1 of 1 linked paper with an AI index.
1 paper
Rahul Krishnan, Volker Schulz
The paper introduces JoLT, a method that compresses the key‑value cache of transformer models by applying a partial Tucker decomposition on token and feature dimensions and adding…