1 paper
Zhuoyuan Li, Rui Zhao, Jin Wang +5
Token compression has become a key technique for reducing the inference cost of large foundation models, with approaches such as token pruning and KV-cache reuse widely adopted in…