1 paper
Sterling Huang, Abigayle Brown, Jiyoo Noh +4
Prompt compression reduces inference cost and context length in large language models, but prior evaluations focus mainly on autoregressive architectures. This study examines wheth…