1 paper · 1 filter
Sterling Huang, Abigayle Brown, Jiyoo Noh +4
Prompt compression reduces inference cost and context length in large language models, but prior evaluations focus mainly on autoregressive architectures. This study examines wheth…