1 paper
Jianfei Ma, Zhaoxin Feng, Emmanuele Chersoni +1
Prompt compression shortens LLM input to reduce inference cost, yet existing methods score token importance through LM forward passes. It remains questionable whether such nuanced,…