1 paper
Zhengpei Hu, Kai Li, Dapeng Fu +5
Hard prompt compression reduces long-context inference cost by independently scoring tokens, sentences, or chunks and retaining the highest-scoring units under a budget. We identif…