1 paper
John Harvill, Ziwei Fan, Hao Wang +4
Existing work on prompt compression for Large Language Models (LLM) focuses on lossy methods that try to maximize the retention of semantic information that is relevant to downstre…