1 paper · 1 filter
Tinghui Zhang, Yifan Wang, Daisy Zhe Wang
A big issue in modern LLM applications is they tend to feed long context to LLM, which results in high inference cost and latency, and may exceed the context limit. Prompt compress…