1 paper · 1 filter
Jiahe Geng, Jinpeng Wang, Kun Yuan
Many long-horizon LLM deployments face tight prompt budgets: latency, cost, and context limits make full-context prompting impractical as interaction length grows. The key question…