2 papers
cs.AI2026
Recovering Wasted Compute in Autoresearch Agents
Au Kwok Chun, Abhigyan Acherjee, Amrutha Rao +4
A slew of recent works develop agents for solving research problems end-to-end, a paradigm increasingly referred to as autoresearch. Such agents have inspired large industry invest…
cs.CL2026
End-to-End Context Compression at Scale
Ang Li, Sean McLeish, Haozhe Chen +12
Long-context language model inference is bottlenecked by memory, as the KV cache grows with context length. Recent techniques to compress the KV cache fall short: they either degra…