2 papers
cs.AI2026
PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling
Dongjie Xu, Julius, Hanchi Dong +6
Reliable evaluation of tool routing is critical as Large Language Models increasingly operate as autonomous agents. Current benchmarks face three structural limitations: data distr…
cs.LG2026
RippleKV: Cross-Layer KV Cache Allocation via Perturbation Propagation
Dongjie Xu, Kai Qian, Julius +6
Long-context LLM inference is bottlenecked by KV cache memory, yet distributing a limited cache budget across layers remains challenging. Existing methods rely on proxies such as l…