Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling
Dongjie Xu, Julius, Hanchi Dong +6
Reliable evaluation of tool routing is critical as Large Language Models increasingly operate as autonomous agents. Current benchmarks face three structural limitations: data distr…
cs.AI2026
LoopGuard: Breaking Self-Reinforcing Attention Loops via Dynamic KV Cache Intervention
Dongjie Xu, Hao Wu, Weijie Shi +7
Through systematic experiments on long-context generation, we observe a damaging failure mode in which decoding can collapse into persistent repetition loops. We find that this deg…