Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Benchmarking at the Edge of Comprehension
Samuele Marro, Jialin Yu, Emanuele La Malfa +8
As frontier Large Language Models (LLMs) increasingly saturate new benchmarks shortly after they are published, benchmarking itself is at a juncture: if frontier models keep improv…
cs.AI2026
BiasBusters: Uncovering and Mitigating Tool Selection Bias in Large Language Models
Thierry Blankenstein, Jialin Yu, Zixuan Li +6
Agents backed by large language models (LLMs) increasingly rely on external tools drawn from marketplaces where multiple providers offer functionally equivalent options. This raise…