2 papers
cs.AI2026
CUBE: A Standard for Unifying Agent Benchmarks
Alexandre Lacoste, Nicolas Gontier, Oleh Shliazhko +23
The proliferation of agent benchmarks has created critical fragmentation that threatens research productivity. Each new benchmark requires substantial custom integration, creating…
cs.AI2026
JEF-Hinter: Leveraging Offline Knowledge for Improving Web Agents Adaptation
Hadi Nekoei, Aman Jaiswal, Patrice Bechard +7
Large language model (LLM) agents perform well in sequential decision-making tasks, but improving them on unfamiliar domains often requires costly online interactions or fine-tunin…