2 citations · 2 across the 1 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents
Davide Paglieri, BartÅomiej CupiaÅ, Jonathan Cook +6
Training large language models (LLMs) to reason via reinforcement learning (RL) significantly improves their problem-solving capabilities. In agentic settings, existing methods lik…
cs.AI2025
Ad-Hoc Human-AI Coordination Challenge
Tin DizdareviÄ, Ravi Hammond, Tobias Gessler +7
Achieving seamless coordination between AI agents and humans is crucial for real-world applications, yet it remains a significant open challenge. Hanabi is a cooperative card game…