Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Shepherd: Enabling Programmable Meta-Agents via Reversible Agentic Execution Traces
Simon Yu, Derek Chong, Ananjan Nandi +4
As LLM agent systems take on more complex tasks, they increasingly rely on meta-agents: higher-order agents that create, operate on and manage other agents. Meta-agent operations s…
cs.AI2025
SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning
Bo Liu, Leon Guertler, Simon Yu +9
Recent advances in reinforcement learning have shown that language models can develop sophisticated reasoning through training on tasks with verifiable rewards, but these approache…