Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Agent^2 RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training?
Wanyi Chen, Xiao Yang, Xu Yang +7
We introduce Agent2 RL-Bench, a compact diagnostic benchmark for evaluating agentic RL post-training, which tests whether LLM agents can autonomously design, implement, debug, and…
cs.AI2024
Can LLMs plan paths in the real world?
Wanyi Chen, Meng-Wen Su, Nafisa Mehjabin +1
As large language models (LLMs) increasingly integrate into vehicle navigation systems, understanding their path-planning capability is crucial. We tested three LLMs through six re…