Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across Domains
Xuzhao Li, Xuchen Li, Shiyu Hu +2
Large language models (LLMs) increasingly rely on reinforcement learning (RL) to enhance their reasoning capabilities through feedback. A critical challenge is verifying the consis…
cs.AI2025
Large Language Models for Planning: A Comprehensive and Systematic Survey
Pengfei Cao, Tianyi Men, Wencan Liu +7
Planning represents a fundamental capability of intelligent agents, requiring comprehensive environmental understanding, rigorous logical reasoning, and effective sequential decisi…