1 paper
Yanan Cai, Ahmed Salem, Besmira Nushi +1
We introduce LogiPlan, a novel benchmark designed to evaluate the capabilities of large language models (LLMs) in logical planning and reasoning over complex relational structures.…