1 paper
Zenghui Zhou, Man Li, Xiaoke Fang +3
Large Language Models (LLMs) achieve strong performance on logical reasoning benchmarks, yet their reliability remains uncertain. Existing evaluations rely on static benchmarks, wh…