1 paper
Xidan Song, Weiqi Wang, Ruifeng Cao +1
The evaluation of Large Language Models (LLMs) in complex reasoning domains typically relies on performance alignment with ground-truth oracles. In the domain of chess, this standa…