1 paper · 1 filter
Avni Mittal, Rauno Arike
Large language models (LLMs) are increasingly used as judges of chain-of-thought (CoT) reasoning, yet it remains unclear whether they can reliably assess process faithfulness rathe…