1 paper · 1 filter
Abinav Rao, Sujan Rachuri, Nikhil Vemuri
LLMs can execute every step of chain-of-thought reasoning correctly and still produce wrong final answers. We introduce the Novel Operator Test, a benchmark that separates operator…