1 paper
Abinav Rao, Sujan Rachuri, Nikhil Vemuri
LLMs can execute every step of chain-of-thought reasoning correctly and still produce wrong final answers. We introduce the Novel Operator Test, a benchmark that separates operator…