1 paper · 1 filter
Qiming Bao, Xiaoxuan Fu, Michael Witbrock
Large language models (LLMs) achieve high accuracy on many reasoning benchmarks but remain brittle under structural perturbations of rule-based systems. We introduce a diagnostic f…