1 paper
Shou-Tzu Han, Rodrigue Rizk, KC Santosh
Large language models demonstrate strong performance on mathematical reasoning benchmarks, yet remain surprisingly fragile to meaning-preserving surface perturbations. We systemati…