1 paper
Ali Ansari, Haoran Sun, Andy Zeyi Liu +48
Low reported scores on leading physics benchmarks, including those featured in the Artificial Analysis Intelligence Index (2026), suggest that frontier language models still strugg…