From the 1 of 1 linked paper with an AI index.
1 paper · 1 filter
Nayera Hasan, Jack Greff, Alvin Grissom
The paper presents a benchmark for testing large language models' ability to reason logically about probability expressions in English, and evaluates 29 models, revealing widesprea…