1 paper · 1 filter
Luca Avena, Gianmarco Bet, Bernardo Busoni
We investigate the probabilistic reasoning capabilities of large language models through a controlled benchmarking study on discrete probability problems. We constructed two datase…