1 citations · 1 across the 1 of their papers we have counts for
1 paper · 1 filter
Youssef Benchekroun, Megi Dervishi, Mark Ibrahim +7
We propose WorldSense, a benchmark designed to assess the extent to which LLMs are consistently able to sustain tacit world models, by testing how they draw simple inferences from…