35 citations · 35 across the 2 of their papers we have counts for
1 paper · 1 filter
Jeremy Schlatter, Benjamin Weinstein-Raun, Jeffrey Ladish
In experiments spanning more than 100,000 trials across thirteen large language models, we show that several state-of-the-art models presented with a simple task (including Grok 4,…