Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Shared circuits predict whether LLMs generalize across formats in arithmetic reasoning
Andrea Gregor de Varda, Sana Pandey, Pengrui Han +2
In many forms of reasoning, including arithmetic reasoning, generalizing across superficial changes in input format is effortless for humans: anyone who can solve 2+5 can also solv…
cs.CL2026
Last Translation Benchmark
Vilém Zouhar, Niyati Bafna, Mukund Choudhary +241
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, stan…
cs.CL2026
Investigating Knowledge Transfer Across Interactive Dialogue Games
Filippo Momentè, Mir Nafis Sharear Shopnil, Andrea de Varda +5
Dialogue games represent a challenging setting where complex cognitive skills are required to accomplish tasks while coordinating with other players. Considering that language repr…