2 papers
cs.CL2026
From Rosetta to Match-Up: A Paired Corpus of Linguistic Puzzles with Human and LLM Benchmarks
Neh Majmudar, Anne Huang, Jinfan Frank Hu +1
In this paper, we examine linguistic puzzles used in high school linguistics competitions, focusing on two common formats: Rosetta Stone and Match-Up. We propose a systematic proce…
cs.SE2025
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
Titouan Duston, Shuo Xin, Yang Sun +26
We introduce AInsteinBench, a large-scale benchmark for evaluating whether large language model (LLM) agents can operate as scientific computing development agents within real rese…