5 citations · 5 across the 1 of their papers we have counts for
2 papers
cs.CL2026★ 5 cited
Lessons from the Trenches on Reproducible Evaluation of Language Models
Stella Biderman, Hailey Schoelkopf, Lintang Sutawika +27
Reliable evaluation of language models (LMs) remains an open challenge. Re- searchers and engineers face methodological issues such as the sensitivity of models to evaluation setup…
cs.CV2025
Deep Neural Networks Can Learn Generalizable Same-Different Visual Relations
Alexa R. Tartaglini, Sheridan Feucht, Michael A. Lepori +4
Although deep neural networks can achieve human-level performance on many object recognition benchmarks, prior work suggests that these same models fail to learn simple abstract re…