2 papers
cs.SE2026
SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks
Pavel Adamenko, Mikhail Ivanov, Aidar Valeev +6
The rapid advancement of Large Language Models (LLMs) in software engineering has revealed critical limitations in existing benchmarks, particularly the widely used SWE-bench datas…
astro-ph.CO2025
The Spectroscopic Stage-5 Experiment
Robert Besuner, Arjun Dey, Alex Drlica-Wagner +106
The existence, properties, and dynamics of the dark sectors of our universe pose fundamental challenges to our current model of physics, and large-scale astronomical surveys may be…