2 papers
cs.AI2026
Learning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimes
Justin Bauer, Thomas Walshe, Derek Pham +4
Fine-tuning Large Language Models (LLMs) typically relies on large quantities of high-quality annotated data, or questions with well-defined ground truth answers in the case of Rei…
cs.SE2025
Automating Benchmark Design
Amanda Dsouza, Harit Vishwakarma, Zhengyang Qi +6
The rapid progress and widespread deployment of LLMs and LLM-powered agents has outpaced our ability to evaluate them. Hand-crafted, static benchmarks are the primary tool for asse…