Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Learning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimes
Justin Bauer, Thomas Walshe, Derek Pham +4
Fine-tuning Large Language Models (LLMs) typically relies on large quantities of high-quality annotated data, or questions with well-defined ground truth answers in the case of Rei…
cs.AI2026
RIFT: A RubrIc Failure Mode Taxonomy and Automated Diagnostics
Zhengyang Qi, Charles Dickens, Derek Pham +4
Rubric-based evaluation is widely used in LLM benchmarks and training pipelines for open-ended, less verifiable tasks. While prior work has demonstrated the effectiveness of rubric…