Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
DNR Bench: Benchmarking Over-Reasoning in Reasoning LLMs
Masoud Hashemi, Oluwanifemi Bamgbose, Sathwik Tejaswi Madhusudhan +3
Test-time scaling has significantly improved large language model performance, enabling deeper reasoning to solve complex problems. However, this increased reasoning capability als…
cs.LG2025
Revitalizing Saturated Benchmarks: A Weighted Metric Approach for Differentiating Large Language Model Performance
Bryan Etzine, Masoud Hashemi, Nishanth Madhusudhan +4
Existing benchmarks are becoming saturated and struggle to separate model performances due to factors like data contamination and advancing LLM capabilities. This paper introduces…