565 citations · 708 across the 34 of their papers we have counts for
Showing 2026 · cs.CLShow all
3 papers · 2 filters
cs.CL2026
Last Translation Benchmark
Vilém Zouhar, Niyati Bafna, Mukund Choudhary +241
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, stan…
cs.CL2026
How Fine-Grained Should a RAG Benchmark Be? A Hierarchical Framework for Synthetic Question Generation
Chase M. Fensore, Kaustubh Dhole, Jason Fan +2
Evaluating retrieval-augmented generation (RAG) systems requires benchmarks that capture diverse question characteristics, yet practitioners lack empirical guidance on which dimens…
cs.CL2026
BabyReasoningBench: Generating Developmentally-Inspired Reasoning Tasks for Evaluating Baby Language Models
Kaustubh D. Dhole
Traditional evaluations of reasoning capabilities of language models are dominated by adult-centric benchmarks that presuppose broad world knowledge, complex instruction following,…