2 papers
cs.CL2026
TurkBench: A Benchmark for Evaluating Turkish Large Language Models
ÃaÄrı Toraman, Ahmet Kaan Sever, Ayse Aysu Cengiz +19
With the recent surge in the development of large language models, the need for comprehensive and language-specific evaluation benchmarks has become critical. While significant pro…
cs.CL2025
Evaluating the Quality of Benchmark Datasets for Low-Resource Languages: A Case Study on Turkish
AyÅe Aysu Cengiz, Ahmet Kaan Sever, Elif Ecem Ãmütlü +6
The reliance on translated or adapted datasets from English or multilingual resources introduces challenges regarding linguistic and cultural suitability. This study addresses the…