2 papers
cs.CL2025
Büyük Dil Modelleri için TR-MMLU Benchmarkı: Performans Değerlendirmesi, Zorluklar ve İyileştirme Fırsatları
M. Ali Bayram, Ali Arda Fincan, Ahmet Semih Gümüş +3
Language models have made significant advancements in understanding and generating human language, achieving remarkable success in various applications. However, evaluating these m…
cs.CL2025
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark
M. Ali Bayram, Ali Arda Fincan, Ahmet Semih Gümüş +3
Tokenization is a fundamental preprocessing step in NLP, directly impacting large language models' (LLMs) ability to capture syntactic, morphosyntactic, and semantic structures. Th…