3 papers
cs.CL2026
Tokens with Meaning: A Hybrid Tokenization Approach for Turkish
M. Ali Bayram, Ali Arda Fincan, Ahmet Semih GümüŠ+4
Tokenization shapes how language models perceive morphology and meaning in NLP, yet widely used frequency-driven subword tokenizers (e.g., Byte Pair Encoding and WordPiece) can fra…
cs.CL2025
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark
M. Ali Bayram, Ali Arda Fincan, Ahmet Semih GümüŠ+3
Tokenization is a fundamental preprocessing step in NLP, directly impacting large language models' (LLMs) ability to capture syntactic, morphosyntactic, and semantic structures. Th…
cs.CL2025
Setting Standards in Turkish NLP: TR-MMLU for Large Language Model Evaluation
M. Ali Bayram, Ali Arda Fincan, Ahmet Semih GümüŠ+3
Language models have made remarkable advancements in understanding and generating human language, achieving notable success across a wide array of applications. However, evaluating…