collaborators

7 papers

cs.CL2026

Cross-Lingual Transfer for Machine Translation in Turkic Languages

Omer Burak Cinar, Mehmet Mert Dalkilic, Cagri Toraman

Cross-lingual transfer is central to low-resource machine translation, but its behavior within closely related language families remains insufficiently characterized. We study tran…

cs.CL2026

OCRTurk: A Comprehensive OCR Benchmark for Turkish

Deniz Yılmaz, Evren Ayberk Munis, Çağrı Toraman +4

Document parsing is now widely used in applications, such as large-scale document digitization, retrieval-augmented generation, and domain-specific pipelines in healthcare and educ…

cs.CL2026

RAGTurk: Best Practices for Retrieval Augmented Generation in Turkish

Süha Kağan Köse, Mehmet Can Baytekin, Burak Aktaş +5

Retrieval-Augmented Generation (RAG) enhances LLM factuality, yet design guidance remains English-centric, limiting insights for morphologically rich languages like Turkish. We add…

cs.CL2026

BIRDTurk: Adaptation of the BIRD Text-to-SQL Dataset to Turkish

Burak Aktaş, Mehmet Can Baytekin, Süha Kağan Köse +4

Text-to-SQL systems have achieved strong performance on English benchmarks, yet their behavior in morphologically rich, low-resource languages remains largely unexplored. We introd…

cs.CL2026

TurkBench: A Benchmark for Evaluating Turkish Large Language Models

Çağrı Toraman, Ahmet Kaan Sever, Ayse Aysu Cengiz +19

With the recent surge in the development of large language models, the need for comprehensive and language-specific evaluation benchmarks has become critical. While significant pro…

cs.CL2026

OpenEthics: A Comprehensive Ethical Evaluation of Open-Source Generative Large Language Models

Yıldırım Özen, Burak Erinç Çetin, Kaan Engür +2

Generative large language models present significant potential but also raise critical ethical concerns, including issues of safety, fairness, robustness, and reliability. Most exi…