collaborators

5 papers

cs.CL2025

Tokenization Disparities as Infrastructure Bias: How Subword Systems Create Inequities in LLM Access and Efficiency

Hailay Kidu Teklehaymanot, Wolfgang Nejdl

Tokenization disparities pose a significant barrier to achieving equitable access to artificial intelligence across linguistically diverse populations. This study conducts a large-…

cs.CL2025

Morphological Synthesizer for Ge'ez Language: Addressing Morphological Complexity and Resource Limitations

Gebrearegawi Gebremariam, Hailay Teklehaymanot, Gebregewergs Mezgebe

Ge'ez is an ancient Semitic language renowned for its unique alphabet. It serves as the script for numerous languages, including Tigrinya and Amharic, and played a pivotal role in…

cs.CL2025

Low-Resource English-Tigrinya MT: Leveraging Multilingual Models, Custom Tokenizers, and Clean Evaluation Benchmarks

Hailay Kidu Teklehaymanot, Gebrearegawi Gidey, Wolfgang Nejdl

Despite advances in Neural Machine Translation (NMT), low-resource languages like Tigrinya remain underserved due to persistent challenges, including limited corpora, inadequate to…

cs.CL2025

MoVoC: Morphology-Aware Subword Construction for Geez Script Languages

Hailay Kidu Teklehaymanot, Dren Fazlija, Wolfgang Nejdl

Subword-based tokenization methods often fail to preserve morphological boundaries, a limitation especially pronounced in low-resource, morphologically complex languages such as th…

cs.IR2025

RAGtifier: Evaluating RAG Generation Approaches of State-of-the-Art RAG Systems for the SIGIR LiveRAG Competition

Tim Cofala, Oleh Astappiev, William Xion +1

Retrieval-Augmented Generation (RAG) enriches Large Language Models (LLMs) by combining their internal, parametric knowledge with external, non-parametric sources, with the goal of…