5 papers
Tokenization Disparities as Infrastructure Bias: How Subword Systems Create Inequities in LLM Access and Efficiency
Hailay Kidu Teklehaymanot, Wolfgang Nejdl
Tokenization disparities pose a significant barrier to achieving equitable access to artificial intelligence across linguistically diverse populations. This study conducts a large-…
Morphological Synthesizer for Ge'ez Language: Addressing Morphological Complexity and Resource Limitations
Gebrearegawi Gebremariam, Hailay Teklehaymanot, Gebregewergs Mezgebe
Ge'ez is an ancient Semitic language renowned for its unique alphabet. It serves as the script for numerous languages, including Tigrinya and Amharic, and played a pivotal role in…
Low-Resource English-Tigrinya MT: Leveraging Multilingual Models, Custom Tokenizers, and Clean Evaluation Benchmarks
Hailay Kidu Teklehaymanot, Gebrearegawi Gidey, Wolfgang Nejdl
Despite advances in Neural Machine Translation (NMT), low-resource languages like Tigrinya remain underserved due to persistent challenges, including limited corpora, inadequate to…
MoVoC: Morphology-Aware Subword Construction for Geez Script Languages
Hailay Kidu Teklehaymanot, Dren Fazlija, Wolfgang Nejdl
Subword-based tokenization methods often fail to preserve morphological boundaries, a limitation especially pronounced in low-resource, morphologically complex languages such as th…
RAGtifier: Evaluating RAG Generation Approaches of State-of-the-Art RAG Systems for the SIGIR LiveRAG Competition
Tim Cofala, Oleh Astappiev, William Xion +1
Retrieval-Augmented Generation (RAG) enriches Large Language Models (LLMs) by combining their internal, parametric knowledge with external, non-parametric sources, with the goal of…