6 papers
AfriMTEB and AfriE5: Benchmarking and Adapting Text Embedding Models for African Languages
Kosei Uemura, Miaoran Zhang, David Ifeoluwa Adelani
Text embeddings are an essential building component of several NLP tasks such as retrieval-augmented generation which is crucial for preventing hallucinations in LLMs. Despite the…
Unveiling the Key Factors for Distilling Chain-of-Thought Reasoning
Xinghao Chen, Zhijing Sun, Wenjin Guo +8
Large Language Models (LLMs) excel in reasoning tasks through Chain-of-Thought (CoT) prompting. However, CoT prompting greatly increases computational demands, which has prompted g…
AFRIDOC-MT: Document-level MT Corpus for African Languages
Jesujoba O. Alabi, Israel Abebe Azime, Miaoran Zhang +13
This paper introduces AFRIDOC-MT, a document-level multi-parallel translation dataset covering English and five African languages: Amharic, Hausa, Swahili, Yorùbá, and Zulu. The da…
MCSE: Multimodal Contrastive Learning of Sentence Embeddings
Miaoran Zhang, Marius Mosbach, David Ifeoluwa Adelani +2
Learning semantically meaningful sentence embeddings is an open problem in natural language processing. In this work, we propose a sentence embedding learning approach that exploit…
Knowledge Base Index Compression via Dimensionality and Precision Reduction
Vilém Zouhar, Marius Mosbach, Miaoran Zhang +1
Recently neural network based approaches to knowledge-intensive NLP tasks, such as question answering, started to rely heavily on the combination of neural retrievers and readers.…
Preventing Author Profiling through Zero-Shot Multilingual Back-Translation
David Ifeoluwa Adelani, Miaoran Zhang, Xiaoyu Shen +3
Documents as short as a single sentence may inadvertently reveal sensitive information about their authors, including e.g. their gender or ethnicity. Style transfer is an effective…