activity
20212025
collaborators

6 papers

cs.CL2025

AfriMTEB and AfriE5: Benchmarking and Adapting Text Embedding Models for African Languages

Kosei Uemura, Miaoran Zhang, David Ifeoluwa Adelani

Text embeddings are an essential building component of several NLP tasks such as retrieval-augmented generation which is crucial for preventing hallucinations in LLMs. Despite the…

cs.CL2025

Unveiling the Key Factors for Distilling Chain-of-Thought Reasoning

Xinghao Chen, Zhijing Sun, Wenjin Guo +8

Large Language Models (LLMs) excel in reasoning tasks through Chain-of-Thought (CoT) prompting. However, CoT prompting greatly increases computational demands, which has prompted g…

cs.CL2025

AFRIDOC-MT: Document-level MT Corpus for African Languages

Jesujoba O. Alabi, Israel Abebe Azime, Miaoran Zhang +13

This paper introduces AFRIDOC-MT, a document-level multi-parallel translation dataset covering English and five African languages: Amharic, Hausa, Swahili, Yorùbá, and Zulu. The da…

cs.CL2022

MCSE: Multimodal Contrastive Learning of Sentence Embeddings

Miaoran Zhang, Marius Mosbach, David Ifeoluwa Adelani +2

Learning semantically meaningful sentence embeddings is an open problem in natural language processing. In this work, we propose a sentence embedding learning approach that exploit…

cs.IR2022

Knowledge Base Index Compression via Dimensionality and Precision Reduction

Vilém Zouhar, Marius Mosbach, Miaoran Zhang +1

Recently neural network based approaches to knowledge-intensive NLP tasks, such as question answering, started to rely heavily on the combination of neural retrievers and readers.…

cs.CL2021

Preventing Author Profiling through Zero-Shot Multilingual Back-Translation

David Ifeoluwa Adelani, Miaoran Zhang, Xiaoyu Shen +3

Documents as short as a single sentence may inadvertently reveal sensitive information about their authors, including e.g. their gender or ethnicity. Style transfer is an effective…