33 citations · 137 across the 39 of their papers we have counts for
70 papers · 1 filter
IDIOLEX: Unified and Continuous Representations for Idiolectal and Stylistic Variation
Anjali Kantharuban, Aarohi Srivastava, Fahim Faisal +5
Existing sentence representations primarily encode what a sentence says, rather than how it is expressed, even though the latter is important for many applications. In contrast, we…
GMU Systems for the IWSLT 2025 Low-Resource Speech Translation Shared Task
Chutong Meng, Antonios Anastasopoulos
This paper describes the GMU systems for the IWSLT 2025 low-resource speech translation shared task. We trained systems for all language pairs, except for Levantine Arabic. We fine…
CMULAB: An Open-Source Framework for Training and Deployment of Natural Language Processing Models
Zaid Sheikh, Antonios Anastasopoulos, Shruti Rijhwani +3
Effectively using Natural Language Processing (NLP) tools in under-resourced languages requires a thorough understanding of the language itself, familiarity with the latest models…
Language and Speech Technology for Central Kurdish Varieties
Sina Ahmadi, Daban Q. Jaff, Md Mahfuz Ibn Alam +1
Kurdish, an Indo-European language spoken by over 30 million speakers, is considered a dialect continuum and known for its diversity in language varieties. Previous studies address…
A Case Study on Filtering for End-to-End Speech Translation
Md Mahfuz Ibn Alam, Antonios Anastasopoulos
It is relatively easy to mine a large parallel corpus for any machine learning task, such as speech-to-text or speech-to-speech translation. Although these mined corpora are large…
A Morphologically-Aware Dictionary-based Data Augmentation Technique for Machine Translation of Under-Represented Languages
Md Mahfuz Ibn Alam, Sina Ahmadi, Antonios Anastasopoulos
The availability of parallel texts is crucial to the performance of machine translation models. However, most of the world's languages face the predominant challenge of data scarci…