3 citations · 5 across the 7 of their papers we have counts for
6 papers · 1 filter
Last Translation Benchmark
Vilém Zouhar, Niyati Bafna, Mukund Choudhary +241
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, stan…
English to Central Kurdish Speech Translation: Corpus Creation, Evaluation, and Orthographic Standardization
Mohammad Mohammadamini, Daban Q. Jaff, Josep Crego +2
We present KUTED, a speech-to-text translation (S2TT) dataset for Central Kurdish, derived from TED and TEDx talks. The corpus comprises 91,000 sentence pairs, including 170 hours…
FLEURS-Kobani: Extending the FLEURS Dataset for Northern Kurdish
Daban Q. Jaff, Mohammad Mohammadamini
FLEURS offers n-way parallel speech for 100+ languages, but Northern Kurdish is not one of them, which limits benchmarking for automatic speech recognition and speech translation t…
From Consensus to Split Decisions: ABC-Stratified Sentiment in Holocaust Oral Histories
Daban Q. Jaff
Polarity detection becomes substantially more challenging under domain shift, particularly in heterogeneous, long-form narratives with complex discourse structure, such as Holocaus…
Language and Speech Technology for Central Kurdish Varieties
Sina Ahmadi, Daban Q. Jaff, Md Mahfuz Ibn Alam +1
Kurdish, an Indo-European language spoken by over 30 million speakers, is considered a dialect continuum and known for its diversity in language varieties. Previous studies address…
Leveraging Multilingual News Websites for Building a Kurdish Parallel Corpus
Sina Ahmadi, Hossein Hassani, Daban Q. Jaff
Machine translation has been a major motivation of development in natural language processing. Despite the burgeoning achievements in creating more efficient machine translation sy…