7 citations · 10 across the 3 of their papers we have counts for
5 papers · 1 filter
Bloom Library: Multimodal Datasets in 300+ Languages for a Variety of Downstream Tasks
Colin Leong, Joshua Nemecek, Jacob Mansdorfer +3
We present Bloom Library, a linguistically diverse set of multimodal and multilingual datasets for language modeling, image captioning, visual storytelling, and speech synthesis/re…
Dyn-ASR: Compact, Multilingual Speech Recognition via Spoken Language and Accent Identification
Sangeeta Ghangam, Daniel Whitenack, Joshua Nemecek
Running automatic speech recognition (ASR) on edge devices is non-trivial due to resource constraints, especially in scenarios that require supporting multiple languages. We propos…
Participatory Research for Low-resourced Machine Translation: A Case Study in African Languages
Wilhelmina Nekoto, Vukosi Marivate, Tshinondiwa Matsila +45
Research in NLP lacks geographic diversity, and the question of how NLP can be scaled to low-resourced languages has not yet been adequately solved. "Low-resourced"-ness is a compl…
Masakhane -- Machine Translation For Africa
Iroro Orife, Julia Kreutzer, Blessing Sibanda +22
Africa has over 2000 languages. Despite this, African languages account for a small portion of available resources and publications in Natural Language Processing (NLP). This is du…
Katecheo: A Portable and Modular System for Multi-Topic Question Answering
Shirish Hirekodi, Seban Sunny, Leonard Topno +4
We introduce a modular system that can be deployed on any Kubernetes cluster for question answering via REST API. This system, called Katecheo, includes three configurable modules…