2 citations · 3 across the 2 of their papers we have counts for
2 papers
cs.CL2022★ 2 cited
Bloom Library: Multimodal Datasets in 300+ Languages for a Variety of Downstream Tasks
Colin Leong, Joshua Nemecek, Jacob Mansdorfer +3
We present Bloom Library, a linguistically diverse set of multimodal and multilingual datasets for language modeling, image captioning, visual storytelling, and speech synthesis/re…
cs.CL2021★ 1 cited
Dyn-ASR: Compact, Multilingual Speech Recognition via Spoken Language and Accent Identification
Sangeeta Ghangam, Daniel Whitenack, Joshua Nemecek
Running automatic speech recognition (ASR) on edge devices is non-trivial due to resource constraints, especially in scenarios that require supporting multiple languages. We propos…