3 papers
eess.AS2023
GPU-Accelerated WFST Beam Search Decoder for CTC-based Speech Recognition
Daniel Galvez, Tim Kaldewey
While Connectionist Temporal Classification (CTC) models deliver state-of-the-art accuracy in automated speech recognition (ASR) pipelines, their performance has been limited by CP…
cs.AI2023
Speech Wikimedia: A 77 Language Multilingual Speech Dataset
Rafael Mosquera Gómez, Julián Eusse, Juan Ciro +4
The Speech Wikimedia Dataset is a publicly available compilation of audio with transcriptions extracted from Wikimedia Commons. It includes 1780 hours (195 GB) of CC-BY-SA licensed…
cs.CL2021
LSH methods for data deduplication in a Wikipedia artificial dataset
Juan Ciro, Daniel Galvez, Tim Schlippe +1
This paper illustrates locality sensitive hasing (LSH) models for the identification and removal of nearly redundant data in a text dataset. To evaluate the different models, we cr…