1 citations · 4 across the 19 of their papers we have counts for
4 papers · 1 filter
Multilingual Word-Level Forced Alignment with Self-Supervised Representations and Learned Dynamic Programming
Roy Weber, Meidan Zehavi, Rotem Rousso +1
We present a method for accurate multilingual word-level forced alignment, consisting of an alignment encoder and a learned alignment decoder. The encoder integrates two representa…
WhisperRT -- Turning Whisper into a Causal Streaming Model
Tomer Krichli, Bhiksha Raj, Joseph Keshet
Automatic Speech Recognition (ASR) has seen remarkable progress, with models like OpenAI Whisper and NVIDIA Canary achieving state-of-the-art (SOTA) performance in offline transcri…
WhisperNER: Unified Open Named Entity and Speech Recognition
Gil Ayache, Menachem Pirchi, Aviv Navon +3
Integrating named entity recognition (NER) with automatic speech recognition (ASR) can significantly enhance transcription accuracy and informativeness. In this paper, we introduce…
Combining Language Models For Specialized Domains: A Colorful Approach
Daniel Eitan, Menachem Pirchi, Neta Glazer +7
General purpose language models (LMs) encounter difficulties when processing domain-specific jargon and terminology, which are frequently utilized in specialized fields such as med…