3 papers
eess.AS2024
Whisper in Medusa's Ear: Multi-head Efficient Decoding for Transformer-based ASR
Yael Segal-Feldman, Aviv Shamsian, Aviv Navon +2
Large transformer-based models have significant potential for speech transcription and translation. Their self-attention mechanisms and parallel processing enable them to capture c…
eess.AS2021
CNN-based Spoken Term Detection and Localization without Dynamic Programming
Tzeviya Sylvia Fuchs, Yael Segal, Joseph Keshet
In this paper, we propose a spoken term detection algorithm for simultaneous prediction and localization of in-vocabulary and out-of-vocabulary terms within an audio segment. The p…
eess.AS2019
SpeechYOLO: Detection and Localization of Speech Objects
Yael Segal, Tzeviya Sylvia Fuchs, Joseph Keshet
In this paper, we propose to apply object detection methods from the vision domain on the speech recognition domain, by treating audio fragments as objects. More specifically, we p…