5 citations · 6 across the 7 of their papers we have counts for
6 papers · 1 filter
Word Timestamps and Speaker Attribution with a Non-Autoregressive LLM
Zvi Kons, Avihu Dekel, Hagai Aronowitz +2
Timestamps and speaker attribution are useful additions to speech recognition, creating a rich text transcript. This information can either be extracted during transcription or ali…
Speaker Attributed Automatic Speech Recognition Using Speech Aware LLMS
Hagai Aronowitz, Zvi Kons, Avihu Dekel +2
Speaker-Attributed Automatic Speech Recognition (SAA) enhances traditional ASR systems by incorporating relative speaker identity tags directly into the transcript (e.g., [Speaker…
Spoken question answering for visual queries
Nimrod Shabtay, Zvi Kons, Avihu Dekel +3
Question answering (QA) systems are designed to answer natural language questions. Visual QA (VQA) and Spoken QA (SQA) systems extend the textual QA system to accept visual and spo…
Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities
George Saon, Avihu Dekel, Alexander Brooks +21
Granite-speech LLMs are compact and efficient speech language models specifically designed for English ASR and automatic speech translation (AST). The models were trained by modali…
Speech Synthesis From Continuous Features Using Per-Token Latent Diffusion
Arnon Turetzky, Avihu Dekel, Nimrod Shabtay +5
We present SALAD, a zero-shot TTS autoregressive model operating over continuous speech representations. SALAD utilizes a per-token diffusion process to refine and predict continuo…
Siamese x-vector reconstruction for domain adapted speaker recognition
Shai Rozenberg, Hagai Aronowitz, Ron Hoory
With the rise of voice-activated applications, the need for speaker recognition is rapidly increasing. The x-vector, an embedding approach based on a deep neural network (DNN), is…