1 citations · 2 across the 4 of their papers we have counts for
4 papers
Spoken question answering for visual queries
Nimrod Shabtay, Zvi Kons, Avihu Dekel +3
Question answering (QA) systems are designed to answer natural language questions. Visual QA (VQA) and Spoken QA (SQA) systems extend the textual QA system to accept visual and spo…
Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities
George Saon, Avihu Dekel, Alexander Brooks +21
Granite-speech LLMs are compact and efficient speech language models specifically designed for English ASR and automatic speech translation (AST). The models were trained by modali…
Exploring the Benefits of Tokenization of Discrete Acoustic Units
Avihu Dekel, Raul Fernandez
Tokenization algorithms that merge the units of a base vocabulary into larger, variable-rate units have become standard in natural language processing tasks. This idea, however, ha…
Speak While You Think: Streaming Speech Synthesis During Text Generation
Avihu Dekel, Slava Shechtman, Raul Fernandez +3
Large Language Models (LLMs) demonstrate impressive capabilities, yet interaction with these models is mostly facilitated through text. Using Text-To-Speech to synthesize LLM outpu…