activity
20242026
most citedGranite-speech: open-source speech-aware LLMs with strong English ASR capabilities

1 citations · 1 across the 7 of their papers we have counts for

collaborators
Showing eess.ASShow all

7 papers · 1 filter

eess.AS2026

Speaker Attributed Automatic Speech Recognition Using Speech Aware LLMS

Hagai Aronowitz, Zvi Kons, Avihu Dekel +2

Speaker-Attributed Automatic Speech Recognition (SAA) enhances traditional ASR systems by incorporating relative speaker identity tags directly into the transcript (e.g., [Speaker…

eess.AS2026

Self-Speculative Decoding for LLM-based ASR with CTC Encoder Drafts

George Saon, Samuel Thomas, Takashi Fukuda +3

We propose self-speculative decoding for speech-aware LLMs by using the CTC encoder as a draft model to accelerate auto-regressive (AR) inference and improve ASR accuracy. Our thre…

eess.AS2026

NLE: Non-autoregressive LLM-based ASR by Transcript Editing

Avihu Dekel, Samuel Thomas, Takashi Fukada +1

While autoregressive (AR) LLM-based ASR systems achieve strong accuracy, their sequential decoding limits parallelism and incurs high latency. We propose NLE, a non-autoregressive…

eess.AS2025

Spoken question answering for visual queries

Nimrod Shabtay, Zvi Kons, Avihu Dekel +3

Question answering (QA) systems are designed to answer natural language questions. Visual QA (VQA) and Spoken QA (SQA) systems extend the textual QA system to accept visual and spo…

eess.AS20251 cited

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities

George Saon, Avihu Dekel, Alexander Brooks +21

Granite-speech LLMs are compact and efficient speech language models specifically designed for English ASR and automatic speech translation (AST). The models were trained by modali…

eess.AS2024

Low Bitrate High-Quality RVQGAN-based Discrete Speech Tokenizer

Slava Shechtman, Avihu Dekel

Discrete Audio codecs (or audio tokenizers) have recently regained interest due to the ability of Large Language Models (LLMs) to learn their compressed acoustic representations. V…