9 papers
Speaker Attributed Automatic Speech Recognition Using Speech Aware LLMS
Hagai Aronowitz, Zvi Kons, Avihu Dekel +2
Speaker-Attributed Automatic Speech Recognition (SAA) enhances traditional ASR systems by incorporating relative speaker identity tags directly into the transcript (e.g., [Speaker…
Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark
Arnon Turetzky, Avihu Dekel, Hagai Aronowitz +2
Spoken meaning often depends not only on what is said, but also on which word is emphasized. The same sentence can convey correction, contrast, or clarification depending on where…
Balanced Thinking: Improving Chain of Thought Training in Vision Language Models
Shaked Perek, Ben Wiesel, Avihu Dekel +2
Multimodal reasoning in vision-language models (VLMs) typically relies on a two-stage process: supervised fine-tuning (SFT) and reinforcement learning (RL). In standard SFT, all to…
Self-Speculative Decoding for LLM-based ASR with CTC Encoder Drafts
George Saon, Samuel Thomas, Takashi Fukuda +3
We propose self-speculative decoding for speech-aware LLMs by using the CTC encoder as a draft model to accelerate auto-regressive (AR) inference and improve ASR accuracy. Our thre…
NLE: Non-autoregressive LLM-based ASR by Transcript Editing
Avihu Dekel, Samuel Thomas, Takashi Fukada +1
While autoregressive (AR) LLM-based ASR systems achieve strong accuracy, their sequential decoding limits parallelism and incurs high latency. We propose NLE, a non-autoregressive…
Speech Synthesis From Continuous Features Using Per-Token Latent Diffusion
Arnon Turetzky, Avihu Dekel, Nimrod Shabtay +5
We present SALAD, a zero-shot TTS autoregressive model operating over continuous speech representations. SALAD utilizes a per-token diffusion process to refine and predict continuo…