most citedTraining and Inference Efficiency of Encoder-Decoder Speech Models

1 citations · 1 across the 3 of their papers we have counts for

collaborators

7 papers

cs.CL2025

Word Level Timestamp Generation for Automatic Speech Recognition and Translation

Ke Hu, Krishna Puvvada, Elena Rastorgueva +7

We introduce a data-driven approach for enabling word-level timestamp prediction in the Canary model. Accurate timestamp information is crucial for a variety of downstream tasks su…

cs.CL2025

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model

Ke Hu, Ehsan Hosseini-Asl, Chen Chen +7

Spoken dialogue is an intuitive form of human-computer interaction, yet current speech language models often remain constrained to turn-based exchanges, lacking real-time adaptabil…

cs.CL20251 cited

Training and Inference Efficiency of Encoder-Decoder Speech Models

Piotr Żelasko, Kunal Dhawan, Daniel Galvez +7

Attention encoder-decoder model architecture is the backbone of several recent top performing foundation speech models: Whisper, Seamless, OWSM, and Canary-1B. However, the reporte…

cs.CL2024

NeKo: Cross-Modality Post-Recognition Error Correction with Tasks-Guided Mixture-of-Experts Language Model

Yen-Ting Lin, Zhehuai Chen, Piotr Zelasko +11

Construction of a general-purpose post-recognition error corrector poses a crucial question: how can we most effectively train a model on a large mixture of domain datasets? The an…

cs.CL2024

VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning

Yifan Peng, Krishna C. Puvvada, Zhehuai Chen +7

Recent studies have augmented large language models (LLMs) with speech capabilities, leading to the development of speech language models (SpeechLMs). Earlier SpeechLMs focused on…

cs.CL2024

EMMeTT: Efficient Multimodal Machine Translation Training

Piotr Żelasko, Zhehuai Chen, Mengru Wang +7

A rising interest in the modality extension of foundation language models warrants discussion on the most effective, and efficient, multimodal training approach. This work focuses…