17 citations · 19 across the 4 of their papers we have counts for
4 papers · 1 filter
Speech ReaLLM -- Real-time Streaming Speech Recognition with Multimodal LLMs by Teaching the Flow of Time
Frank Seide, Morrie Doulaty, Yangyang Shi +3
We introduce Speech ReaLLM, a new ASR architecture that marries "decoder-only" ASR with the RNN-T to make multimodal LLM architectures capable of real-time streaming. This is the f…
Leveraging Timestamp Information for Serialized Joint Streaming Recognition and Translation
Sara Papi, Peidong Wang, Junkun Chen +4
The growing need for instant spoken language transcription and translation is driven by increased global communication and cross-lingual interactions. This has made offering transl…
VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation
Tianrui Wang, Long Zhou, Ziqiang Zhang +6
Recent research shows a big convergence in model architecture, training objectives, and inference methods across various tasks for different modalities. In this paper, we propose V…
Sequence-level self-learning with multiple hypotheses
Kenichi Kumatani, Dimitrios Dimitriadis, Yashesh Gaur +4
In this work, we develop new self-learning techniques with an attention-based sequence-to-sequence (seq2seq) model for automatic speech recognition (ASR). For untranscribed speech…