27 citations · 31 across the 6 of their papers we have counts for
6 papers
Gemma 4 Technical Report
Gemma Team, Sherif El Abd, Vaibhav Aggarwal +320
We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemm…
Cross-attention conformer for context modeling in speech enhancement for ASR
Arun Narayanan, Chung-Cheng Chiu, Tom O'Malley +2
This work introduces \emph{cross-attention conformer}, an attention-based architecture for context modeling in speech enhancement. Given that the context information can often be s…
Tied & Reduced RNN-T Decoder
Rami Botros, Tara N. Sainath, Robert David +3
Previous works on the Recurrent Neural Network-Transducer (RNN-T) models have shown that, under some conditions, it is possible to simplify its prediction network with little or no…
Multi-user VoiceFilter-Lite via Attentive Speaker Embedding
Rajeev Rikhye, Quan Wang, Qiao Liang +2
In this paper, we propose a solution to allow speaker conditioned speech models, such as VoiceFilter-Lite, to support an arbitrary number of enrolled users in a single pass. This i…
Multi-Task Learning for End-to-End ASR Word and Utterance Confidence with Deletion Prediction
David Qiu, Yanzhang He, Qiujia Li +3
Confidence scores are very useful for downstream applications of automatic speech recognition (ASR) systems. Recent works have proposed using neural networks to learn word or utter…
Personalized Keyphrase Detection using Speaker and Environment Information
Rajeev Rikhye, Quan Wang, Qiao Liang +6
In this paper, we introduce a streaming keyphrase detection system that can be easily customized to accurately detect any phrase composed of words from a large vocabulary. The syst…