297 citations · 299 across the 5 of their papers we have counts for
3 papers · 1 filter
IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages
Tahir Javed, Janki Atul Nawale, Eldho Ittan George +18
We present INDICVOICES, a dataset of natural and spontaneous speech containing a total of 7348 hours of read (9%), extempore (74%) and conversational (17%) audio from 16237 speaker…
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Petko Georgiev, Ving Ian Lei +1132
In this report, we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning over…
Multilingual and Fully Non-Autoregressive ASR with Large Language Model Fusion: A Comprehensive Study
W. Ronny Huang, Cyril Allauzen, Tongzhou Chen +7
In the era of large models, the autoregressive nature of decoding often results in latency serving as a significant bottleneck. We propose a non-autoregressive LM-fused ASR system…