1 citations · 2 across the 3 of their papers we have counts for
4 papers
IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages
Tahir Javed, Janki Atul Nawale, Eldho Ittan George +18
We present INDICVOICES, a dataset of natural and spontaneous speech containing a total of 7348 hours of read (9%), extempore (74%) and conversational (17%) audio from 16237 speaker…
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Petko Georgiev, Ving Ian Lei +1132
In this report, we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning over…
E3 TTS: Easy End-to-End Diffusion-based Text to Speech
Yuan Gao, Nobuyuki Morioka, Yu Zhang +1
We propose Easy End-to-End Diffusion-based Text to Speech, a simple and efficient end-to-end text-to-speech model based on diffusion. E3 TTS directly takes plain text as input and…
SLM: Bridge the thin gap between speech and text foundation models
Mingqiu Wang, Wei Han, Izhak Shafran +15
We present a joint Speech and Language Model (SLM), a multitask, multilingual, and dual-modal model that takes advantage of pretrained foundational speech and language models. SLM…