most citedMultilingual and Fully Non-Autoregressive ASR with Large Language Model Fusion: A Comprehensive Study

1 citations · 2 across the 3 of their papers we have counts for

collaborators

5 papers

eess.AS2025

Towards General Auditory Intelligence: Large Multimodal Models for Machine Listening and Speaking

Siyin Wang, Zengrui Jin, Changli Tang +26

In the era of large language models (LLMs) and artificial general intelligence (AGI), computer audition must evolve beyond traditional paradigms to fully leverage the capabilities…

cs.CV2024

VISTA: A Visual and Textual Attention Dataset for Interpreting Multimodal Models

Harshit, Tolga Tasdizen

The recent developments in deep learning led to the integration of natural language processing (NLP) with computer vision, resulting in powerful integrated Vision and Language Mode…

cs.CL20241 cited

IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages

Tahir Javed, Janki Atul Nawale, Eldho Ittan George +18

We present INDICVOICES, a dataset of natural and spontaneous speech containing a total of 7348 hours of read (9%), extempore (74%) and conversational (17%) audio from 16237 speaker…

cs.CL2024

Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Gemini Team, Petko Georgiev, Ving Ian Lei +1132

In this report, we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning over…

cs.CL20241 cited

Multilingual and Fully Non-Autoregressive ASR with Large Language Model Fusion: A Comprehensive Study

W. Ronny Huang, Cyril Allauzen, Tongzhou Chen +7

In the era of large models, the autoregressive nature of decoding often results in latency serving as a significant bottleneck. We propose a non-autoregressive LM-fused ASR system…