most citedIndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages

1 citations · 1 across the 3 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2025

FairI Tales: Evaluation of Fairness in Indian Contexts with a Focus on Bias and Stereotypes

Janki Atul Nawale, Mohammed Safi Ur Rahman Khan, Janani D +3

Existing studies on fairness are largely Western-focused, making them inadequate for culturally diverse countries such as India. To address this gap, we introduce INDIC-BIAS, a com…

cs.CL2024

LAHAJA: A Robust Multi-accent Benchmark for Evaluating Hindi ASR Systems

Tahir Javed, Janki Nawale, Sakshi Joshi +4

Hindi, one of the most spoken language of India, exhibits a diverse array of accents due to its usage among individuals from diverse linguistic origins. To enable a robust evaluati…

cs.CL20241 cited

IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages

Tahir Javed, Janki Atul Nawale, Eldho Ittan George +18

We present INDICVOICES, a dataset of natural and spontaneous speech containing a total of 7348 hours of read (9%), extempore (74%) and conversational (17%) audio from 16237 speaker…

cs.CL20233 cited

Svarah: Evaluating English ASR Systems on Indian Accents

Tahir Javed, Sakshi Joshi, Vignesh Nagarajan +6

India is the second largest English-speaking country in the world with a speaker base of roughly 130 million. Thus, it is imperative that automatic speech recognition (ASR) systems…

cs.CL2023

IndicTrans2: Towards High-Quality and Accessible Machine Translation Models for all 22 Scheduled Indian Languages

Jay Gala, Pranjal A. Chitale, Raghavan AK +11

India has a rich linguistic landscape with languages from 4 major language families spoken by over a billion people. 22 of these languages are listed in the Constitution of India (…