343 citations · 355 across the 13 of their papers we have counts for
10 papers · 1 filter
Language Complexity and Speech Recognition Accuracy: Orthographic Complexity Hurts, Phonological Complexity Doesn't
Chihiro Taguchi, David Chiang
We investigate what linguistic factors affect the performance of Automatic Speech Recognition (ASR) models. We hypothesize that orthographic and phonological complexities both degr…
PILA: A Historical-Linguistic Dataset of Proto-Italic and Latin
Stephen Bothwell, Brian DuSell, David Chiang +1
Computational historical linguistics seeks to systematically understand processes of sound change, including during periods at which little to no formal recording of language is at…
Killkan: The Automatic Speech Recognition Dataset for Kichwa with Morphosyntactic Information
Chihiro Taguchi, Jefferson Saransig, Dayana Velásquez +1
This paper presents Killkan, the first dataset for automatic speech recognition (ASR) in the Kichwa language, an indigenous language of Ecuador. Kichwa is an extremely low-resource…
Nostra Domina at EvaLatin 2024: Improving Latin Polarity Detection through Data Augmentation
Stephen Bothwell, Abigail Swenor, David Chiang
This paper describes submissions from the team Nostra Domina to the EvaLatin 2024 shared task of emotion polarity detection. Given the low-resource environment of Latin and the com…
BERTwich: Extending BERT's Capabilities to Model Dialectal and Noisy Text
Aarohi Srivastava, David Chiang
Real-world NLP applications often deal with nonstandard text (e.g., dialectal, informal, or misspelled text). However, language models like BERT deteriorate in the face of dialect…
Universal Automatic Phonetic Transcription into the International Phonetic Alphabet
Chihiro Taguchi, Yusuke Sakai, Parisa Haghani +1
This paper presents a state-of-the-art model for transcribing speech in any language into the International Phonetic Alphabet (IPA). Transcription of spoken languages into IPA is a…