collaborators

5 papers

cs.SD2026

CS-YODAS: A Mined Dataset of In-the-Wild Code-Switched Speech

Brian Yan, Qingzheng Wang, Matthew Wiesner +9

We present CS-YODAS, a Creative Commons-licensed dataset of in-the-wild code-switched speech mined from multilingual YouTube data. Code-switching (CS), or the alternation between l…

cs.CL2026

Beyond Transcripts: Iterative Peer-Editing with Audio Unlocks High-Quality Human Summaries of Conversational Speech

Kaavya Chaparala, Thomas Thebaud, Jesús Villalba López +3

There are not enough established benchmarks for the task fo speech summarization. Creating new benchmarks demands human annotation, as LLMs could embed systemic errors and bias int…

cs.CL2025

Enhancing Dialogue Annotation with Speaker Characteristics Leveraging a Frozen LLM

Thomas Thebaud, Yen-Ju Lu, Matthew Wiesner +2

In dialogue transcription pipelines, Large Language Models (LLMs) are frequently employed in post-processing to improve grammar, punctuation, and readability. We explore a compleme…

cs.CL2025

Measurement of the Granularity of Vowel Production Space By Just Producible Different (JPD) Limens

Peter Viechnicki

A body of work over the past several decades has demonstrated that the complex and coordinated articulatory movements of human vowel production are governed (at least in part)by co…

stat.ML2025

Asymptotically perfect seeded graph matching without edge correlation (and applications to inference)

Tong Qi, Vera Andersson, Peter Viechnicki +1

We present the OmniMatch algorithm for seeded multiple graph matching. In the setting of -dimensional Random Dot Product Graphs (RDPG), we prove that under mild assumptions, Omn…