5 citations · 9 across the 3 of their papers we have counts for
4 papers
A Multi-View Approach To Audio-Visual Speaker Verification
Leda Sarı, Kritika Singh, Jiatong Zhou +3
Although speaker verification has conventionally been an audio-only task, some practical applications provide both audio and visual streams of input. In these cases, the visual str…
Deep F-measure Maximization for End-to-End Speech Understanding
Leda Sarı, Mark Hasegawa-Johnson
Spoken language understanding (SLU) datasets, like many other machine learning datasets, usually suffer from the label imbalance problem. Label imbalance usually causes the learned…
Identify Speakers in Cocktail Parties with End-to-End Attention
Junzhe Zhu, Mark Hasegawa-Johnson, Leda Sari
In scenarios where multiple speakers talk at the same time, it is important to be able to identify the talkers accurately. This paper presents an end-to-end system that integrates…
Unsupervised Speaker Adaptation using Attention-based Speaker Memory for End-to-End ASR
Leda Sarı, Niko Moritz, Takaaki Hori +1
We propose an unsupervised speaker adaptation method inspired by the neural Turing machine for end-to-end (E2E) automatic speech recognition (ASR). The proposed model contains a me…