17 citations · 23 across the 7 of their papers we have counts for
6 papers · 1 filter
Audio Retrieval with WavText5K and CLAP Training
Soham Deshmukh, Benjamin Elizalde, Huaming Wang
Audio-Text retrieval takes a natural language query to retrieve relevant audio files in a database. Conversely, Text-Audio retrieval takes an audio file as a query to retrieve rele…
Fast Real-time Personalized Speech Enhancement: End-to-End Enhancement Network (E3Net) and Knowledge Distillation
Manthan Thakker, Sefik Emre Eskimez, Takuya Yoshioka +1
This paper investigates how to improve the runtime speed of personalized speech enhancement (PSE) networks while maintaining the model quality. Our approach includes two aspects: a…
One model to enhance them all: array geometry agnostic multi-channel personalized speech enhancement
Hassan Taherian, Sefik Emre Eskimez, Takuya Yoshioka +3
With the recent surge of video conferencing tools usage, providing high-quality speech signals and accurate captions have become essential to conduct day-to-day business or connect…
Personalized Speech Enhancement: New Models and Comprehensive Evaluation
Sefik Emre Eskimez, Takuya Yoshioka, Huaming Wang +3
Personalized speech enhancement (PSE) models utilize additional cues, such as speaker embeddings like d-vectors, to remove background noise and interfering speech in real-time and…
Human Listening and Live Captioning: Multi-Task Training for Speech Enhancement
Sefik Emre Eskimez, Xiaofei Wang, Min Tang +5
With the surge of online meetings, it has become more critical than ever to provide high-quality speech audio and live captioning under various noise conditions. However, most mona…
Advances in Online Audio-Visual Meeting Transcription
Takuya Yoshioka, Igor Abramovski, Cem Aksoylar +23
This paper describes a system that generates speaker-annotated transcripts of meetings by using a microphone array and a 360-degree camera. The hallmark of the system is its abilit…