activity
20182022
most citedAdvances in Online Audio-Visual Meeting Transcription

17 citations · 23 across the 7 of their papers we have counts for

collaborators
Showing eess.ASShow all

6 papers · 1 filter

eess.AS2022

Audio Retrieval with WavText5K and CLAP Training

Soham Deshmukh, Benjamin Elizalde, Huaming Wang

Audio-Text retrieval takes a natural language query to retrieve relevant audio files in a database. Conversely, Text-Audio retrieval takes an audio file as a query to retrieve rele…

eess.AS2022

Fast Real-time Personalized Speech Enhancement: End-to-End Enhancement Network (E3Net) and Knowledge Distillation

Manthan Thakker, Sefik Emre Eskimez, Takuya Yoshioka +1

This paper investigates how to improve the runtime speed of personalized speech enhancement (PSE) networks while maintaining the model quality. Our approach includes two aspects: a…

eess.AS2021

One model to enhance them all: array geometry agnostic multi-channel personalized speech enhancement

Hassan Taherian, Sefik Emre Eskimez, Takuya Yoshioka +3

With the recent surge of video conferencing tools usage, providing high-quality speech signals and accurate captions have become essential to conduct day-to-day business or connect…

eess.AS20211 cited

Personalized Speech Enhancement: New Models and Comprehensive Evaluation

Sefik Emre Eskimez, Takuya Yoshioka, Huaming Wang +3

Personalized speech enhancement (PSE) models utilize additional cues, such as speaker embeddings like d-vectors, to remove background noise and interfering speech in real-time and…

eess.AS20212 cited

Human Listening and Live Captioning: Multi-Task Training for Speech Enhancement

Sefik Emre Eskimez, Xiaofei Wang, Min Tang +5

With the surge of online meetings, it has become more critical than ever to provide high-quality speech audio and live captioning under various noise conditions. However, most mona…

eess.AS201917 cited

Advances in Online Audio-Visual Meeting Transcription

Takuya Yoshioka, Igor Abramovski, Cem Aksoylar +23

This paper describes a system that generates speaker-annotated transcripts of meetings by using a microphone array and a 360-degree camera. The hallmark of the system is its abilit…