activity
20162023
most citedVLP: A Survey on Vision-Language Pre-training

271 citations · 429 across the 17 of their papers we have counts for

collaborators
Showing 2020 · eess.ASShow all

9 papers · 2 filters

eess.AS2020★ 6 cited

The 2020 ESPnet update: new features, broadened applications, performance improvements, and future plans

Shinji Watanabe, Florian Boyer, Xuankai Chang +12

This paper describes the recent development of ESPnet (https://github.com/espnet/espnet), an end-to-end speech processing toolkit. This project was initiated in December 2017 to ma…

eess.AS2020

ESPnet-se: end-to-end speech enhancement and separation toolkit designed for asr integration

Chenda Li, Jing Shi, Wangyou Zhang +8

We present ESPnet-SE, which is designed for the quick development of speech enhancement and speech separation systems in a single framework, along with the optional downstream spee…

eess.AS2020★ 40 cited

Recent Developments on ESPnet Toolkit Boosted by Conformer

Pengcheng Guo, Florian Boyer, Xuankai Chang +12

In this study, we present recent developments on ESPnet: End-to-End Speech Processing toolkit, which mainly involves a recently proposed architecture called Conformer, Convolution-…

eess.AS2020

Training Noisy Single-Channel Speech Separation With Noisy Oracle Sources: A Large Gap and A Small Step

Matthew Maciejewski, Jing Shi, Shinji Watanabe +1

As the performance of single-channel speech separation systems has improved, there has been a desire to move to more challenging conditions than the clean, near-field speech that i…

eess.AS2020

Multi-task Metric Learning for Text-independent Speaker Verification

Yafeng Chen, Wu Guo, Jingjing Shi +2

In this work, we introduce metric learning (ML) to enhance the deep embedding learning for text-independent speaker verification (SV). Specifically, the deep speaker embedding netw…

eess.AS2020

Exploring Universal Speech Attributes for Speaker Verification with an Improved Cross-stitch Network

Jiajun Qi, Wu Guo, Jingjing Shi +2

The universal speech attributes for x-vector based speaker verification (SV) are addressed in this paper. The manner and place of articulation form the fundamental speech attribute…