activity
20152021
most citedCodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation

416 citations · 662 across the 9 of their papers we have counts for

collaborators

23 papers

cs.SE2021416 cited

CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation

Shuai Lu, Daya Guo, Shuo Ren +19

Benchmark datasets have a significant impact on accelerating research in programming language tasks. In this paper, we introduce CodeXGLUE, a benchmark dataset to foster machine le…

cs.CL2021

UniSpeech: Unified Speech Representation Learning with Labeled and Unlabeled Data

Chengyi Wang, Yu Wu, Yao Qian +5

In this paper, we propose a unified pre-training approach called UniSpeech to learn speech representations with both unlabeled and labeled data, in which supervised phonetic CTC le…

eess.AS2020

Microsoft Speaker Diarization System for the VoxCeleb Speaker Recognition Challenge 2020

Xiong Xiao, Naoyuki Kanda, Zhuo Chen +10

This paper describes the Microsoft speaker diarization system for monaural multi-talker recordings in the wild, evaluated at the diarization track of the VoxCeleb Speaker Recogniti…

cs.SE2020189 cited

CodeBLEU: a Method for Automatic Evaluation of Code Synthesis

Shuo Ren, Daya Guo, Shuai Lu +7

Evaluation metrics play a vital role in the growth of an area as it defines the standard of distinguishing between good and bad models. In the area of code synthesis, the commonly…

eess.AS20203 cited

MoBoAligner: a Neural Alignment Model for Non-autoregressive TTS with Monotonic Boundary Search

Naihan Li, Shujie Liu, Yanqing Liu +3

To speed up the inference of neural speech synthesis, non-autoregressive models receive increasing attention recently. In non-autoregressive models, additional durations of text to…

eess.AS202018 cited

On the Comparison of Popular End-to-End Models for Large Scale Speech Recognition

Jinyu Li, Yu Wu, Yashesh Gaur +3

Recently, there has been a strong push to transition from hybrid models to end-to-end (E2E) models for automatic speech recognition. Currently, there are three promising E2E method…