activity
20192022
most citedend-to-end training of a large vocabulary end-to-end speech recognition system

9 citations · 13 across the 6 of their papers we have counts for

collaborators

10 papers

cs.CV20221 cited

Transformer-based Global 3D Hand Pose Estimation in Two Hands Manipulating Objects Scenarios

Hoseong Cho, Donguk Kim, Chanwoo Kim +2

This report describes our 1st place solution to ECCV 2022 challenge on Human Body, Hands, and Activities (HBHA) from Egocentric and Multi-view Cameras (hand pose estimation). In th…

cs.SD2021

Streaming end-to-end speech recognition with jointly trained neural feature enhancement

Chanwoo Kim, Abhinav Garg, Dhananjaya Gowda +2

In this paper, we present a streaming end-to-end speech recognition model based on Monotonic Chunkwise Attention (MoCha) jointly trained with enhancement layers. Even though the Mo…

cs.CL2020

Faster Re-translation Using Non-Autoregressive Model For Simultaneous Neural Machine Translation

Hyojung Han, Sathish Indurthi, Mohd Abbas Zaidi +5

Recently, simultaneous translation has gathered a lot of attention since it enables compelling applications such as subtitle translation for a live event or real-time video-call tr…

cs.LG2020

A review of on-device fully neural end-to-end automatic speech recognition algorithms

Chanwoo Kim, Dhananjaya Gowda, Dongsoo Lee +5

In this paper, we review various end-to-end automatic speech recognition algorithms and their optimization techniques for on-device applications. Conventional speech recognition sy…

eess.AS2020

Small energy masking for improved neural network training for end-to-end speech recognition

Chanwoo Kim, Kwangyoun Kim, Sathish Reddy Indurthi

In this paper, we present a Small Energy Masking (SEM) algorithm, which masks inputs having values below a certain threshold. More specifically, a time-frequency bin is masked if t…

eess.AS2020

Attention based on-device streaming speech recognition with large speech corpus

Kwangyoun Kim, Kyungmin Lee, Dhananjaya Gowda +10

In this paper, we present a new on-device automatic speech recognition (ASR) system based on monotonic chunk-wise attention (MoChA) models trained with large (> 10K hours) corpus.…