12 citations · 19 across the 14 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2021
Audio-Visual Scene-Aware Dialog and Reasoning using Audio-Visual Transformers with Joint Student-Teacher Learning
Ankit P. Shah, Shijie Geng, Peng Gao +5
In previous work, we have proposed the Audio-Visual Scene-Aware Dialog (AVSD) task, collected an AVSD dataset, developed AVSD technologies, and hosted an AVSD challenge track at bo…
cs.CL2021★ 1 cited
Advanced Long-context End-to-end Speech Recognition Using Context-expanded Transformers
Takaaki Hori, Niko Moritz, Chiori Hori +1
This paper addresses end-to-end automatic speech recognition (ASR) for long audio recordings such as lecture and conversational speeches. Most end-to-end ASR models are designed to…