1 paper
Jeongkyun Park, Jung-Wook Hwang, Kwanghee Choi +4
Inspired by humans comprehending speech in a multi-modal manner, various audio-visual datasets have been constructed. However, most existing datasets focus on English, induce depen…