92 citations · 256 across the 9 of their papers we have counts for
15 papers
Identifying Actions for Sound Event Classification
Benjamin Elizalde, Radu Revutchi, Samarjit Das +3
In Psychology, actions are paramount for humans to identify sound events. In Machine Learning (ML), action recognition achieves high accuracy; however, it has not been asked whethe…
Learning Question-Guided Video Representation for Multi-Turn Video Question Answering
Guan-Lin Chao, Abhinav Rastogi, Semih Yavuz +3
Understanding and conversing about dynamic scenes is one of the key capabilities of AI agents that navigate the environment and convey useful information to humans. Video question…
BERT-DST: Scalable End-to-End Dialogue State Tracking with Bidirectional Encoder Representations from Transformer
Guan-Lin Chao, Ian Lane
An important yet rarely tackled problem in dialogue state tracking (DST) is scalability for dynamic ontology (e.g., movie, restaurant) and unseen slot values. We focus on a specifi…
Speaker-Targeted Audio-Visual Models for Speech Recognition in Cocktail-Party Environments
Guan-Lin Chao, William Chan, Ian Lane
Speech recognition in cocktail-party environments remains a significant challenge for state-of-the-art speech recognition systems, as it is extremely difficult to extract an acoust…
Speaker Diarization With Lexical Information
Tae Jin Park, Kyu Han, Ian Lane +1
This work presents a novel approach to leverage lexical information for speaker diarization. We introduce a speaker diarization system that can directly integrate lexical as well a…
Understanding and Improving Recurrent Networks for Human Activity Recognition by Continuous Attention
Ming Zeng, Haoxiang Gao, Tong Yu +4
Deep neural networks, including recurrent networks, have been successfully applied to human activity recognition. Unfortunately, the final representation learned by recurrent netwo…