activity
20152021
most citedAn End-to-End Trainable Neural Network Model with Belief Tracking for Task-Oriented Dialog

92 citations · 256 across the 9 of their papers we have counts for

collaborators

15 papers

cs.SD2021

Identifying Actions for Sound Event Classification

Benjamin Elizalde, Radu Revutchi, Samarjit Das +3

In Psychology, actions are paramount for humans to identify sound events. In Machine Learning (ML), action recognition achieves high accuracy; however, it has not been asked whethe…

cs.CL2019

Learning Question-Guided Video Representation for Multi-Turn Video Question Answering

Guan-Lin Chao, Abhinav Rastogi, Semih Yavuz +3

Understanding and conversing about dynamic scenes is one of the key capabilities of AI agents that navigate the environment and convey useful information to humans. Video question…

cs.CL201911 cited

BERT-DST: Scalable End-to-End Dialogue State Tracking with Bidirectional Encoder Representations from Transformer

Guan-Lin Chao, Ian Lane

An important yet rarely tackled problem in dialogue state tracking (DST) is scalability for dynamic ontology (e.g., movie, restaurant) and unseen slot values. We focus on a specifi…

eess.AS20193 cited

Speaker-Targeted Audio-Visual Models for Speech Recognition in Cocktail-Party Environments

Guan-Lin Chao, William Chan, Ian Lane

Speech recognition in cocktail-party environments remains a significant challenge for state-of-the-art speech recognition systems, as it is extremely difficult to extract an acoust…

cs.CL2018

Speaker Diarization With Lexical Information

Tae Jin Park, Kyu Han, Ian Lane +1

This work presents a novel approach to leverage lexical information for speaker diarization. We introduce a speaker diarization system that can directly integrate lexical as well a…

cs.LG2018

Understanding and Improving Recurrent Networks for Human Activity Recognition by Continuous Attention

Ming Zeng, Haoxiang Gao, Tong Yu +4

Deep neural networks, including recurrent networks, have been successfully applied to human activity recognition. Unfortunately, the final representation learned by recurrent netwo…