53 citations · 112 across the 9 of their papers we have counts for
14 papers
(2.5+1)D Spatio-Temporal Scene Graphs for Video Question Answering
Anoop Cherian, Chiori Hori, Tim K. Marks +1
Spatio-temporal scene-graph approaches to video-based reasoning tasks, such as video question-answering (QA), typically construct such graphs for every video frame. These approache…
Audio-Visual Scene-Aware Dialog and Reasoning using Audio-Visual Transformers with Joint Student-Teacher Learning
Ankit P. Shah, Shijie Geng, Peng Gao +5
In previous work, we have proposed the Audio-Visual Scene-Aware Dialog (AVSD) task, collected an AVSD dataset, developed AVSD technologies, and hosted an AVSD challenge track at bo…
Optimizing Latency for Online Video CaptioningUsing Audio-Visual Transformers
Chiori Hori, Takaaki Hori, Jonathan Le Roux
Video captioning is an essential technology to understand scenes and describe events in natural language. To apply it to real-time monitoring, a system needs not only to describe e…
Advanced Long-context End-to-end Speech Recognition Using Context-expanded Transformers
Takaaki Hori, Niko Moritz, Chiori Hori +1
This paper addresses end-to-end automatic speech recognition (ASR) for long audio recordings such as lecture and conversational speeches. Most end-to-end ASR models are designed to…
Multi-Pass Transformer for Machine Translation
Peng Gao, Chiori Hori, Shijie Geng +2
In contrast with previous approaches where information flows only towards deeper layers of a stack, we consider a multi-pass transformer (MPT) architecture in which earlier layers…
Dynamic Graph Representation Learning for Video Dialog via Multi-Modal Shuffled Transformers
Shijie Geng, Peng Gao, Moitreya Chatterjee +5
Given an input video, its associated audio, and a brief caption, the audio-visual scene aware dialog (AVSD) task requires an agent to indulge in a question-answer dialog with a hum…