3 citations · 3 across the 2 of their papers we have counts for
7 papers
Reasoning Over History: Context Aware Visual Dialog
Muhammad A. Shah, Shikib Mehri, Tejas Srinivasan
While neural models have been shown to exhibit strong performance on single-turn visual question answering (VQA) tasks, extending VQA to a multi-turn, conversational setting remain…
Multimodal Speech Recognition with Unstructured Audio Masking
Tejas Srinivasan, Ramon Sanabria, Florian Metze +1
Visual context has been shown to be useful for automatic speech recognition (ASR) systems when the speech signal is noisy or corrupted. Previous work, however, has only demonstrate…
Fine-Grained Grounding for Multimodal Speech Recognition
Tejas Srinivasan, Ramon Sanabria, Florian Metze +1
Multimodal automatic speech recognition systems integrate information from images to improve speech recognition quality, by grounding the speech in the visual context. While visual…
Looking Enhances Listening: Recovering Missing Speech Using Images
Tejas Srinivasan, Ramon Sanabria, Florian Metze
Speech is understood better by using visual context; for this reason, there have been many attempts to use images to adapt automatic speech recognition (ASR) systems. Current work,…
Multitask Learning For Different Subword Segmentations In Neural Machine Translation
Tejas Srinivasan, Ramon Sanabria, Florian Metze
In Neural Machine Translation (NMT) the usage of subwords and characters as source and target units offers a simple and flexible solution for translation of rare and unseen words.…
Structured Fusion Networks for Dialog
Shikib Mehri, Tejas Srinivasan, Maxine Eskenazi
Neural dialog models have exhibited strong performance, however their end-to-end nature lacks a representation of the explicit structure of dialog. This results in a loss of genera…