9 citations · 10 across the 4 of their papers we have counts for
6 papers
Navigating Connected Memories with a Task-oriented Dialog System
Seungwhan Moon, Satwik Kottur, Alborz Geramifard +1
Recent years have seen an increasing trend in the volume of personal media captured by users, thanks to the advent of smartphones and smart glasses, resulting in large media collec…
Tell Your Story: Task-Oriented Dialogs for Interactive Content Creation
Satwik Kottur, Seungwhan Moon, Aram H. Markosyan +3
People capture photos and videos to relive and share memories of personal significance. Recently, media montages (stories) have become a popular mode of sharing these memories due…
IMU2CLIP: Multimodal Contrastive Learning for IMU Motion Sensors from Egocentric Videos and Text
Seungwhan Moon, Andrea Madotto, Zhaojiang Lin +4
We present IMU2CLIP, a novel pre-training approach to align Inertial Measurement Unit (IMU) motion sensor recordings with video and text, by projecting them into the joint represen…
Connecting What to Say With Where to Look by Modeling Human Attention Traces
Zihang Meng, Licheng Yu, Ning Zhang +4
We introduce a unified framework to jointly model images, text, and human attention traces. Our work is built on top of the recent Localized Narratives annotation framework [30], w…
SIMMC 2.0: A Task-oriented Dialog Dataset for Immersive Multimodal Conversations
Satwik Kottur, Seungwhan Moon, Alborz Geramifard +1
Next generation task-oriented dialog systems need to understand conversational contexts with their perceived surroundings, to effectively help users in the real-world multimodal en…
NN-grams: Unifying neural network and n-gram language models for Speech Recognition
Babak Damavandi, Shankar Kumar, Noam Shazeer +1
We present NN-grams, a novel, hybrid language model integrating n-grams and neural networks (NN) for speech recognition. The model takes as input both word histories as well as n-g…