53 citations · 94 across the 7 of their papers we have counts for
6 papers · 1 filter
(2.5+1)D Spatio-Temporal Scene Graphs for Video Question Answering
Anoop Cherian, Chiori Hori, Tim K. Marks +1
Spatio-temporal scene-graph approaches to video-based reasoning tasks, such as video question-answering (QA), typically construct such graphs for every video frame. These approache…
MOST-GAN: 3D Morphable StyleGAN for Disentangled Face Image Manipulation
Safa C. Medin, Bernhard Egger, Anoop Cherian +4
Recent advances in generative adversarial networks (GANs) have led to remarkable achievements in face image synthesis. While methods that use style-based GANs can generate striking…
LUVLi Face Alignment: Estimating Landmarks' Location, Uncertainty, and Visibility Likelihood
Abhinav Kumar, Tim K. Marks, Wenxuan Mou +6
Modern face alignment methods have become quite accurate at predicting the locations of facial landmarks, but they do not typically estimate the uncertainty of their predicted loca…
Spatio-Temporal Ranked-Attention Networks for Video Captioning
Anoop Cherian, Jue Wang, Chiori Hori +1
Generating video descriptions automatically is a challenging task that involves a complex interplay between spatio-temporal visual features and language models. Given that videos c…
Audio-Visual Scene-Aware Dialog
Huda Alamri, Vincent Cartillier, Abhishek Das +9
We introduce the task of scene-aware dialog. Our goal is to generate a complete and natural response to a question about a scene, given video and audio of the scene and the history…
Class Subset Selection for Transfer Learning using Submodularity
Varun Manjunatha, Srikumar Ramalingam, Tim K. Marks +1
In recent years, it is common practice to extract fully-connected layer (fc) features that were learned while performing image classification on a source dataset, such as ImageNet,…