activity
20152022
most citedMake-A-Video: Text-to-Video Generation without Text-Video Data

315 citations · 955 across the 38 of their papers we have counts for

collaborators
Showing 2021Show all

6 papers · 1 filter

cs.HC20213 cited

Telling Creative Stories Using Generative Visual Aids

Safinah Ali, Devi Parikh

Can visual artworks created using generative visual algorithms inspire human creativity in storytelling? We asked writers to write creative stories from a starting prompt, and prov…

cs.SD20218 cited

Dance2Music: Automatic Dance-driven Music Generation

Gunjan Aggarwal, Devi Parikh

Dance and music typically go hand in hand. The complexities in dance, music, and their synchronisation make them fascinating to study from a computational creativity perspective. W…

cs.CL20214 cited

Visual Conceptual Blending with Large-scale Language and Vision Models

Songwei Ge, Devi Parikh

We ask the question: to what extent can recent large-scale language and image generation models blend visual concepts? Given an arbitrary object, we identify a relevant object and…

cs.CV202122 cited

Human-Adversarial Visual Question Answering

Sasha Sheng, Amanpreet Singh, Vedanuj Goswami +4

Performance on the most commonly used Visual Question Answering dataset (VQA v2) is starting to approach human accuracy. However, in interacting with state-of-the-art VQA models, i…

cs.LG202125 cited

ForceNet: A Graph Neural Network for Large-Scale Quantum Calculations

Weihua Hu, Muhammed Shuaibi, Abhishek Das +5

With massive amounts of atomic simulation data available, there is a huge opportunity to develop fast and accurate machine learning models to approximate expensive physics-based ca…

cs.CV20218 cited

VX2TEXT: End-to-End Learning of Video-Based Text Generation From Multimodal Inputs

Xudong Lin, Gedas Bertasius, Jue Wang +3

We present \textsc{Vx2Text}, a framework for text generation from multimodal inputs consisting of video plus text, speech, or audio. In order to leverage transformer networks, whic…