73 citations · 309 across the 34 of their papers we have counts for
9 papers · 1 filter
Improved Few-Shot Visual Classification
Peyman Bateni, Raghav Goyal, Vaden Masrani +2
Few-shot learning is a fundamental task in computer vision that carries the promise of alleviating the need for exhaustively labeled data. Most few-shot learning approaches to date…
Generating Videos of Zero-Shot Compositions of Actions and Objects
Megha Nawhal, Mengyao Zhai, Andreas Lehrmann +2
Human activity videos involve rich, varied interactions between people and objects. In this paper we develop methods for generating such videos -- making progress toward addressing…
OptiBox: Breaking the Limits of Proposals for Visual Grounding
Zicong Fan, Si Yi Meng, Leonid Sigal +1
The problem of language grounding has attracted much attention in recent years due to its pivotal role in more general image-lingual high level reasoning tasks (e.g., image caption…
Watch, Listen and Tell: Multi-modal Weakly Supervised Dense Event Captioning
Tanzila Rahman, Bicheng Xu, Leonid Sigal
Multi-modal learning, particularly among imaging and linguistic modalities, has made amazing strides in many high-level fundamental visual understanding problems, ranging from lang…
DwNet: Dense warp-based network for pose-guided human video generation
Polina Zablotskaia, Aliaksandr Siarohin, Bo Zhao +1
Generation of realistic high-resolution videos of human subjects is a challenging and important task in computer vision. In this paper, we focus on human motion transfer - generati…
LayoutVAE: Stochastic Scene Layout Generation From a Label Set
Akash Abdu Jyothi, Thibaut Durand, Jiawei He +2
Recently there is an increasing interest in scene generation within the research community. However, models used for generating scene layouts from textual description largely ignor…