4 citations · 5 across the 3 of their papers we have counts for
9 papers · 1 filter
Temporal Query Networks for Fine-grained Video Understanding
Chuhan Zhang, Ankush Gupta, Andrew Zisserman
Our objective in this work is fine-grained classification of actions in untrimmed videos, where the actions may be temporally extended or may span only a few frames of the video. W…
Adaptive Text Recognition through Visual Matching
Chuhan Zhang, Ankush Gupta, Andrew Zisserman
In this work, our objective is to address the problems of generalization and flexibility for text recognition in documents. We introduce a new model that exploits the repetitive na…
CrossTransformers: spatially-aware few-shot transfer
Carl Doersch, Ankush Gupta, Andrew Zisserman
Given new tasks with very little datasuch as new classes in a classification problem or a domain shift in the inputperformance of modern vision systems degrades remarkably qu…
Self-supervised Learning of Interpretable Keypoints from Unlabelled Videos
Tomas Jakab, Ankush Gupta, Hakan Bilen +1
We propose KeypointGAN, a new method for recognizing the pose of objects from a single image that for learning uses only unlabelled videos and a weak empirical prior on the object…
Unsupervised Learning of Object Keypoints for Perception and Control
Tejas Kulkarni, Ankush Gupta, Catalin Ionescu +4
The study of object representations in computer vision has primarily focused on developing representations that are useful for image classification, object detection, or semantic s…
Learning to Read by Spelling: Towards Unsupervised Text Recognition
Ankush Gupta, Andrea Vedaldi, Andrew Zisserman
This work presents a method for visual text recognition without using any paired supervisory data. We formulate the text recognition task as one of aligning the conditional distrib…