activity
20172021
most citedSee, Hear, and Read: Deep Aligned Representations

68 citations · 127 across the 10 of their papers we have counts for

collaborators

17 papers

cs.CV2021

With a Little Help from My Friends: Nearest-Neighbor Contrastive Learning of Visual Representations

Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson +2

Self-supervised learning algorithms based on instance discrimination train encoders to be invariant to pre-defined transformations of the same instance. While most methods treat di…

cs.RO20211 cited

Manipulator-Independent Representations for Visual Imitation

Yuxiang Zhou, Yusuf Aytar, Konstantinos Bousmalis

Imitation learning is an effective tool for robotic learning tasks where specifying a reinforcement learning (RL) reward is not feasible or where the exploration problem is particu…

cs.RO20212 cited

Learning rich touch representations through cross-modal self-supervision

Martina Zambelli, Yusuf Aytar, Francesco Visin +2

The sense of touch is fundamental in several manipulation tasks, but rarely used in robot manipulation. In this work we tackle the problem of learning rich touch features from cros…

cs.LG20207 cited

Semi-supervised reward learning for offline reinforcement learning

Ksenia Konyushkova, Konrad Zolna, Yusuf Aytar +4

In offline reinforcement learning (RL) agents are trained using a logged dataset. It appears to be the most natural route to attack real-life applications because in domains such a…

cs.LG202014 cited

Offline Learning from Demonstrations and Unlabeled Experience

Konrad Zolna, Alexander Novikov, Ksenia Konyushkova +6

Behavior cloning (BC) is often practical for robot learning because it allows a policy to be trained offline without rewards, by supervised learning on expert demonstrations. Howev…

cs.CV202012 cited

Large-scale multilingual audio visual dubbing

Yi Yang, Brendan Shillingford, Yannis Assael +9

We describe a system for large-scale audiovisual translation and dubbing, which translates videos from one language to another. The source language's speech content is transcribed…