53 citations · 97 across the 4 of their papers we have counts for
5 papers
NeRF-VAE: A Geometry Aware 3D Scene Generative Model
Adam R. Kosiorek, Heiko Strathmann, Daniel Zoran +4
We propose NeRF-VAE, a 3D scene generative model that incorporates geometric structure via NeRF and differentiable volume rendering. In contrast to NeRF, our model takes into accou…
Decoupling the Role of Data, Attention, and Losses in Multimodal Transformers
Lisa Anne Hendricks, John Mellor, Rosalia Schneider +2
Recently multimodal transformer models have gained popularity because their performance on language and vision tasks suggest they learn rich visual-linguistic representations. Focu…
Self-Supervised MultiModal Versatile Networks
Jean-Baptiste Alayrac, Adrià Recasens, Rosalia Schneider +6
Videos are a rich source of multi-modal supervision. In this work, we learn representations using self-supervision by leveraging three modalities naturally present in videos: visua…
Probing Emergent Semantics in Predictive Agents via Question Answering
Abhishek Das, Federico Carnevale, Hamza Merzic +8
Recent work has shown how predictive modeling can endow agents with rich knowledge of their surroundings, improving their ability to act in complex environments. We propose questio…
Environmental drivers of systematicity and generalization in a situated agent
Felix Hill, Andrew Lampinen, Rosalia Schneider +4
The question of whether deep neural networks are good at generalising beyond their immediate training experience is of critical importance for learning-based approaches to AI. Here…