118 citations · 217 across the 3 of their papers we have counts for
3 papers
cs.CV2023★ 118 cited
Scaling Vision Transformers to 22 Billion Parameters
Mostafa Dehghani, Josip Djolonga, Basil Mustafa +39
The scaling of Transformers has driven breakthrough capabilities for language models. At present, the largest large language models (LLMs) contain upwards of 100B parameters. Visio…
cs.CV2016★ 14 cited
ResearchDoom and CocoDoom: Learning Computer Vision with Games
A. Mahendran, H. Bilen, J. F. Henriques +1
In this short note we introduce ResearchDoom, an implementation of the Doom first-person shooter that can extract detailed metadata from the game. We also introduce the CocoDoom da…
cs.CV2014★ 85 cited
Understanding Deep Image Representations by Inverting Them
Aravindh Mahendran, Andrea Vedaldi
Image representations, from SIFT and Bag of Visual Words to Convolutional Neural Networks (CNNs), are a crucial component of almost any image understanding system. Nevertheless, ou…