29 citations · 36 across the 5 of their papers we have counts for
12 papers · 1 filter
Encyclopedic VQA: Visual questions about detailed properties of fine-grained categories
Thomas Mensink, Jasper Uijlings, Lluis Castrejon +6
We propose Encyclopedic-VQA, a large scale visual question answering (VQA) dataset featuring visual questions about detailed properties of fine-grained categories and instances. It…
Infinite Class Mixup
Thomas Mensink, Pascal Mettes
Mixup is a widely adopted strategy for training deep networks, where additional samples are augmented by interpolating inputs and labels of training pairs. Mixup has shown to impro…
Scaling Vision Transformers to 22 Billion Parameters
Mostafa Dehghani, Josip Djolonga, Basil Mustafa +39
The scaling of Transformers has driven breakthrough capabilities for language models. At present, the largest large language models (LLMs) contain upwards of 100B parameters. Visio…
EDEN: Multimodal Synthetic Dataset of Enclosed GarDEN Scenes
Hoang-An Le, Thomas Mensink, Partha Das +2
Multimodal large-scale datasets for outdoor scenes are mostly designed for urban driving problems. The scenes are highly structured and semantically different from scenarios seen i…
Novel View Synthesis from Single Images via Point Cloud Transformation
Hoang-An Le, Thomas Mensink, Partha Das +1
In this paper the argument is made that for true novel view synthesis of objects, where the object can be synthesized from any viewpoint, an explicit 3D shape representation isdesi…
Multi-Loss Weighting with Coefficient of Variations
Rick Groenendijk, Sezer Karaoglu, Theo Gevers +1
Many interesting tasks in machine learning and computer vision are learned by optimising an objective function defined as a weighted linear combination of multiple losses. The fina…