193 citations · 211 across the 8 of their papers we have counts for
4 papers · 1 filter
SODA: Bottleneck Diffusion Models for Representation Learning
Drew A. Hudson, Daniel Zoran, Mateusz Malinowski +6
We introduce SODA, a self-supervised diffusion model, designed for representation learning. The model incorporates an image encoder, which distills a source view into a compact rep…
Leveraging VLM-Based Pipelines to Annotate 3D Objects
Rishabh Kabra, Loic Matthey, Alexander Lerchner +1
Pretrained vision language models (VLMs) present an opportunity to caption unlabeled 3D objects at scale. The leading approach to summarize VLM descriptions from different views of…
AlignNet: Unsupervised Entity Alignment
Antonia Creswell, Kyriacos Nikiforou, Oriol Vinyals +8
Recently developed deep learning models are able to learn to segment scenes into component objects without supervision. This opens many new and exciting avenues of research, allowi…
MONet: Unsupervised Scene Decomposition and Representation
Christopher P. Burgess, Loic Matthey, Nicholas Watters +4
The ability to decompose scenes in terms of abstract building blocks is crucial for general intelligence. Where those basic building blocks share meaningful properties, interaction…