activity
20172024
most citedDINOv2: Learning Robust Visual Features without Supervision

1.1k citations · 1.6k across the 21 of their papers we have counts for

collaborators

24 papers

cs.AI2024

The Llama 3 Herd of Models

Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri +556

Modern artificial intelligence (AI) systems are powered by foundation models. This paper presents a new set of foundation models, called Llama 3. It is a herd of language models th…

eess.AS2024

VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild

Puyuan Peng, Po-Yao Huang, Shang-Wen Li +2

We introduce VoiceCraft, a token infilling neural codec language model, that achieves state-of-the-art performance on both speech editing and zero-shot text-to-speech (TTS) on audi…

cs.CV2023★ 22 cited

Demystifying CLIP Data

Hu Xu, Saining Xie, Xiaoqing Ellen Tan +7

Contrastive Language-Image Pre-training (CLIP) is an approach that has advanced research and applications in computer vision, fueling modern recognition systems and generative mode…

cs.CV2023★ 61 cited

Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles

Chaitanya Ryali, Yuan-Ting Hu, Daniel Bolya +10

Modern hierarchical vision transformers have added several vision-specific components in the pursuit of supervised classification performance. While these components lead to effect…

cs.CV2023★ 2 cited

Diffusion Models as Masked Autoencoders

Chen Wei, Karttikeya Mangalam, Po-Yao Huang +7

There has been a longstanding belief that generation can facilitate a true understanding of visual data. In line with this, we revisit generatively pre-training visual representati…

cs.CV2023★ 1.1k cited

DINOv2: Learning Robust Visual Features without Supervision

Maxime Oquab, Timothée Darcet, Théo Moutakanni +23

The recent breakthroughs in natural language processing for model pretraining on large quantities of data have opened the way for similar foundation models in computer vision. Thes…