most citedEfficient Vision-Language Pretraining with Visual Concepts and Hierarchical Alignment

2 citations · 3 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV20241 cited

What Makes Multimodal In-Context Learning Work?

Folco Bertini Baldassini, Mustafa Shukor, Matthieu Cord +2

Large Language Models have demonstrated remarkable performance across various tasks, exhibiting the capacity to swiftly acquire new skills, such as through In-Context Learning (ICL…

cs.CV2024

Improved Baselines for Data-efficient Perceptual Augmentation of LLMs

Théophane Vallaeys, Mustafa Shukor, Matthieu Cord +1

The abilities of large language models (LLMs) have recently progressed to unprecedented levels, paving the way to novel applications in a wide variety of areas. In computer vision,…

cs.CV20222 cited

Efficient Vision-Language Pretraining with Visual Concepts and Hierarchical Alignment

Mustafa Shukor, Guillaume Couairon, Matthieu Cord

Vision and Language Pretraining has become the prevalent approach for tackling multimodal downstream tasks. The current trend is to move towards ever larger models and pretraining…

eess.IV2022

Video Coding Using Learned Latent GAN Compression

Mustafa Shukor, Bharath Bhushan Damodaran, Xu Yao +1

We propose in this paper a new paradigm for facial video compression. We leverage the generative capacity of GANs such as StyleGAN to represent and compress a video, including intr…

cs.CV2022

Semantic Unfolding of StyleGAN Latent Space

Mustafa Shukor, Xu Yao, Bharath Bushan Damodaran +1

Generative adversarial networks (GANs) have proven to be surprisingly efficient for image editing by inverting and manipulating the latent code corresponding to an input real image…