4 papers
Domain Adaptation with a Single Vision-Language Embedding
Mohammad Fahes, Tuan-Hung Vu, Andrei Bursuc +2
Domain adaptation has been extensively investigated in computer vision but still requires access to target data at the training time, which might be difficult to obtain in real-wor…
CLIP's Visual Embedding Projector is a Few-shot Cornucopia
Mohammad Fahes, Tuan-Hung Vu, Andrei Bursuc +2
We introduce ProLIP, a simple and architecture-agnostic method for adapting contrastively pretrained vision-language models, such as CLIP, to few-shot classification. ProLIP fine-t…
Vision-Speech Models: Teaching Speech Models to Converse about Images
Amélie Royer, Moritz Böhle, Gabriel de Marmiesse +4
The recent successes of Vision-Language models raise the question of how to equivalently imbue a pretrained speech model with vision understanding, an important milestone towards b…
High-Fidelity Simultaneous Speech-To-Speech Translation
Tom Labiausse, Laurent Mazaré, Edouard Grave +3
We introduce Hibiki, a decoder-only model for simultaneous speech translation. Hibiki leverages a multistream language model to synchronously process source and target speech, and…