2 papers
cs.CV2024
PaliGemma: A versatile 3B VLM for transfer
Lucas Beyer, Andreas Steiner, André Susano Pinto +32
PaliGemma is an open Vision-Language Model (VLM) that is based on the SigLIP-So400m vision encoder and the Gemma-2B language model. It is trained to be a versatile and broadly know…
cs.CV2024
Conditional Diffusion on Web-Scale Image Pairs leads to Diverse Image Variations
Manoj Kumar, Neil Houlsby, Emiel Hoogeboom
Generating image variations, where a model produces variations of an input image while preserving the semantic context has gained increasing attention. Current image variation tech…