47 citations · 47 across the 3 of their papers we have counts for
5 papers · 1 filter
Image Generators are Generalist Vision Learners
Valentin Gabeur, Shangbang Long, Songyou Peng +22
Recent works show that image and video generators exhibit zero-shot visual understanding behaviors, in a way reminiscent of how LLMs develop emergent capabilities of language under…
Unified Autoregressive Visual Generation and Understanding with Continuous Tokens
Lijie Fan, Luming Tang, Siyang Qin +11
We present UniFluid, a unified autoregressive framework for joint visual generation and understanding leveraging continuous visual tokens. Our unified autoregressive architecture p…
ImageInWords: Unlocking Hyper-Detailed Image Descriptions
Roopal Garg, Andrea Burns, Burcu Karagol Ayan +7
Despite the longstanding adage "an image is worth a thousand words," generating accurate hyper-detailed image descriptions remains unsolved. Trained on short web-scraped image text…
PaliGemma: A versatile 3B VLM for transfer
Lucas Beyer, Andreas Steiner, André Susano Pinto +32
PaliGemma is an open Vision-Language Model (VLM) that is based on the SigLIP-So400m vision encoder and the Gemma-2B language model. It is trained to be a versatile and broadly know…
Wavelet-Based Image Tokenizer for Vision Transformers
Zhenhai Zhu, Radu Soricut
Non-overlapping patch-wise convolution is the default image tokenizer for all state-of-the-art vision Transformer (ViT) models. Even though many ViT variants have been proposed to…