activity
20172026
most citedRT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

273 citations · 410 across the 21 of their papers we have counts for

collaborators
Showing cs.CVShow all

16 papers · 1 filter

cs.CV2026

Image Generators are Generalist Vision Learners

Valentin Gabeur, Shangbang Long, Songyou Peng +22

Recent works show that image and video generators exhibit zero-shot visual understanding behaviors, in a way reminiscent of how LLMs develop emergent capabilities of language under…

cs.CV2025

Unified Autoregressive Visual Generation and Understanding with Continuous Tokens

Lijie Fan, Luming Tang, Siyang Qin +11

We present UniFluid, a unified autoregressive framework for joint visual generation and understanding leveraging continuous visual tokens. Our unified autoregressive architecture p…

cs.CV2024

PaliGemma: A versatile 3B VLM for transfer

Lucas Beyer, Andreas Steiner, André Susano Pinto +32

PaliGemma is an open Vision-Language Model (VLM) that is based on the SigLIP-So400m vision encoder and the Gemma-2B language model. It is trained to be a versatile and broadly know…

cs.CV2024

Wavelet-Based Image Tokenizer for Vision Transformers

Zhenhai Zhu, Radu Soricut

Non-overlapping patch-wise convolution is the default image tokenizer for all state-of-the-art vision Transformer (ViT) models. Even though many ViT variants have been proposed to…

cs.CV2024

ImageInWords: Unlocking Hyper-Detailed Image Descriptions

Roopal Garg, Andrea Burns, Burcu Karagol Ayan +7

Despite the longstanding adage "an image is worth a thousand words," generating accurate hyper-detailed image descriptions remains unsolved. Trained on short web-scraped image text…

cs.CV2023

GeomVerse: A Systematic Evaluation of Large Models for Geometric Reasoning

Mehran Kazemi, Hamidreza Alvari, Ankit Anand +3

Large language models have shown impressive results for multi-hop mathematical reasoning when the input question is only textual. Many mathematical reasoning problems, however, con…