activity
20242026
collaborators

5 papers

cs.CV2026

ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations

Junke Wang, Xiao Wang, Jiacheng Pan +16

This paper introduces ARM, a discrete representation-based AutoRegressive Model that unifies image understanding, generation, and editing within a next-token prediction framework.…

cs.CV2026

Scaling Pre-training to One Hundred Billion Data for Vision Language Models

Xiao Wang, Ibrahim Alabdulmohsin, Daniel Salz +3

We provide an empirical investigation of the potential of pre-training vision-language models on an unprecedented scale: 100 billion examples. We find that model performance tends…

cond-mat.mtrl-sci2025

Zero-shot Autonomous Microscopy for Scalable and Intelligent Characterization of 2D Materials

Jingyun Yang, Ruoyan Avery Yin, Chi Jiang +14

Characterization of atomic-scale materials traditionally requires human experts with months to years of specialized training. Even for trained human operators, accurate and reliabl…

cs.CV2025

SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Michael Tschannen, Alexey Gritsenko, Xiao Wang +11

We introduce SigLIP 2, a family of new multilingual vision-language encoders that build on the success of the original SigLIP. In this second iteration, we extend the original imag…

cs.CV2024

PaliGemma 2: A Family of Versatile VLMs for Transfer

Andreas Steiner, André Susano Pinto, Michael Tschannen +15

PaliGemma 2 is an upgrade of the PaliGemma open Vision-Language Model (VLM) based on the Gemma 2 family of language models. We combine the SigLIP-So400m vision encoder that was als…