activity
20242026
collaborators

7 papers

cs.CV2026

Vision Foundation Models as Generalist Tokenizers for Image Generation

Anlin Zheng, Qi Han, Xin Wen +5

In this work, we explore the largely unexplored direction of building a generalist image tokenizer directly on top of a frozen vision foundation model (VFM). To build this tokenize…

cs.CV2026

Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Representation Tokens

Yuqing Wang, Chuofan Ma, Zhijie Lin +7

Visual generation with discrete tokens has gained significant attention as it enables a unified token prediction paradigm shared with language models, promising seamless multimodal…

cs.CV2025

Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation

Anlin Zheng, Xin Wen, Xuanyang Zhang +5

In this work, we present a novel direction to build an image tokenizer directly on top of a frozen vision foundation model, which is a largely underexplored area. Specifically, we…

cs.CV2025

UniTok: A Unified Tokenizer for Visual Generation and Understanding

Chuofan Ma, Yi Jiang, Junfeng Wu +5

Visual generative and understanding models typically rely on distinct tokenizers to process images, presenting a key challenge for unifying them within a single framework. Recent s…

cs.CV2025

Recognize Any Regions

Haosen Yang, Chuofan Ma, Bin Wen +3

Understanding the semantics of individual regions or patches of unconstrained images, such as open-world object detection, remains a critical yet challenging task in computer visio…

cs.CV2025

Liquid: Language Models are Scalable and Unified Multi-modal Generators

Junfeng Wu, Yi Jiang, Chuofan Ma +5

We present Liquid, an auto-regressive generation paradigm that seamlessly integrates visual comprehension and generation by tokenizing images into discrete codes and learning these…