5 papers · 1 filter
DINO-Tok: Adapting DINO for Visual Tokenizers
Mingkai Jia, Mingxiao Li, Zhijian Shu +12
Recent advances in visual generation have emphasized the importance of Latent Generative Models (LGMs), which critically depend on effective visual tokenizers to bridge pixels and…
3D and 4D World Modeling: A Survey
Lingdong Kong, Yu Yang, Jianbiao Mei +20
World modeling has become a cornerstone in AI research, enabling agents to understand, represent, and predict the dynamic environments they inhabit. While prior work largely emphas…
2D Gaussians Meet Visual Tokenizer
Yiang Shi, Xiaoyang Guo, Wei Yin +5
The image tokenizer is a critical component in AR image generation, as it determines how rich and structured visual content is encoded into compact representations. Existing quanti…
MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization
Mingkai Jia, Wei Yin, Xiaotao Hu +5
Vector Quantized Variational Autoencoders (VQ-VAEs) are fundamental models that compress continuous visual data into discrete tokens. Existing methods have tried to improve the qua…
DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT
Xiaotao Hu, Mingkai Jia, Xiaoyang Guo +3
Recent successes in autoregressive (AR) generation models, such as the GPT series in natural language processing, have motivated efforts to replicate this success in visual tasks.…