Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
Autoregressive Video Generation beyond Next Frames Prediction
Sucheng Ren, Chen Chen, Zhenbang Wang +5
Autoregressive models for video generation typically operate frame-by-frame, extending next-token prediction from language to video's temporal dimension. We question that unlike wo…
cs.CV2025
MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
Yanghao Li, Rui Qian, Bowen Pan +24
Unified multimodal Large Language Models (LLMs) that can both understand and generate visual content hold immense potential. However, existing open-source models often suffer from…
cs.CV2025
AToken: A Unified Tokenizer for Vision
Jiasen Lu, Liangchen Song, Mingze Xu +5
We present AToken, the first unified visual tokenizer that achieves both high-fidelity reconstruction and semantic understanding across images, videos, and 3D assets. Unlike existi…