From the 1 of 4 linked papers with an AI index.
4 papers
VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders
Zhihao Xie, Junfeng Wu, Xinting Hu +2
VideoRAE is a representation autoencoder that leverages frozen video foundation model features to create compact, generation‑friendly video latents, supporting both continuous diff…
UniTok: A Unified Tokenizer for Visual Generation and Understanding
Chuofan Ma, Yi Jiang, Junfeng Wu +5
Visual generative and understanding models typically rely on distinct tokenizers to process images, presenting a key challenge for unifying them within a single framework. Recent s…
TokBench: Evaluating Your Visual Tokenizer before Visual Generation
Junfeng Wu, Dongliang Luo, Weizhi Zhao +6
In this work, we reveal the limitations of visual tokenizers and VAEs in preserving fine-grained features, and propose a benchmark to evaluate reconstruction performance for two ch…
Liquid: Language Models are Scalable and Unified Multi-modal Generators
Junfeng Wu, Yi Jiang, Chuofan Ma +5
We present Liquid, an auto-regressive generation paradigm that seamlessly integrates visual comprehension and generation by tokenizing images into discrete codes and learning these…