2 papers
cs.CV2025
SFTok: Bridging the Performance Gap in Discrete Tokenizers
Qihang Rao, Borui Zhang, Wenzhao Zheng +2
Recent advances in multimodal models highlight the pivotal role of image tokenization in high-resolution image generation. By compressing images into compact latent representations…
cs.CV2025
Quantize-then-Rectify: Efficient VQ-VAE Training
Borui Zhang, Qihang Rao, Wenzhao Zheng +2
Visual tokenizers are pivotal in multimodal large models, acting as bridges between continuous inputs and discrete tokens. Nevertheless, training high-compression-rate VQ-VAEs rema…