4 papers
ResTok: Learning Hierarchical Residuals in 1D Visual Tokenizers for Autoregressive Image Generation
Xu Zhang, Cheng Da, Huan Yang +3
Existing 1D visual tokenizers for autoregressive (AR) generation largely follow the design principles of language modeling, as they are built directly upon transformers whose prior…
Perception-Oriented Latent Coding for High-Performance Compressed Domain Semantic Inference
Xu Zhang, Ming Lu, Yan Chen +1
In recent years, compressed domain semantic inference has primarily relied on learned image coding models optimized for mean squared error (MSE). However, MSE-oriented optimization…
Ultra Lowrate Image Compression with Semantic Residual Coding and Compression-aware Diffusion
Anle Ke, Xu Zhang, Tong Chen +4
Existing multimodal large model-based image compression frameworks often rely on a fragmented integration of semantic retrieval, latent compression, and generative models, resultin…
All-in-One Image Coding for Joint Human-Machine Vision with Multi-Path Aggregation
Xu Zhang, Peiyao Guo, Ming Lu +1
Image coding for multi-task applications, catering to both human perception and machine vision, has been extensively investigated. Existing methods often rely on multiple task-spec…