1 citations · 2 across the 12 of their papers we have counts for
7 papers · 1 filter
ResTok: Learning Hierarchical Residuals in 1D Visual Tokenizers for Autoregressive Image Generation
Xu Zhang, Cheng Da, Huan Yang +3
Existing 1D visual tokenizers for autoregressive (AR) generation largely follow the design principles of language modeling, as they are built directly upon transformers whose prior…
Neural B-frame Video Compression with Bi-directional Reference Harmonization
Yuxi Liu, Dengchao Jin, Shuai Huo +5
Neural video compression (NVC) has made significant progress in recent years, while neural B-frame video compression (NBVC) remains underexplored compared to P-frame compression. N…
Perception-Oriented Latent Coding for High-Performance Compressed Domain Semantic Inference
Xu Zhang, Ming Lu, Yan Chen +1
In recent years, compressed domain semantic inference has primarily relied on learned image coding models optimized for mean squared error (MSE). However, MSE-oriented optimization…
Ultra Lowrate Image Compression with Semantic Residual Coding and Compression-aware Diffusion
Anle Ke, Xu Zhang, Tong Chen +4
Existing multimodal large model-based image compression frameworks often rely on a fragmented integration of semantic retrieval, latent compression, and generative models, resultin…
DreamInsert: Zero-Shot Image-to-Video Object Insertion from A Single Image
Qi Zhao, Zhan Ma, Pan Zhou
Recent developments in generative diffusion models have turned many dreams into realities. For video object insertion, existing methods typically require additional information, su…
Towards Loss-Resilient Image Coding for Unstable Satellite Networks
Hongwei Sha, Muchen Dong, Quanyou Luo +3
Geostationary Earth Orbit (GEO) satellite communication demonstrates significant advantages in emergency short burst data services. However, unstable satellite networks, particular…