5 papers
Visual Token Codec: Unleashing Spatial Redundancy for ViT Feature Coding
Donghui Feng, Fengxi Zhang, Changsheng Gao +6
Distributed deployment of large vision foundation models often partitions a ViT backbone and exchanges intermediate token features between computing nodes, making efficient feature…
Adaptive Learned Image Compression with Graph Neural Networks
Yunuo Chen, Bing He, Zezheng Lyu +4
Efficient image compression relies on modeling both local and global redundancy. Most state-of-the-art (SOTA) learned image compression (LIC) methods are based on CNNs or Transform…
OmniZip: Learning a Unified and Lightweight Lossless Compressor for Multi-Modal Data
Yan Zhao, Zhengxue Cheng, Junxuan Zhang +4
Lossless compression is essential for efficient data storage and transmission. Although learning-based lossless compressors achieve strong results, most of them are designed for a…
H3D-DGS: Exploring Heterogeneous 3D Motion Representation for Deformable 3D Gaussian Splatting
Bing He, Yunuo Chen, Guo Lu +5
Dynamic scene reconstruction poses a persistent challenge in 3D vision. Deformable 3D Gaussian Splatting has emerged as an effective method for this task, offering real-time render…
DualComp: End-to-End Learning of a Unified Dual-Modality Lossless Compressor
Yan Zhao, Zhengxue Cheng, Junxuan Zhang +3
Most learning-based lossless compressors are designed for a single modality, requiring separate models for multi-modal data and lacking flexibility. However, different modalities v…