collaborators

7 papers

cs.CV2026

Visual Token Codec: Unleashing Spatial Redundancy for ViT Feature Coding

Donghui Feng, Fengxi Zhang, Changsheng Gao +6

Distributed deployment of large vision foundation models often partitions a ViT backbone and exchanges intermediate token features between computing nodes, making efficient feature…

cs.CV2026

Content-Aware Mamba for Learned Image Compression

Yunuo Chen, Zezheng Lyu, Bing He +6

Recent learned image compression (LIC) leverages Mamba-style state-space models (SSMs) for global receptive fields with linear complexity. However, the standard Mamba adopts conten…

cs.LG2026

OmniZip: Learning a Unified and Lightweight Lossless Compressor for Multi-Modal Data

Yan Zhao, Zhengxue Cheng, Junxuan Zhang +4

Lossless compression is essential for efficient data storage and transmission. Although learning-based lossless compressors achieve strong results, most of them are designed for a…

cs.CV2026

Spatio-Temporal Token Pruning for Efficient High-Resolution GUI Agents

Zhou Xu, Bowen Zhou, Qi Wang +2

Pure-vision GUI agents provide universal interaction capabilities but suffer from severe efficiency bottlenecks due to the massive spatiotemporal redundancy inherent in high-resolu…

cs.CV2025

H3D-DGS: Exploring Heterogeneous 3D Motion Representation for Deformable 3D Gaussian Splatting

Bing He, Yunuo Chen, Guo Lu +5

Dynamic scene reconstruction poses a persistent challenge in 3D vision. Deformable 3D Gaussian Splatting has emerged as an effective method for this task, offering real-time render…

cs.CV2025

DualComp: End-to-End Learning of a Unified Dual-Modality Lossless Compressor

Yan Zhao, Zhengxue Cheng, Junxuan Zhang +3

Most learning-based lossless compressors are designed for a single modality, requiring separate models for multi-modal data and lacking flexibility. However, different modalities v…