7 citations · 17 across the 51 of their papers we have counts for
6 papers · 1 filter
Towards Compact 3D Representations via Point Feature Enhancement Masked Autoencoders
Yaohua Zha, Huizhen Ji, Jinmin Li +5
Learning 3D representation plays a critical role in masked autoencoder (MAE) based pre-training methods for point cloud, including single-modal and cross-modal based MAE. Specifica…
Parameter Efficient Adaptation for Image Restoration with Heterogeneous Mixture-of-Experts
Hang Guo, Tao Dai, Yuanchao Bai +4
Designing single-task image restoration models for specific degradation has seen great success in recent years. To achieve generalized image restoration, all-in-one methods have re…
Perceptual Image Compression with Cooperative Cross-Modal Side Information
Shiyu Qin, Bin Chen, Yujun Huang +3
The explosion of data has resulted in more and more associated text being transmitted along with images. Inspired by from distributed source coding, many works utilize image side i…
Progressive Learning with Visual Prompt Tuning for Variable-Rate Image Compression
Shiyu Qin, Yimin Zhou, Jinpeng Wang +4
In this paper, we propose a progressive learning paradigm for transformer-based variable-rate image compression. Our approach covers a wide range of compression rates with the assi…
Editable-DeepSC: Cross-Modal Editable Semantic Communication Systems
Wenbo Yu, Bin Chen, Qinshan Zhang +1
Different from data-oriented communication systems that primarily focus on how to accurately transmit every bit of data, task-oriented semantic communication systems only transmit…
GMMFormer: Gaussian-Mixture-Model Based Transformer for Efficient Partially Relevant Video Retrieval
Yuting Wang, Jinpeng Wang, Bin Chen +2
Given a text query, partially relevant video retrieval (PRVR) seeks to find untrimmed videos containing pertinent moments in a database. For PRVR, clip modeling is essential to cap…