6 papers · 1 filter
Just Noticeable Difference Modeling for Token Compression in Vision-Language-Action Models
Zhuoyuan Li, Rui Zhao, Jin Wang +5
Token compression has become a key technique for reducing the inference cost of large foundation models, with approaches such as token pruning and KV-cache reuse widely adopted in…
Current Injection Spiking Neural Network for Infrared and Visible Image Fusion
Rui Zhao, Zhuoyuan Li, Wenrui Li +4
Infrared and visible image fusion (IVIF) integrates the complementary information of two modalities into a single image with richer scene content. While existing methods are largel…
RAVE: Rate-Adaptive Visual Encoding for 3D Gaussian Splatting
Hoang-Nhat Tran, Francesco Di Sario, Gabriele Spadaro +2
Recent advances in neural scene representations have transformed immersive multimedia, with 3D Gaussian Splatting (3DGS) enabling real-time photorealistic rendering. Despite its ef…
Generative Models at the Frontier of Compression: A Survey on Generative Face Video Coding
Bolin Chen, Shanzhi Yin, Goluck Konuko +4
The rise of deep generative models has greatly advanced video compression, reshaping the paradigm of face video coding through their powerful capability for semantic-aware represen…
Denoising Diffusion Probabilistic Model for Point Cloud Compression at Low Bit-Rates
Gabriele Spadaro, Alberto Presta, Jhony H. Giraldo +5
Efficient compression of low-bit-rate point clouds is critical for bandwidth-constrained applications. However, existing techniques mainly focus on high-fidelity reconstruction, re…
Language-Guided Visual Perception Disentanglement for Image Quality Assessment and Conditional Image Generation
Zhichao Yang, Leida Li, Pengfei Chen +2
Contrastive vision-language models, such as CLIP, have demonstrated excellent zero-shot capability across semantic recognition tasks, mainly attributed to the training on a large-s…