9 papers
Just Noticeable Difference Modeling for Token Compression in Vision-Language-Action Models
Zhuoyuan Li, Rui Zhao, Jin Wang +5
Token compression has become a key technique for reducing the inference cost of large foundation models, with approaches such as token pruning and KV-cache reuse widely adopted in…
Plug-and-play endoscopic OCT enabled by a fully fiber-integrated 3D-printed probe
Yihan Wang, Ruilin You, Jiabin Chen +5
Endoscopic optical coherence tomography (OCT) brings depth-resolved imaging into small lumens and confined spaces, but each probe must be assembled from spliced, cleaved, and align…
Current Injection Spiking Neural Network for Infrared and Visible Image Fusion
Rui Zhao, Zhuoyuan Li, Wenrui Li +4
Infrared and visible image fusion (IVIF) integrates the complementary information of two modalities into a single image with richer scene content. While existing methods are largel…
RAVE: Rate-Adaptive Visual Encoding for 3D Gaussian Splatting
Hoang-Nhat Tran, Francesco Di Sario, Gabriele Spadaro +2
Recent advances in neural scene representations have transformed immersive multimedia, with 3D Gaussian Splatting (3DGS) enabling real-time photorealistic rendering. Despite its ef…
Generative Models at the Frontier of Compression: A Survey on Generative Face Video Coding
Bolin Chen, Shanzhi Yin, Goluck Konuko +4
The rise of deep generative models has greatly advanced video compression, reshaping the paradigm of face video coding through their powerful capability for semantic-aware represen…
Denoising Diffusion Probabilistic Model for Point Cloud Compression at Low Bit-Rates
Gabriele Spadaro, Alberto Presta, Jhony H. Giraldo +5
Efficient compression of low-bit-rate point clouds is critical for bandwidth-constrained applications. However, existing techniques mainly focus on high-fidelity reconstruction, re…