6 papers · 1 filter
TORINO: Token Reduction via Interpretable Concept Overlap in Vision-Language Models
Riccardo Renzulli, Gabriele Spadaro, Shruthi Gowda +2
Vision-Language Models (VLMs) have demonstrated impressive capabilities across different tasks, but their computational cost is dominated by the large number of visual tokens fed t…
RAVE: Rate-Adaptive Visual Encoding for 3D Gaussian Splatting
Hoang-Nhat Tran, Francesco Di Sario, Gabriele Spadaro +2
Recent advances in neural scene representations have transformed immersive multimedia, with 3D Gaussian Splatting (3DGS) enabling real-time photorealistic rendering. Despite its ef…
Denoising Diffusion Probabilistic Model for Point Cloud Compression at Low Bit-Rates
Gabriele Spadaro, Alberto Presta, Jhony H. Giraldo +5
Efficient compression of low-bit-rate point clouds is critical for bandwidth-constrained applications. However, existing techniques mainly focus on high-fidelity reconstruction, re…
FOLDER: Accelerating Multi-modal Large Language Models with Enhanced Performance
Haicheng Wang, Zhemeng Yu, Gabriele Spadaro +4
Recently, Multi-modal Large Language Models (MLLMs) have shown remarkable effectiveness for multi-modal tasks due to their abilities to generate and understand cross-modal data. Ho…
WiGNet: Windowed Vision Graph Neural Network
Gabriele Spadaro, Marco Grangetto, Attilio Fiandrotti +2
In recent years, Graph Neural Networks (GNNs) have demonstrated strong adaptability to various real-world challenges, with architectures such as Vision GNN (ViG) achieving state-of…
Domain Adaptation for Learned Image Compression with Supervised Adapters
Alberto Presta, Gabriele Spadaro, Enzo Tartaglione +2
In Learned Image Compression (LIC), a model is trained at encoding and decoding images sampled from a source domain, often outperforming traditional codecs on natural images; yet i…