papers

Publications (24)

cs.CV2024

Enhancing the Rate-Distortion-Perception Flexibility of Learned Image Codecs with Conditional Diffusion Decoders

Daniele Mari, Simone Milani

Learned image compression codecs have recently achieved impressive compression performances surpassing the most efficient image coding architectures. However, most approaches are t…

cs.SD2022

The Sound of Silence: Efficiency of First Digit Features in Synthetic Audio Detection

Daniele Mari, Federica Latora, Simone Milani

The recent integration of generative neural strategies and audio processing techniques have fostered the widespread of synthetic speech synthesis or transformation algorithms. This…

eess.SP2019

Seq2Seq RNN based Gait Anomaly Detection from Smartphone Acquired Multimodal Motion Data

Riccardo Bonetto, Mattia Soldan, Alberto Lanaro +2

Smartphones and wearable devices are fast growing technologies that, in conjunction with advances in wireless sensor hardware, are enabling ubiquitous sensing applications. Wearabl…

cs.MM2026

Code Division Modulation Layers Against Forgetting and Inference in Continual Gait Identification

Simone Milani

Continual learning (CL) has been recently employed in biometric identification systems thanks to its ability to integrate new knowledge within a pre-trained model and to the possib…

cs.CV2026

Traceback Translators Against Forgetting in Continual Fake Speech Detection

Enrico Gottardis, Mattia Tamiazzo, Simone Milani

The paper proposes a method that uses a domain‑translator network to map new fake‑speech data back into the feature space of an existing detector, allowing continual learning witho…

#continual learning#fake speech detection#catastrophic forgetting#domain translation
eess.IV2024

Effectiveness of learning-based image codecs on fingerprint storage

Daniele Mari, Saverio Cavasin, Simone Milani +1

The success of learning-based coding techniques and the development of learning-based image coding standards, such as JPEG-AI, point towards the adoption of such solutions in diffe…

cs.GR2025

SAGE: Semantic-Driven Adaptive Gaussian Splatting in Extended Reality

Chiara Schiavo, Elena Camuffo, Leonardo Badia +1

3D Gaussian Splatting (3DGS) has significantly improved the efficiency and realism of three-dimensional scene visualization in several applications, ranging from robotics to eXtend…

cs.GR2026

Split&Splat: Zero-Shot Panoptic Segmentation via Explicit Instance Modeling and 3D Gaussian Splatting

Leonardo Monchieri, Elena Camuffo, Francesco Barbato +2

3D Gaussian Splatting (GS) enables fast and high-quality scene reconstruction, but it lacks an object-consistent and semantically aware structure. We propose Split&Splat, a framewo…

cs.CV2025

Point Cloud Geometry Scalable Coding Using a Resolution and Quality-conditioned Latents Probability Estimator

Daniele Mari, André F. R. Guarda, Nuno M. M. Rodrigues +2

In the current age, users consume multimedia content in very heterogeneous scenarios in terms of network, hardware, and display capabilities. A naive solution to this problem is to…

cs.SD2026

Explainable-by-Design Audio Deepfake Detection via Wiener-Hopf Linear Prediction

Mattia Tamiazzo, Simone Milani, Massimo Iuliani +1

The paper introduces a lightweight, explainable audio deepfake detector that uses Wiener‑Hopf linear prediction combined with a 2D CNN, achieving competitive accuracy with lower co…

#audio deepfake detection#explainable AI#linear prediction#lightweight CNN
cs.CV2026

MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment

Elena Camuffo, Francesco Barbato, Mete Ozay +2

Personalized object detection aims to adapt a general-purpose detector to recognize user-specific instances from only a few examples. Lightweight models often struggle in this sett…

cs.CV2024

Fingerprint Membership and Identity Inference Against Generative Adversarial Networks

Saverio Cavasin, Daniele Mari, Simone Milani +1

Generative models are gaining significant attention as potential catalysts for a novel industrial revolution. Since automated sample generation can be useful to solve privacy and d…

cs.CV2025

TeLL Me what you cant see

Saverio Cavasin, Pietro Biasetton, Mattia Tamiazzo +2

During criminal investigations, images of persons of interest directly influence the success of identification procedures. However, law enforcement agencies often face challenges r…

cs.CV2022

Real or Virtual: A Video Conferencing Background Manipulation-Detection System

Ehsan Nowroozi, Yassine Mekdad, Mauro Conti +3

Recently, the popularity and wide use of the last-generation video conferencing technologies created an exponential growth in its market size. Such technology allows participants i…

cs.LG2024

Point Cloud Geometry Scalable Coding with a Quality-Conditioned Latents Probability Estimator

Daniele Mari, André F. R. Guarda, Nuno M. M. Rodrigues +2

The widespread usage of point clouds (PC) for immersive visual applications has resulted in the use of very heterogeneous receiving conditions and devices, notably in terms of netw…

cs.CV2023

Continual Learning for LiDAR Semantic Segmentation: Class-Incremental and Coarse-to-Fine strategies on Sparse Data

Elena Camuffo, Simone Milani

During the last few years, continual learning (CL) strategies for image classification and segmentation have been widely investigated designing innovative solutions to tackle catas…

cs.CR2021

Do Not Deceive Your Employer with a Virtual Background: A Video Conferencing Manipulation-Detection System

Mauro Conti, Simone Milani, Ehsan Nowroozi +1

The last-generation video conferencing software allows users to utilize a virtual background to conceal their personal environment due to privacy concerns, especially in official m…

cs.CV2020

On the use of Benford's law to detect GAN-generated images

Nicolò Bonettini, Paolo Bestagini, Simone Milani +1

The advent of Generative Adversarial Network (GAN) architectures has given anyone the ability of generating incredibly realistic synthetic imagery. The malicious diffusion of GAN-g…

cs.CR2021

Hand Me Your PIN! Inferring ATM PINs of Users Typing with a Covered Hand

Matteo Cardaioli, Stefano Cecconello, Mauro Conti +3

Automated Teller Machines (ATMs) represent the most used system for withdrawing cash. The European Central Bank reported more than 11 billion cash withdrawals and loading/unloading…

cs.CV2020

FOCAL: A Forgery Localization Framework based on Video Coding Self-Consistency

Sebastiano Verde, Paolo Bestagini, Simone Milani +2

Forgery operations on video contents are nowadays within the reach of anyone, thanks to the availability of powerful and user-friendly editing software. Integrity verification and…

cs.CV2024

Enhanced Model Robustness to Input Corruptions by Per-corruption Adaptation of Normalization Statistics

Elena Camuffo, Umberto Michieli, Simone Milani +2

Developing a reliable vision system is a fundamental challenge for robotic technologies (e.g., indoor service robots and outdoor autonomous robots) which can ensure reliable naviga…

cs.CV2024

Continual Road-Scene Semantic Segmentation via Feature-Aligned Symmetric Multi-Modal Network

Francesco Barbato, Elena Camuffo, Simone Milani +1

State-of-the-art multimodal semantic segmentation strategies combining LiDAR and color data are usually designed on top of asymmetric information-sharing schemes and assume that bo…

cs.CV2023

Learning from Mistakes: Self-Regularizing Hierarchical Representations in Point Cloud Semantic Segmentation

Elena Camuffo, Umberto Michieli, Simone Milani

Recent advances in autonomous robotic technologies have highlighted the growing need for precise environmental analysis. LiDAR semantic segmentation has gained attention to accompl…

cs.SD2023

All-for-One and One-For-All: Deep learning-based feature fusion for Synthetic Speech Detection

Daniele Mari, Davide Salvi, Paolo Bestagini +1

Recent advances in deep learning and computer vision have made the synthesis and counterfeiting of multimedia content more accessible than ever, leading to possible threats and dan…