4 papers
Efficient Audio-Visual Event Recognition via Knowledge Distillation and Dynamic INT8 Quantization of a Hybrid Cross-Attention Network
Parinaz Binandeh Dehaghani, Danilo Pena, A. Pedro Aguiar
Audio-visual event recognition (AVER) has achieved significant performance improvements through transformer-based multimodal architectures. However, the high computational complexi…
Stable Hybrid Cross-Attention Fusion for Audio-Visual Event Recognition
Parinaz Binandeh Dehaghani, Danilo Pena, A. Pedro Aguiar
Audio-Visual Event Recognition (AVER) is essential for intelligent urban monitoring systems, where robust multimodal understanding of complex environments is required. This paper p…
Evaluating CNN with Stacked Feature Representations and Audio Spectrogram Transformer Models for Sound Classification
Parinaz Binandeh Dehaghania, Danilo Penab, A. Pedro Aguiar
Environmental sound classification (ESC) has gained significant attention due to its diverse applications in smart city monitoring, fault detection, acoustic surveillance, and manu…
Investigation of Feature Selection and Pooling Methods for Environmental Sound Classification
Parinaz Binandeh Dehaghani, Danilo Pena, A. Pedro Aguiar
This paper explores the impact of dimensionality reduction and pooling methods for Environmental Sound Classification (ESC) using lightweight CNNs. We evaluate Sparse Salient Regio…