5 papers
TAMFormer: Multi-Modal Transformer with Learned Attention Mask for Early Intent Prediction
Nada Osman, Guglielmo Camporese, Lamberto Ballan
Human intention prediction is a growing area of research where an activity in a video has to be anticipated by a vision-based system. To this end, the model creates a representatio…
SlowFast Rolling-Unrolling LSTMs for Action Anticipation in Egocentric Videos
Nada Osman, Guglielmo Camporese, Pasquale Coscia +1
Action anticipation in egocentric videos is a difficult task due to the inherently multi-modal nature of human actions. Additionally, some actions happen faster or slower than othe…
Conditional Variational Capsule Network for Open Set Recognition
Yunrui Guo, Guglielmo Camporese, Wenjing Yang +2
In open set recognition, a classifier has to detect unknown classes that are not known at training time. In order to recognize new categories, the classifier has to project the inp…
Improved Robustness to Disfluencies in RNN-Transducer Based Speech Recognition
Valentin Mendelev, Tina Raissi, Guglielmo Camporese +1
Automatic Speech Recognition (ASR) based on Recurrent Neural Network Transducers (RNN-T) is gaining interest in the speech community. We investigate data selection and preparation…
Knowledge Distillation for Action Anticipation via Label Smoothing
Guglielmo Camporese, Pasquale Coscia, Antonino Furnari +2
Human capability to anticipate near future from visual observations and non-verbal cues is essential for developing intelligent systems that need to interact with people. Several r…