papers

Publications (68)

cs.CV2014

Robust Temporally Coherent Laplacian Protrusion Segmentation of 3D Articulated Bodies

Fabio Cuzzolin, Diana Mateus, Radu Horaud

In motion analysis and understanding it is important to be able to fit a suitable model or structure to the temporal series of observed data, in order to describe motion patterns i…

cs.AI2026

A formal definition and meta-model for a machine theory of mind

Fabio Cuzzolin

This paper proposes, for the first time, a rigorous formal definition of the concept of Machine Theory of Mind, based on principles supported by evidence from cognitive psychology,…

cs.RO2021

Unsupervised anomaly detection for a Smart Autonomous Robotic Assistant Surgeon (SARAS)using a deep residual autoencoder

Dinesh Jackson Samuel, Fabio Cuzzolin

Anomaly detection in Minimally-Invasive Surgery (MIS) traditionally requires a human expert monitoring the procedure from a console. Data scarcity, on the other hand, hinders what…

cs.AI2021

A geometric approach to conditioning belief functions

Fabio Cuzzolin

Conditioning is crucial in applied science when inference involving time series is involved. Belief calculus is an effective way of handling such inference in the presence of epist…

cs.LG2026

Learning Credal Ensembles via Distributionally Robust Optimization

Kaizheng Wang, Ghifari Adam Faza, Fabio Cuzzolin +3

Credal predictors are models that are aware of epistemic uncertainty and produce a convex set of probabilistic predictions. They offer a principled way to quantify predictive epist…

eess.IV2018

TraMNet - Transition Matrix Network for Efficient Action Tube Proposals

Gurkirt Singh, Suman Saha, Fabio Cuzzolin

Current state-of-the-art methods solve spatiotemporal action localisation by extending 2D anchors to 3D-cuboid proposals on stacks of frames, to generate sets of temporally connect…

cs.CV2022

Situation Awareness for Automated Surgical Check-listing in AI-Assisted Operating Room

Tochukwu Onyeogulu, Salman Khan, Izzeddin Teeti +6

Nowadays, there are more surgical procedures that are being performed using minimally invasive surgery (MIS). This is due to its many benefits, such as minimal post-operative probl…

cs.LG2022

Epistemic Deep Learning

Shireen Kudukkil Manchingal, Fabio Cuzzolin

The belief function approach to uncertainty quantification as proposed in the Demspter-Shafer theory of evidence is established upon the general mathematical models for set-valued…

cs.CV2020

Two-Stream AMTnet for Action Detection

Suman Saha, Gurkirt Singh, Fabio Cuzzolin

In this paper, we propose Two-Stream AMTnet, which leverages recent advances in video-based action representation[1] and incremental action tube generation[2]. Majority of the pres…

cs.CV2020

Articulated Shape Matching Using Laplacian Eigenfunctions and Unsupervised Point Registration

Diana Mateus, Radu Horaud, David Knossow +2

Matching articulated shapes represented by voxel-sets reduces to maximal sub-graph isomorphism when each set is described by a weighted graph. Spectral graph theory can be used to…

cs.LG2025

Generalized Decision Focused Learning under Imprecise Uncertainty--Theoretical Study

Keivan Shariatmadar, Neil Yorke-Smith, Ahmad Osman +3

Decision Focused Learning has emerged as a critical paradigm for integrating machine learning with downstream optimisation. Despite its promise, existing methodologies predominantl…

cs.CV2026

ROAD-Waymo: A Large-Scale Action Awareness Dataset for Autonomous Driving

Salman Khan, Izzeddin Teeti, Reza Javanmard Alitappeh +5

Autonomous Vehicle (AV) perception systems require more than simply seeing, via e.g., object detection or scene segmentation. They need a holistic understanding of what is happenin…

cs.CV2024

Feature boosting with efficient attention for scene parsing

Vivek Singh, Shailza Sharma, Fabio Cuzzolin

The complexity of scene parsing grows with the number of object and scene classes, which is higher in unrestricted open scenes. The biggest challenge is to model the spatial relati…

cs.LG2025

Credal Wrapper of Model Averaging for Uncertainty Estimation in Classification

Kaizheng Wang, Fabio Cuzzolin, Keivan Shariatmadar +2

This paper presents an innovative approach, called credal wrapper, to formulating a credal set representation of model averaging for Bayesian neural networks (BNNs) and deep ensemb…

cs.CV2022

Spatiotemporal Deformable Scene Graphs for Complex Activity Detection

Salman Khan, Fabio Cuzzolin

Long-term complex activity recognition and localisation can be crucial for decision making in autonomous systems such as smart cars and surgical robots. Here we address the problem…

cs.LG2026

Epistemic Generative Adversarial Networks

Muhammad Mubashar, Fabio Cuzzolin

Generative models, particularly Generative Adversarial Networks (GANs), often suffer from a lack of output diversity, frequently generating similar samples rather than a wide range…

cs.CV2019

Recurrent Convolutions for Causal 3D CNNs

Gurkirt Singh, Fabio Cuzzolin

Recently, three dimensional (3D) convolutional neural networks (CNNs) have emerged as dominant methods to capture spatiotemporal representations in videos, by adding to pre-existin…

math.ST2021

Uncertainty measures: The big picture

Fabio Cuzzolin

Probability theory is far from being the most general mathematical theory of uncertainty. A number of arguments point at its inability to describe second-order ('Knightian') uncert…

math.ST2026

Statistical inference with belief functions: A survey

Fabio Cuzzolin

Belief functions are a powerful and popular framework for the mathematical characterisation of uncertainty, in particular in situations in which lack of data renders learning a pro…

cs.CV2017

Spatio-temporal Human Action Localisation and Instance Segmentation in Temporally Untrimmed Videos

Suman Saha, Gurkirt Singh, Michael Sapienza +2

Current state-of-the-art human action recognition is focused on the classification of temporally trimmed videos in which only one action occurs per frame. In this work we address t…

cs.CV2019

End-to-End Video Captioning

Silvio Olivastri, Gurkirt Singh, Fabio Cuzzolin

Building correspondences across different modalities, such as video and language, has recently become critical in many visual recognition applications, such as video captioning. In…

cs.LG2026

Direct Interval Propagation Methods using Neural-Network Surrogates for Uncertainty Quantification in Physical Systems Surrogate Model

Ghifari Adam Faza, Jolan Wauters, Fabio Cuzzolin +2

In engineering, uncertainty propagation aims to characterise system outputs under uncertain inputs. For interval uncertainty, the goal is to determine output bounds given interval-…

cs.LG2025

Random-Set Neural Networks (RS-NN)

Shireen Kudukkil Manchingal, Muhammad Mubashar, Kaizheng Wang +2

Machine learning is increasingly deployed in safety-critical domains where erroneous predictions may lead to potentially catastrophic consequences, highlighting the need for learni…

cs.CV2018

Action Detection from a Robot-Car Perspective

Valentina Fontana, Gurkirt Singh, Stephen Akrigg +3

We present the new Road Event and Activity Detection (READ) dataset, designed and created from an autonomous vehicle perspective to take action detection challenges to autonomous d…

cs.CV2026

A neurosymbolic Approach with Epistemic Deep Learning for Hierarchical Image Classification

Ezel Kilicdere, Shireen Kudukkil Manchingal, Fabio Cuzzolin

Deep neural networks achieve high accuracy on image classification tasks. Yet, they often produce overconfident predictions as which fail to express epistemic uncertainty, and freq…

cs.CV2017

AMTnet: Action-Micro-Tube Regression by End-to-end Trainable Deep Architecture

Suman Saha, Gurkirt Singh, Fabio Cuzzolin

Dominant approaches to action detection can only provide sub-optimal solutions to the problem, as they rely on seeking frame-level detections, to later compose them into "action tu…

cs.CV2021

The SARAS Endoscopic Surgeon Action Detection (ESAD) dataset: Challenges and methods

Vivek Singh Bawa, Gurkirt Singh, Francis KapingA +16

For an autonomous robotic system, monitoring surgeon actions and assisting the main surgeon during a procedure can be very challenging. The challenges come from the peculiar struct…

cs.CV2018

Incremental Tube Construction for Human Action Detection

Harkirat Singh Behl, Michael Sapienza, Gurkirt Singh +3

Current state-of-the-art action detection systems are tailored for offline batch-processing applications. However, for online applications like human-robot interaction, current sys…

cs.AI2025

Epistemic Artificial Intelligence is Essential for Machine Learning Models to Truly 'Know When They Do Not Know'

Shireen Kudukkil Manchingal, Andrew Bradley, Julian F. P. Kooij +3

Despite AI's impressive achievements, including recent advances in generative and large language models, there remains a significant gap in the ability of AI systems to handle unce…

cs.LG2025

CreINNs: Credal-Set Interval Neural Networks for Uncertainty Estimation in Classification Tasks

Kaizheng Wang, Keivan Shariatmadar, Shireen Kudukkil Manchingal +3

Effective uncertainty estimation is becoming increasingly attractive for enhancing the reliability of neural networks. This work presents a novel approach, termed Credal-Set Interv…

cs.LG2025

A Unified Evaluation Framework for Epistemic Predictions

Shireen Kudukkil Manchingal, Muhammad Mubashar, Kaizheng Wang +1

Predictions of uncertainty-aware models are diverse, ranging from single point estimates (often averaged over prediction samples) to predictive distributions, to set-valued or cred…

cs.CV2023

A Hybrid Graph Network for Complex Activity Detection in Video

Salman Khan, Izzeddin Teeti, Andrew Bradley +2

Interpretation and understanding of video presents a challenging computer vision task in numerous fields - e.g. autonomous driving and sports analytics. Existing approaches to inte…

cs.CV2017

Online Real-time Multiple Spatiotemporal Action Localisation and Prediction

Gurkirt Singh, Suman Saha, Michael Sapienza +2

We present a deep-learning framework for real-time multiple spatio-temporal (S/T) action localisation, classification and early prediction. Current state-of-the-art approaches work…

cs.LG2025

Epistemic Wrapping for Uncertainty Quantification

Maryam Sultana, Neil Yorke-Smith, Kaizheng Wang +3

Uncertainty estimation is pivotal in machine learning, especially for classification tasks, as it improves the robustness and reliability of models. We introduce a novel `Epistemic…

cs.LG2022

ROAD-R: The Autonomous Driving Dataset with Logical Requirements

Eleonora Giunchiglia, Mihaela Cătălina Stoian, Salman Khan +2

Neural networks have proven to be very powerful at computer vision tasks. However, they often exhibit unexpected behaviours, violating known requirements expressing background know…

cs.CV2018

Visions of a generalized probability theory

Fabio Cuzzolin

In this Book we argue that the fruitful interaction of computer vision and belief calculus is capable of stimulating significant advances in both fields. From a methodological poin…

cs.RO2025

Uncertainty-Aware Autonomous Vehicles: Predicting the Road Ahead

Shireen Kudukkil Manchingal, Armand Amaritei, Mihir Gohad +4

Autonomous Vehicle (AV) perception systems have advanced rapidly in recent years, providing vehicles with the ability to accurately interpret their environment. Perception systems…

cs.CV2020

ESAD: Endoscopic Surgeon Action Detection Dataset

Vivek Singh Bawa, Gurkirt Singh, Francis KapingA +8

In this work, we take aim towards increasing the effectiveness of surgical assistant robots. We intended to make assistant robots safer by making them aware about the actions of su…

cs.AI2014

Consistent transformations of belief functions

Fabio Cuzzolin

Consistent belief functions represent collections of coherent or non-contradictory pieces of evidence, but most of all they are the counterparts of consistent knowledge bases in be…

math.ST2018

Belief likelihood function for generalised logistic regression

Fabio Cuzzolin

The notion of belief likelihood function of repeated trials is introduced, whenever the uncertainty for individual trials is encoded by a belief measure (a finite random set). This…

cs.CV2016

Untrimmed Video Classification for Activity Detection: submission to ActivityNet Challenge

Gurkirt Singh, Fabio Cuzzolin

Current state-of-the-art human activity recognition is focused on the classification of temporally trimmed videos in which only one action occurs per frame. We propose a simple, ye…

cs.LG2025

Credal and Interval Deep Evidential Classifications

Michele Caprio, Shireen K. Manchingal, Fabio Cuzzolin

Uncertainty Quantification (UQ) presents a pivotal challenge in the field of Artificial Intelligence (AI), profoundly impacting decision-making, risk assessment and model reliabili…

cs.LG2024

Generalising realisability in statistical learning theory under epistemic uncertainty

Fabio Cuzzolin

The purpose of this paper is to look into how central notions in statistical learning theory, such as realisability, generalise under the assumption that train and test distributio…

cs.CV2023

Vision in adverse weather: Augmentation using CycleGANs with various object detectors for robust perception in autonomous racing

Izzeddin Teeti, Valentina Musat, Salman Khan +3

In an autonomous driving system, perception - identification of features and objects from the environment - is crucial. In autonomous racing, high speeds and small margins demand r…

cs.LG2022

Identification of Cognitive Workload during Surgical Tasks with Multimodal Deep Learning

Kaizhe Jin, Adrian Rubio-Solis, Ravi Naik +8

The operating room (OR) is a dynamic and complex environment consisting of a multidisciplinary team working together in a high take environment to provide safe and efficient patien…

cs.CV2019

Spatio-Temporal Action Localization in a Weakly Supervised Setting

Kurt Degiorgio, Fabio Cuzzolin

Enabling computational systems with the ability to localize actions in video-based content has manifold applications. Traditionally, such a problem is approached in a fully-supervi…

cs.CV2016

Deep Learning for Detecting Multiple Space-Time Action Tubes in Videos

Suman Saha, Gurkirt Singh, Michael Sapienza +2

In this work, we propose an approach to the spatiotemporal localisation (detection) and classification of multiple concurrent actions within temporally untrimmed videos. Our framew…

cs.LG2024

Anomaly detection using Diffusion-based methods

Aryan Bhosale, Samrat Mukherjee, Biplab Banerjee +1

This paper explores the utility of diffusion-based models for anomaly detection, focusing on their efficacy in identifying deviations in both compact and high-resolution datasets.…

cs.CV2023

YOLO-Z: Improving small object detection in YOLOv5 for autonomous vehicles

Aduen Benjumea, Izzeddin Teeti, Fabio Cuzzolin +1

As autonomous vehicles and autonomous racing rise in popularity, so does the need for faster and more accurate detectors. While our naked eyes are able to extract contextual inform…

cs.CV2021

International Workshop on Continual Semi-Supervised Learning: Introduction, Benchmarks and Baselines

Ajmal Shahbaz, Salman Khan, Mohammad Asiful Hossain +4

The aim of this paper is to formalize a new continual semi-supervised learning (CSSL) paradigm, proposed to the attention of the machine learning community via the IJCAI 2021 Inter…

cs.AI2025

Proceedings of 1st Workshop on Advancing Artificial Intelligence through Theory of Mind

Mouad Abrini, Omri Abend, Dina Acklin +105

This volume includes a selection of papers presented at the Workshop on Advancing Artificial Intelligence through Theory of Mind held at AAAI 2025 in Philadelphia US on 3rd March 2…

cs.CV2023

Temporal DINO: A Self-supervised Video Strategy to Enhance Action Prediction

Izzeddin Teeti, Rongali Sai Bhargav, Vivek Singh +3

The emerging field of action prediction plays a vital role in various computer vision applications such as autonomous driving, activity analysis and human-computer interaction. Des…

cs.CV2020

Challenges and Opportunities for Computer Vision in Real-life Soccer Analytics

Neha Bhargava, Fabio Cuzzolin

In this paper, we explore some of the applications of computer vision to sports analytics. Sport analytics deals with understanding and discovering patterns from a corpus of sports…

cs.CV2014

Feature sampling and partitioning for visual vocabulary generation on large action classification datasets

Michael Sapienza, Fabio Cuzzolin, Philip H. S. Torr

The recent trend in action recognition is towards larger datasets, an increasing number of action classes and larger visual vocabularies. State-of-the-art human action classificati…

math.ST2023

Reasoning with random sets: An agenda for the future

Fabio Cuzzolin

In this paper, we discuss a potential agenda for future work in the theory of random sets and belief functions, touching upon a number of focal issues: the development of a fully-f…

cs.LG2021

Avalanche: an End-to-End Library for Continual Learning

Vincenzo Lomonaco, Lorenzo Pellegrini, Andrea Cossu +25

Learning continually from non-stationary data streams is a long-standing goal and a challenging problem in machine learning. Recently, we have witnessed a renewed and fast-growing…

stat.ML2016

Active Learning for Online Recognition of Human Activities from Streaming Videos

Rocco De Rosa, Ilaria Gori, Fabio Cuzzolin +2

Recognising human activities from streaming videos poses unique challenges to learning algorithms: predictive models need to be scalable, incrementally trainable, and must remain b…

cs.AI2022

The intersection probability: betting with probability intervals

Fabio Cuzzolin

Probability intervals are an attractive tool for reasoning under uncertainty. Unlike belief functions, though, they lack a natural probability transformation to be used for decisio…

cs.CV2025

ASTRA: A Scene-aware TRAnsformer-based model for trajectory prediction

Izzeddin Teeti, Aniket Thomas, Munish Monga +5

We present ASTRA (A} Scene-aware TRAnsformer-based model for trajectory prediction), a light-weight pedestrian trajectory forecasting model that integrates the scene context, spati…

cs.LG2026

Set-based v.s. Distribution-based Representations of Epistemic Uncertainty: A Comparative Study

Kaizheng Wang, Yunjia Wang, Fabio Cuzzolin +3

Epistemic uncertainty in neural networks is commonly modeled using two second-order paradigms: distribution-based representations, which rely on posterior parameter distributions,…

cs.LG2024

Credal Learning Theory

Michele Caprio, Maryam Sultana, Eleni Elia +1

Statistical learning theory is the foundation of machine learning, providing theoretical bounds for the risk of models learned from a (single) training set, assumed to issue from a…

cs.CL2025

Random-Set Large Language Models

Muhammad Mubashar, Shireen Kudukkil Manchingal, Fabio Cuzzolin

Large Language Models (LLMs) are known to produce very high-quality tests and responses to our queries. But how much can we trust this generated text? In this paper, we study the p…

cs.SE2019

Datamorphic Testing: A Methodology for Testing AI Applications

Hong Zhu, Dongmei Liu, Ian Bayley +2

With the rapid growth of the applications of machine learning (ML) and other artificial intelligence (AI) techniques, adequate testing has become a necessity to ensure their qualit…

cs.LG2024

Deep evolving semi-supervised anomaly detection

Jack Belham, Aryan Bhosale, Samrat Mukherjee +2

The aim of this paper is to formalise the task of continual semi-supervised anomaly detection (CSAD), with the aim of highlighting the importance of such a problem formulation whic…

cs.LG2025

Credal Ensemble Distillation for Uncertainty Quantification

Kaizheng Wang, Fabio Cuzzolin, David Moens +1

Deep ensembles (DE) have emerged as a powerful approach for quantifying predictive uncertainty and distinguishing its aleatoric and epistemic components, thereby enhancing model ro…

cs.AI2026

Random-Set Graph Neural Networks

Tommy Woodley, Shireen Kudukkil Manchingal, Matteo Tolloso +2

Uncertainty quantification has become an important factor in understanding the data representations produced by Graph Neural Networks (GNNs). Despite their predictive capabilities…

cs.CV2022

ROAD: The ROad event Awareness Dataset for Autonomous Driving

Gurkirt Singh, Stephen Akrigg, Manuele Di Maio +13

Humans drive in a holistic fashion which entails, in particular, understanding dynamic road events and their evolution. Injecting these capabilities in autonomous vehicles can thus…

cs.CV2018

Predicting Action Tubes

Gurkirt Singh, Suman Saha, Fabio Cuzzolin

In this work, we present a method to predict an entire `action tube' (a set of temporally linked bounding boxes) in a trimmed video just by observing a smaller subset of it. Predic…