Publications (176)
AG-VPReID 2025: Aerial-Ground Video-based Person Re-identification Challenge Results
Kien Nguyen, Clinton Fookes, Sridha Sridharan +20
Person re-identification (ReID) across aerial and ground vantage points has become crucial for large-scale surveillance and public safety applications. Although significant progres…
Geometry-constrained Car Recognition Using a 3D Perspective Network
Rui Zeng, Zongyuan Ge, Simon Denman +2
We present a novel learning framework for vehicle recognition from a single RGB image. Unlike existing methods which only use attention mechanisms to locate 2D discriminative infor…
Domain Generalization in Biosignal Classification
Theekshana Dissanayake, Tharindu Fernando, Simon Denman +3
Objective: When training machine learning models, we often assume that the training data and evaluation data are sampled from the same distribution. However, this assumption is vio…
Improving the Generation of VAEs with High Dimensional Latent Spaces by the use of Hyperspherical Coordinates
Alejandro Ascarate, Leo Lebrat, Rodrigo Santa Cruz +2
Variational autoencoders (VAE) encode data into lower-dimensional latent vectors before decoding those vectors back to data. Once trained, decoding a random latent vector from the…
Temporarily-Aware Context Modelling using Generative Adversarial Networks for Speech Activity Detection
Tharindu Fernando, Sridha Sridharan, Mitchell McLaren +3
This paper presents a novel framework for Speech Activity Detection (SAD). Inspired by the recent success of multi-task learning approaches in the speech processing domain, we prop…
Automatic Event Detection for Signal-based Surveillance
Jingxin Xu, Clinton Fookes, Sridha Sridharan
Signal-based Surveillance systems such as Closed Circuits Televisions (CCTV) have been widely installed in public places. Those systems are normally used to find the events with se…
High-Dimensional Latents Should Be Diagnosed Through Phase Structure
Alejandro Ascarate, Leo Lebrat, Rodrigo Santa Cruz +2
We study autoencoder and variational-autoencoder latent spaces through the lens of spin-glass theory. The paper has two components. First, we formalize a latent-space spin-glass di…
A Deep Four-Stream Siamese Convolutional Neural Network with Joint Verification and Identification Loss for Person Re-detection
Amena Khatun, Simon Denman, Sridha Sridharan +1
State-of-the-art person re-identification systems that employ a triplet based deep network suffer from a poor generalization capability. In this paper, we propose a four stream Sia…
In Depth We Trust: Reliable Monocular Depth Supervision for Gaussian Splatting
Wenhui Xiao, Ethan Goan, Rodrigo Santa Cruz +4
Using accurate depth priors in 3D Gaussian Splatting helps mitigate artifacts caused by sparse training data and textureless surfaces. However, acquiring accurate depth maps requir…
Syn3DWound: A Synthetic Dataset for 3D Wound Bed Analysis
Léo Lebrat, Rodrigo Santa Cruz, Remi Chierchia +12
Wound management poses a significant challenge, particularly for bedridden patients and the elderly. Accurate diagnostic and healing monitoring can significantly benefit from moder…
Patient-independent Epileptic Seizure Prediction using Deep Learning Models
Theekshana Dissanayake, Tharindu Fernando, Simon Denman +2
Objective: Epilepsy is one of the most prevalent neurological diseases among humans and can lead to severe brain injuries, strokes, and brain tumors. Early detection of seizures ca…
Image2Mesh: A Learning Framework for Single Image 3D Reconstruction
Jhony K. Pontes, Chen Kong, Sridha Sridharan +3
One challenge that remains open in 3D deep learning is how to efficiently represent 3D data to feed deep networks. Recent works have relied on volumetric or point cloud representat…
First Shape, Then Meaning: Efficient Geometry and Semantics Learning for Indoor Reconstruction
Remi Chierchia, Léo Lebrat, David Ahmedt-Aristizabal +3
Neural Surface Reconstruction has become a standard methodology for indoor 3D reconstruction, with Signed Distance Functions (SDFs) proving particularly effective for representing…
Soft + Hardwired Attention: An LSTM Framework for Human Trajectory Prediction and Abnormal Event Detection
Tharindu Fernando, Simon Denman, Sridha Sridharan +1
As humans we possess an intuitive ability for navigation which we master through years of practice; however existing approaches to model this trait for diverse tasks including moni…
Robust and Interpretable Temporal Convolution Network for Event Detection in Lung Sound Recordings
Tharindu Fernando, Sridha Sridharan, Simon Denman +2
This paper proposes a novel framework for lung sound event detection, segmenting continuous lung sound recordings into discrete events and performing recognition on each event. Exp…
Two Stream LSTM: A Deep Fusion Framework for Human Action Recognition
Harshala Gammulle, Simon Denman, Sridha Sridharan +1
In this paper we address the problem of human action recognition from video sequences. Inspired by the exemplary results obtained via automatic feature learning and deep learning a…
Face Deepfakes -- A Comprehensive Review
Tharindu Fernando, Darshana Priyasad, Sridha Sridharan +2
In recent years, remarkable advancements in deep-fake generation technology have led to unprecedented leaps in its realism and capabilities. Despite these advances, we observe a no…
Multi-stage Learning for Radar Pulse Activity Segmentation
Zi Huang, Akila Pemasiri, Simon Denman +2
Radio signal recognition is a crucial function in electronic warfare. Precise identification and localisation of radar pulse activities are required by electronic warfare systems t…
Learning Dense Correspondence from Synthetic Environments
Mithun Lal, Anthony Paproki, Nariman Habili +3
Estimation of human shape and pose from a single image is a challenging task. It is an even more difficult problem to map the identified human shape onto a 3D human model. Existing…
InCloud: Incremental Learning for Point Cloud Place Recognition
Joshua Knights, Peyman Moghadam, Milad Ramezani +2
Place recognition is a fundamental component of robotics, and has seen tremendous improvements through the use of deep learning models in recent years. Networks can experience sign…
DeepForgeSeal: Latent Space-Driven Semi-Fragile Watermarking for Deepfake Detection Using Adversarial Reinforcement Learning
Tharindu Fernando, Clinton Fookes, Sridha Sridharan
Rapid advances in generative AI have led to increasingly realistic deepfakes, posing growing challenges for law enforcement and public trust. Existing passive deepfake detectors st…
LoGG3D-Net: Locally Guided Global Descriptor Learning for 3D Place Recognition
Kavisha Vidanapathirana, Milad Ramezani, Peyman Moghadam +2
Retrieval-based place recognition is an efficient and effective solution for re-localization within a pre-built map, or global data association for Simultaneous Localization and Ma…
A Robust Interpretable Deep Learning Classifier for Heart Anomaly Detection Without Segmentation
Theekshana Dissanayake, Tharindu Fernando, Simon Denman +3
Traditionally, abnormal heart sound classification is framed as a three-stage process. The first stage involves segmenting the phonocardiogram to detect fundamental heart sounds; a…
Cross-Branch Orthogonality for Improved Generalization in Face Deepfake Detection
Tharindu Fernando, Clinton Fookes, Sridha Sridharan +1
Remarkable advancements in generative AI technology have given rise to a spectrum of novel deepfake categories with unprecedented leaps in their realism, and deepfakes are increasi…
Probabilistic Surfel Fusion for Dense LiDAR Mapping
Chanoh Park, Soohwan Kim, Peyman Moghadam +2
With the recent development of high-end LiDARs, more and more systems are able to continuously map the environment while moving and producing spatially redundant information. Howev…
AI, Entrepreneurs, and Privacy: Deep Learning Outperforms Humans in Detecting Entrepreneurs from Image Data
Martin Obschonka, Christian Fisch, Tharindu Fernando +1
Occupational outcomes like entrepreneurship are generally considered personal information that individuals should have the autonomy to disclose. With the advancing capability of ar…
Aerial-Ground Person Re-ID
Huy Nguyen, Kien Nguyen, Sridha Sridharan +1
Person re-ID matches persons across multiple non-overlapping cameras. Despite the increasing deployment of airborne platforms in surveillance, current existing person re-ID benchma…
Wound3DAssist: A Practical Framework for 3D Wound Assessment
Remi Chierchia, Rodrigo Santa Cruz, Léo Lebrat +7
Managing chronic wounds remains a major healthcare challenge, with clinical assessment often relying on subjective and time-consuming manual documentation methods. Although 2D digi…
Wild-Places: A Large-Scale Dataset for Lidar Place Recognition in Unstructured Natural Environments
Joshua Knights, Kavisha Vidanapathirana, Milad Ramezani +3
Many existing datasets for lidar place recognition are solely representative of structured urban environments, and have recently been saturated in performance by deep learning base…
Fast & Slow Learning: Incorporating Synthetic Gradients in Neural Memory Controllers
Tharindu Fernando, Simon Denman, Sridha Sridharan +1
Neural Memory Networks (NMNs) have received increased attention in recent years compared to deep architectures that use a constrained memory. Despite their new appeal, the success…
Two-Stream Deep Feature Modelling for Automated Video Endoscopy Data Analysis
Harshala Gammulle, Simon Denman, Sridha Sridharan +1
Automating the analysis of imagery of the Gastrointestinal (GI) tract captured during endoscopy procedures has substantial potential benefits for patients, as it can provide diagno…
A Survey on Physics Informed Reinforcement Learning: Review and Open Problems
Chayan Banerjee, Kien Nguyen, Clinton Fookes +1
The inclusion of physical information in machine learning frameworks has revolutionized many application areas. This involves enhancing the learning process by incorporating physic…
Learning Test-time Augmentation for Content-based Image Retrieval
Osman Tursun, Simon Denman, Sridha Sridharan +1
Off-the-shelf convolutional neural network features achieve outstanding results in many image retrieval tasks. However, their invariance to target data is pre-defined by the networ…
MTRNet: A Generic Scene Text Eraser
Osman Tursun, Rui Zeng, Simon Denman +3
Text removal algorithms have been proposed for uni-lingual scripts with regular shapes and layouts. However, to the best of our knowledge, a generic text removal method which is ab…
Testing the Test: Score-Direction Instability in Class-Split Anomaly Detection
Alejandro Ascarate, Leo Lebrat, Rodrigo Santa Cruz +2
Within-dataset class-split evaluation is widely used as a proxy for fully unconditional out-of-distribution anomaly detection. We show that this protocol can become ill-posed when…
Point-PNG: Conditional Pseudo-Negatives Generation for Point Cloud Pre-Training
Sutharsan Mahendren, Saimunur Rahman, Piotr Koniusz +4
We propose Point-PNG, a novel self-supervised learning framework that generates conditional pseudo-negatives in the latent space to learn point cloud representations that are both…
FactoFormer: Factorized Hyperspectral Transformers with Self-Supervised Pretraining
Shaheer Mohamed, Maryam Haghighat, Tharindu Fernando +3
Hyperspectral images (HSIs) contain rich spectral and spatial information. Motivated by the success of transformers in the field of natural language processing and computer vision…
SALVE: A 3D Reconstruction Benchmark of Wounds from Consumer-grade Videos
Remi Chierchia, Leo Lebrat, David Ahmedt-Aristizabal +3
Managing chronic wounds is a global challenge that can be alleviated by the adoption of automatic systems for clinical wound assessment from consumer-grade videos. While 2D image a…
Locus: LiDAR-based Place Recognition using Spatiotemporal Higher-Order Pooling
Kavisha Vidanapathirana, Peyman Moghadam, Ben Harwood +3
Place Recognition enables the estimation of a globally consistent map and trajectory by providing non-local constraints in Simultaneous Localisation and Mapping (SLAM). This paper…
Learning Regional Attention over Multi-resolution Deep Convolutional Features for Trademark Retrieval
Osman Tursun, Simon Denman, Sridha Sridharan +1
Large-scale trademark retrieval is an important content-based image retrieval task. A recent study shows that off-the-shelf deep features aggregated with Regional-Maximum Activatio…
Biomechanically Accurate Gait Analysis: A 3d Human Reconstruction Framework for Markerless Estimation of Gait Parameters
Akila Pemasiri, Ethan Goan, Glen Lichtwark +3
This paper presents a biomechanically interpretable framework for gait analysis using 3D human reconstruction from video data. Unlike conventional keypoint based approaches, the pr…
A Survey on Graph-Based Deep Learning for Computational Histopathology
David Ahmedt-Aristizabal, Mohammad Ali Armin, Simon Denman +2
With the remarkable success of representation learning for prediction problems, we have witnessed a rapid expansion of the use of machine learning and deep learning for the analysi…
Im2Mesh GAN: Accurate 3D Hand Mesh Recovery from a Single RGB Image
Akila Pemasiri, Kien Nguyen Thanh, Sridha Sridharan +1
This work addresses hand mesh recovery from a single RGB image. In contrast to most of the existing approaches where the parametric hand models are employed as the prior, we show t…
Deep Decision Trees for Discriminative Dictionary Learning with Adversarial Multi-Agent Trajectories
Tharindu Fernando, Sridha Sridharan, Clinton Fookes +1
With the explosion in the availability of spatio-temporal tracking data in modern sports, there is an enormous opportunity to better analyse, learn and predict important events in…
Compact Model Representation for 3D Reconstruction
Jhony K. Pontes, Chen Kong, Anders Eriksson +3
3D reconstruction from 2D images is a central problem in computer vision. Recent works have been focusing on reconstruction directly from a single image. It is well known however t…
Does Interference Exist When Training a Once-For-All Network?
Jordan Shipard, Arnold Wiliem, Clinton Fookes
The Once-For-All (OFA) method offers an excellent pathway to deploy a trained neural network model into multiple target platforms by utilising the supernet-subnet architecture. Onc…
A Comparative Analysis of Registration Tools: Traditional vs Deep Learning Approach on High Resolution Tissue Cleared Data
Abdullah Nazib, Clinton Fookes, Dimitri Perrin
Image registration plays an important role in comparing images. It is particularly important in analyzing medical images like CT, MRI, PET, etc. to quantify different biological sa…
Multi-Level Sequence GAN for Group Activity Recognition
Harshala Gammulle, Simon Denman, Sridha Sridharan +1
We propose a novel semi-supervised, Multi-Level Sequential Generative Adversarial Network (MLS-GAN) architecture for group activity recognition. In contrast to previous works which…
Deep Auto-Encoders with Sequential Learning for Multimodal Dimensional Emotion Recognition
Dung Nguyen, Duc Thanh Nguyen, Rui Zeng +5
Multimodal dimensional emotion recognition has drawn a great attention from the affective computing community and numerous schemes have been extensively investigated, making a sign…
Mining--Gym: A Configurable RL Benchmarking Environment for Truck Dispatch Scheduling
Chayan Banerjee, Kien Nguyen, Clinton Fookes
Optimizing the mining process -- particularly truck dispatch scheduling -- is a key driver of efficiency in open-pit operations. However, the dynamic and stochastic nature of these…
Learning Temporal Strategic Relationships using Generative Adversarial Imitation Learning
Tharindu Fernando, Simon Denman, Sridha Sridharan +1
This paper presents a novel framework for automatic learning of complex strategies in human decision making. The task that we are interested in is to better facilitate long term pl…
DIS2: Disentanglement Meets Distillation with Classwise Attention for Robust Remote Sensing Segmentation under Missing Modalities
Nhi Kieu, Kien Nguyen, Arnold Wiliem +2
The efficacy of multimodal learning in remote sensing (RS) is severely undermined by missing modalities. The challenge is exacerbated by the RS highly heterogeneous data and huge s…
Uncertainty in Real-Time Semantic Segmentation on Embedded Systems
Ethan Goan, Clinton Fookes
Application for semantic segmentation models in areas such as autonomous vehicles and human computer interaction require real-time predictive capabilities. The challenges of addres…
Multi-component Image Translation for Deep Domain Generalization
Mohammad Mahfujur Rahman, Clinton Fookes, Mahsa Baktashmotlagh +1
Domain adaption (DA) and domain generalization (DG) are two closely related methods which are both concerned with the task of assigning labels to an unlabeled data set. The only di…
Multi-modal Fusion for Single-Stage Continuous Gesture Recognition
Harshala Gammulle, Simon Denman, Sridha Sridharan +1
Gesture recognition is a much studied research area which has myriad real-world applications including robotics and human-machine interaction. Current gesture recognition methods h…
Using Auxiliary Information for Person Re-Identification -- A Tutorial Overview
Tharindu Fernando, Clinton Fookes, Sridha Sridharan +1
Person re-identification (re-id) is a pivotal task within an intelligent surveillance pipeline and there exist numerous re-id frameworks that achieve satisfactory performance in ch…
Elastic LiDAR Fusion: Dense Map-Centric Continuous-Time SLAM
Chanoh Park, Peyman Moghadam, Soohwan Kim +3
The concept of continuous-time trajectory representation has brought increased accuracy and efficiency to multi-modal sensor fusion in modern SLAM. However, regardless of these adv…
VAE with Hyperspherical Coordinates: Improving Anomaly Detection from Hypervolume-Compressed Latent Space
Alejandro Ascarate, Leo Lebrat, Rodrigo Santa Cruz +2
Variational autoencoders (VAE) encode data into lower-dimensional latent vectors before decoding those vectors back to data. Once trained, one can hope to detect out-of-distributio…
Complex-valued Iris Recognition Network
Kien Nguyen, Clinton Fookes, Sridha Sridharan +1
In this work, we design a fully complex-valued neural network for the task of iris recognition. Unlike the problem of general object recognition, where real-valued neural networks…
Predicting the Future: A Jointly Learnt Model for Action Anticipation
Harshala Gammulle, Simon Denman, Sridha Sridharan +1
Inspired by human neurological structures for action anticipation, we present an action anticipation model that enables the prediction of plausible future actions by forecasting bo…
Efficient Localization of Directional Emitters via Joint Beampattern Estimation
Fraser Williams, Akila Pemasiri, Dhammika Jayalath +2
The localization of directional RF emitters presents significant challenges for electronic warfare applications. Traditional localization methods, designed for omnidirectional emit…
Physics Augmented Tuple Transformer for Autism Severity Level Detection
Chinthaka Ranasingha, Harshala Gammulle, Tharindu Fernando +2
Early diagnosis of Autism Spectrum Disorder (ASD) is an effective and favorable step towards enhancing the health and well-being of children with ASD. Manual ASD diagnosis testing…
Tracking by Prediction: A Deep Generative Model for Mutli-Person localisation and Tracking
Tharindu Fernando, Simon Denman, Sridha Sridharan +1
Current multi-person localisation and tracking systems have an over reliance on the use of appearance models for target re-identification and almost no approaches employ a complete…
MTRNet++: One-stage Mask-based Scene Text Eraser
Osman Tursun, Simon Denman, Rui Zeng +3
A precise, controllable, interpretable and easily trainable text removal approach is necessary for both user-specific and large-scale text removal applications. To achieve this, we…
NeRF Director: Revisiting View Selection in Neural Volume Rendering
Wenhui Xiao, Rodrigo Santa Cruz, David Ahmedt-Aristizabal +3
Neural Rendering representations have significantly contributed to the field of 3D computer vision. Given their potential, considerable efforts have been invested to improve their…
Task Specific Visual Saliency Prediction with Memory Augmented Conditional Generative Adversarial Networks
Tharindu Fernando, Simon Denman, Sridha Sridharan +1
Visual saliency patterns are the result of a variety of factors aside from the image being parsed, however existing approaches have ignored these. To address this limitation, we pr…
Filling the Gaps: A Multitask Hybrid Multiscale Generative Framework for Missing Modality in Remote Sensing Semantic Segmentation
Nhi Kieu, Kien Nguyen, Arnold Wiliem +2
Multimodal learning has shown significant performance boost compared to ordinary unimodal models across various domains. However, in real-world scenarios, multimodal signals are su…
Corrupting Attention: Evasion-Based Adversarial Attacks on Encoder Attention in Detection Transformers
Ridma Jayasundara, Shaheer Mohamed, Tharindu Fernando +6
Adversarial vulnerabilities remain a major concern for the safe deployment of neural networks, particularly in object detection, a core task embedded in many safety-critical system…
Part-based Quantitative Analysis for Heatmaps
Osman Tursun, Sinan Kalkan, Simon Denman +2
Heatmaps have been instrumental in helping understand deep network decisions, and are a common approach for Explainable AI (XAI). While significant progress has been made in enhanc…
Multi-task Learning for Radar Signal Characterisation
Zi Huang, Akila Pemasiri, Simon Denman +2
Radio signal recognition is a crucial task in both civilian and military applications, as accurate and timely identification of unknown signals is an essential part of spectrum man…
General-Purpose Multimodal Transformer meets Remote Sensing Semantic Segmentation
Nhi Kieu, Kien Nguyen, Sridha Sridharan +1
The advent of high-resolution multispectral/hyperspectral sensors, LiDAR DSM (Digital Surface Model) information and many others has provided us with an unprecedented wealth of dat…
Point Cloud Segmentation Using Sparse Temporal Local Attention
Joshua Knights, Peyman Moghadam, Clinton Fookes +1
Point clouds are a key modality used for perception in autonomous vehicles, providing the means for a robust geometric understanding of the surrounding environment. However despite…
Size and Smoothness Aware Adaptive Focal Loss for Small Tumor Segmentation
Md Rakibul Islam, Riad Hassan, Abdullah Nazib +3
Deep learning has achieved remarkable accuracy in medical image segmentation, particularly for larger structures with well-defined boundaries. However, its effectiveness can be cha…
Discriminative Domain-Invariant Adversarial Network for Deep Domain Generalization
Mohammad Mahfujur Rahman, Clinton Fookes, Sridha Sridharan
Domain generalization approaches aim to learn a domain invariant prediction model for unknown target domains from multiple training source domains with different distributions. Sig…
Heart Sound Segmentation using Bidirectional LSTMs with Attention
Tharindu Fernando, Houman Ghaemmaghami, Simon Denman +3
This paper proposes a novel framework for the segmentation of phonocardiogram (PCG) signals into heart states, exploiting the temporal evolution of the PCG as well as considering t…
Neural Memory Plasticity for Anomaly Detection
Tharindu Fernando, Simon Denman, David Ahmedt-Aristizabal +4
In the domain of machine learning, Neural Memory Networks (NMNs) have recently achieved impressive results in a variety of application areas including visual question answering, tr…
Sparse Over-complete Patch Matching
Akila Pemasiri, Kien Nguyen, Sridha Sridharan +1
Image patch matching, which is the process of identifying corresponding patches across images, has been used as a subroutine for many computer vision and image processing tasks. St…
HOTFLoc++: End-to-End Hierarchical LiDAR Place Recognition, Re-Ranking, and 6-DoF Metric Localisation in Forests
Ethan Griffiths, Maryam Haghighat, Simon Denman +2
This article presents HOTFLoc++, an end-to-end hierarchical framework for LiDAR place recognition, re-ranking, and 6-DoF metric localisation in forests. Leveraging an octree-based…
AG-ReID.v2: Bridging Aerial and Ground Views for Person Re-identification
Huy Nguyen, Kien Nguyen, Sridha Sridharan +1
Aerial-ground person re-identification (Re-ID) presents unique challenges in computer vision, stemming from the distinct differences in viewpoints, poses, and resolutions between h…
Improving Short Utterance PLDA Speaker Verification using SUV Modelling and Utterance Partitioning Approach
Ahilan Kanagasundaram, David Dean, Sridha Sridharan +1
This paper analyses the short utterance probabilistic linear discriminant analysis (PLDA) speaker verification with utterance partitioning and short utterance variance (SUV) modell…
Person Recognition in Aerial Surveillance: A Decade Survey
Kien Nguyen, Feng Liu, Clinton Fookes +3
The rapid emergence of airborne platforms and imaging sensors is enabling new forms of aerial surveillance due to their unprecedented advantages in scale, mobility, deployment, and…
Revisiting the Role of Texture in 3D Person Re-identification
Huy Nguyen, Kien Nguyen, Akila Pemasiri +2
This study introduces a new framework for 3D person re-identification (re-ID) that leverages readily available high-resolution texture data in 3D reconstruction to improve the perf…
Automatic Radar Signal Detection and FFT Estimation using Deep Learning
Akila Pemasiri, Zi Huang, Fraser Williams +4
This paper addresses a critical preliminary step in radar signal processing: detecting the presence of a radar signal and robustly estimating its bandwidth. Existing methods which…
Few-Shot Radar Signal Recognition through Self-Supervised Learning and Radio Frequency Domain Adaptation
Zi Huang, Simon Denman, Akila Pemasiri +2
Radar signal recognition (RSR) plays a pivotal role in electronic warfare (EW), as accurately classifying radar signals is critical for informing decision-making. Recent advances i…
Dense Deformation Network for High Resolution Tissue Cleared Image Registration
Abdullah Nazib, Clinton Fookes, Dimitri Perrin
The recent application of deep learning in various areas of medical image analysis has brought excellent performance gains. In particular, technologies based on deep learning in me…
AG-VPReID: A Challenging Large-Scale Benchmark for Aerial-Ground Video-based Person Re-Identification
Huy Nguyen, Kien Nguyen, Akila Pemasiri +3
We introduce AG-VPReID, a new large-scale dataset for aerial-ground video-based person re-identification (ReID) that comprises 6,632 subjects, 32,321 tracklets and over 9.6 million…
Non-rigid Reconstruction with a Single Moving RGB-D Camera
Shafeeq Elanattil, Peyman Moghadam, Sridha Sridharan +2
We present a novel non-rigid reconstruction method using a moving RGB-D camera. Current approaches use only non-rigid part of the scene and completely ignore the rigid background.…
Domain adaptation based Speaker Recognition on Short Utterances
Ahilan Kanagasundaram, David Dean, Sridha Sridharan +1
This paper explores how the in- and out-domain probabilistic linear discriminant analysis (PLDA) speaker verification behave when enrolment and verification lengths are reduced. Ex…
On Minimum Discrepancy Estimation for Deep Domain Adaptation
Mohammad Mahfujur Rahman, Clinton Fookes, Mahsa Baktashmotlagh +1
In the presence of large sets of labeled data, Deep Learning (DL) has accomplished extraordinary triumphs in the avenue of computer vision, particularly in object classification an…
PDV: Prompt Directional Vectors for Zero-shot Composed Image Retrieval
Osman Tursun, Sinan Kalkan, Simon Denman +1
Zero-shot Composed Image Retrieval (ZS-CIR) enables image search using a reference image and a text prompt without requiring specialized text-image composition networks trained on…
Spatiotemporal Camera-LiDAR Calibration: A Targetless and Structureless Approach
Chanoh Park, Peyman Moghadam, Soohwan Kim +2
The demand for multimodal sensing systems for robotics is growing due to the increase in robustness, reliability and accuracy offered by these systems. These systems also need to b…
AG-VPReID.VIR: Bridging Aerial and Ground Platforms for Video-based Visible-Infrared Person Re-ID
Huy Nguyen, Kien Nguyen, Akila Pemasiri +3
Person re-identification (Re-ID) across visible and infrared modalities is crucial for 24-hour surveillance systems, but existing datasets primarily focus on ground-level perspecti…
Deep Classification of Epileptic Signals
David Ahmedt-Aristizabal, Clinton Fookes, Kien Nguyen +1
Electrophysiological observation plays a major role in epilepsy evaluation. However, human interpretation of brain signals is subjective and prone to misdiagnosis. Automating this…
Performance of Image Registration Tools on High-Resolution 3D Brain Images
Abdullah Nazib, James Galloway, Clinton Fookes +1
Recent progress in tissue clearing has allowed for the imaging of entire organs at single-cell resolution. These methods produce very large 3D images (several gigabytes for a whole…
Dual-Domain Masked Image Modeling: A Self-Supervised Pretraining Strategy Using Spatial and Frequency Domain Masking for Hyperspectral Data
Shaheer Mohamed, Tharindu Fernando, Sridha Sridharan +2
Hyperspectral images (HSIs) capture rich spectral signatures that reveal vital material properties, offering broad applicability across various domains. However, the scarcity of la…
Bayesian Neural Networks: An Introduction and Survey
Ethan Goan, Clinton Fookes
Neural Networks (NNs) have provided state-of-the-art results for many challenging machine learning tasks such as detection, regression and classification across the domains of comp…
Divide and Conquer: Rethinking the Training Paradigm of Neural Radiance Fields
Rongkai Ma, Leo Lebrat, Rodrigo Santa Cruz +4
Neural radiance fields (NeRFs) have exhibited potential in synthesizing high-fidelity views of 3D scenes but the standard training paradigm of NeRF presupposes an equal importance…
Memory Augmented Deep Generative models for Forecasting the Next Shot Location in Tennis
Tharindu Fernando, Simon Denman, Sridha Sridharan +1
This paper presents a novel framework for predicting shot location and type in tennis. Inspired by recent neuroscience discoveries we incorporate neural memory modules to model the…
Piecewise Deterministic Markov Processes for Bayesian Neural Networks
Ethan Goan, Dimitri Perrin, Kerrie Mengersen +1
Inference on modern Bayesian Neural Networks (BNNs) often relies on a variational inference treatment, imposing violated assumptions of independence and the form of the posterior.…
Towards Self-Explainability of Deep Neural Networks with Heatmap Captioning and Large-Language Models
Osman Tursun, Simon Denman, Sridha Sridharan +1
Heatmaps are widely used to interpret deep neural networks, particularly for computer vision tasks, and the heatmap-based explainable AI (XAI) techniques are a well-researched topi…