Publications (191)
PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos
Meng Cao, Haoran Tang, Haoze Zhao +7
Recent advancements in video-based large language models (Video LLMs) have witnessed the emergence of diverse capabilities to reason and interpret dynamic visual content. Among the…
Residual Pattern Learning for Pixel-wise Out-of-Distribution Detection in Semantic Segmentation
Yuyuan Liu, Choubo Ding, Yu Tian +4
Semantic segmentation models classify pixels into a set of known (``in-distribution'') visual classes. When deployed in an open world, the reliability of these models depends on th…
Depth Based Semantic Scene Completion with Position Importance Aware Loss
Yu Liu, Jie Li, Xia Yuan +4
Semantic Scene Completion (SSC) refers to the task of inferring the 3D semantic segmentation of a scene while simultaneously completing the 3D shapes. We propose PALNet, a novel hy…
Rotation Coordinate Descent for Fast Globally Optimal Rotation Averaging
Ãlvaro Parra, Shin-Fang Chng, Tat-Jun Chin +2
Under mild conditions on the noise level of the measurements, rotation averaging satisfies strong duality, which enables global solutions to be obtained via semidefinite programmin…
Hyperspectral Classification Based on 3D Asymmetric Inception Network with Data Fusion Transfer Learning
Haokui Zhang, Yu Liu, Bei Fang +3
Hyperspectral image(HSI) classification has been improved with convolutional neural network(CNN) in very recent years. Being different from the RGB datasets, different HSI datasets…
EvidentialMix: Learning with Combined Open-set and Closed-set Noisy Labels
Ragav Sachdeva, Filipe R. Cordeiro, Vasileios Belagiannis +2
The efficacy of deep learning depends on large-scale data sets that have been carefully curated with reliable data acquisition and annotation processes. However, acquiring such lar…
MatchBench: An Evaluation of Feature Matchers
JiaWang Bian, Ruihan Yang, Yun Liu +4
Feature matching is one of the most fundamental and active research areas in computer vision. A comprehensive evaluation of feature matchers is necessary, since it would advance bo…
SceneCut: Joint Geometric and Object Segmentation for Indoor Scenes
Trung Pham, Thanh-Toan Do, Niko Sünderhauf +1
This paper presents SceneCut, a novel approach to jointly discover previously unseen objects and non-object surfaces using a single RGB-D image. SceneCut's joint reasoning over sce…
Semantic Segmentation on 3D Point Clouds with High Density Variations
Ryan Faulkner, Luke Haub, Simon Ratcliffe +2
LiDAR scanning for surveying applications acquire measurements over wide areas and long distances, which produces large-scale 3D point clouds with significant local density variati…
AffordanceNet: An End-to-End Deep Learning Approach for Object Affordance Detection
Thanh-Toan Do, Anh Nguyen, Ian Reid
We propose AffordanceNet, a new deep learning approach to simultaneously detect multiple objects and their affordances from RGB images. Our AffordanceNet has two branches: an objec…
Past, Present, and Future of Simultaneous Localization And Mapping: Towards the Robust-Perception Age
Cesar Cadena, Luca Carlone, Henry Carrillo +5
Simultaneous Localization and Mapping (SLAM)consists in the concurrent construction of a model of the environment (the map), and the estimation of the state of the robot moving wit…
Light-Weight RefineNet for Real-Time Semantic Segmentation
Vladimir Nekrasov, Chunhua Shen, Ian Reid
We consider an important task of effective and efficient semantic image segmentation. In particular, we adapt a powerful semantic segmentation architecture, called RefineNet, into…
Learning Depth from Single Monocular Images Using Deep Convolutional Neural Fields
Fayao Liu, Chunhua Shen, Guosheng Lin +1
In this article, we tackle the problem of depth estimation from single monocular images. Compared with depth estimation using multiple images such as stereo depth perception, depth…
Just-in-Time Reconstruction: Inpainting Sparse Maps using Single View Depth Predictors as Priors
Chamara Saroj Weerasekera, Thanuja Dharmasiri, Ravi Garg +2
We present ``just-in-time reconstruction" as real-time image-guided inpainting of a map with arbitrary scale and sparsity to generate a fully dense depth map for the image. In part…
Scalable Deep -Subspace Clustering
Tong Zhang, Pan Ji, Mehrtash Harandi +2
Subspace clustering algorithms are notorious for their scalability issues because building and processing large affinity matrices are demanding. In this paper, we introduce a metho…
Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling
Meng Cao, Haokun Lin, Haoyuan Li +6
Spatial reasoning, the ability to understand and interpret the 3D structure of the world, is a critical yet underdeveloped capability in Multimodal Large Language Models (MLLMs). C…
Attention-guided Network for Ghost-free High Dynamic Range Imaging
Qingsen Yan, Dong Gong, Qinfeng Shi +4
Ghosting artifacts caused by moving objects or misalignments is a key challenge in high dynamic range (HDR) imaging for dynamic scenes. Previous methods first register the input lo…
ManipArena: Comprehensive Real-world Evaluation of Reasoning-Oriented Generalist Robot Manipulation
Yu Sun, Meng Cao, Yang Ping +24
Vision-Language-Action (VLA) models and world-action models have emerged as central paradigms for general-purpose robotic intelligence, yet their empirical progress remains constra…
"Maximizing rigidity" revisited: a convex programming approach for generic 3D shape reconstruction from multiple perspective views
Pan Ji, Hongdong Li, Yuchao Dai +1
Rigid structure-from-motion (RSfM) and non-rigid structure-from-motion (NRSfM) have long been treated in the literature as separate (different) problems. Inspired by a previous wor…
Simultaneous Diffusion Sampling for Conditional LiDAR Generation
Ryan Faulkner, Luke Haub, Simon Ratcliffe +3
By enabling capturing of 3D point clouds that reflect the geometry of the immediate environment, LiDAR has emerged as a primary sensor for autonomous systems. If a LiDAR scan is to…
Binary Constrained Deep Hashing Network for Image Retrieval without Manual Annotation
Thanh-Toan Do, Tuan Hoang, Dang-Khoa Le Tan +4
Learning compact binary codes for image retrieval task using deep neural networks has attracted increasing attention recently. However, training deep hashing networks for the task…
Action Tokenizer Matters in In-Context Imitation Learning
An Dinh Vuong, Minh Nhat Vu, Dong An +1
In-context imitation learning (ICIL) is a new paradigm that enables robots to generalize from demonstrations to unseen tasks without retraining. A well-structured action representa…
Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models
Meng Cao, Pengfei Hu, Yingyao Wang +12
Recent advancements in Large Video Language Models (LVLMs) have highlighted their potential for multi-modal understanding, yet evaluating their factual grounding in videos remains…
Order from Chaos: Physical World Understanding from Glitchy Gameplay Videos
Meng Cao, Haoran Tang, Haoze Zhao +6
Understanding the physical world, including object dynamics, material properties, and causal interactions, remains a core challenge in artificial intelligence. Although recent mult…
MOTChallenge 2015: Towards a Benchmark for Multi-Target Tracking
Laura Leal-Taixé, Anton Milan, Ian Reid +2
In the recent past, the computer vision community has developed centralized benchmarks for the performance evaluation of a variety of tasks, including generic object and pedestrian…
Constraint-Aware Zero-Shot Vision-Language Navigation in Continuous Environments
Kehan Chen, Dong An, Yan Huang +5
We address the task of Vision-Language Navigation in Continuous Environments (VLN-CE) under the zero-shot setting. Zero-shot VLN-CE is particularly challenging due to the absence o…
Improved Visual Localization via Graph Smoothing
Carlos Lassance, Yasir Latif, Ravi Garg +2
Vision based localization is the problem of inferring the pose of the camera given a single image. One solution to this problem is to learn a deep neural network to infer the pose…
V2CNet: A Deep Learning Framework to Translate Videos to Commands for Robotic Manipulation
Anh Nguyen, Thanh-Toan Do, Ian Reid +2
We propose V2CNet, a new deep learning framework to automatically translate the demonstration videos to commands that can be directly used in robotic applications. Our V2CNet has t…
Joint Learning of Set Cardinality and State Distribution
S. Hamid Rezatofighi, Anton Milan, Qinfeng Shi +2
We present a novel approach for learning to predict sets using deep learning. In recent years, deep neural networks have shown remarkable results in computer vision, natural langua…
Deep-6DPose: Recovering 6D Object Pose from a Single RGB Image
Thanh-Toan Do, Ming Cai, Trung Pham +1
Detecting objects and their 6D poses from only RGB images is an important task for many robotic applications. While deep learning methods have made significant progress in visual o…
Structured Binary Neural Networks for Image Recognition
Bohan Zhuang, Chunhua Shen, Mingkui Tan +3
We propose methods to train convolutional neural networks (CNNs) with both binarized weights and activations, leading to quantized models that are specifically friendly to mobile d…
Asynchronous Optimisation for Event-based Visual Odometry
Daqi Liu, Alvaro Parra, Yasir Latif +3
Event cameras open up new possibilities for robotic perception due to their low latency and high dynamic range. On the other hand, developing effective event-based vision algorithm…
Adaptive Low-Rank Kernel Subspace Clustering
Pan Ji, Ian Reid, Ravi Garg +2
In this paper, we present a kernel subspace clustering method that can handle non-linear models. In contrast to recent kernel subspace clustering methods which use predefined kerne…
Self-supervised Mean Teacher for Semi-supervised Chest X-ray Classification
Fengbei Liu, Yu Tian, Filipe R. Cordeiro +3
The training of deep learning models generally requires a large amount of annotated data for effective convergence and generalisation. However, obtaining high-quality annotations i…
Unsupervised Learning of Monocular Depth Estimation and Visual Odometry with Deep Feature Reconstruction
Huangying Zhan, Ravi Garg, Chamara Saroj Weerasekera +3
Despite learning based methods showing promising results in single view depth estimation and visual odometry, most existing approaches treat the tasks in a supervised manner. Recen…
ROSflight 2.0: Lean ROS 2-Based Autopilot for Unmanned Aerial Vehicles
Jacob Moore, Phil Tokumaru, Ian Reid +4
ROSflight is a lean, open-source autopilot ecosystem for unmanned aerial vehicles (UAVs). Designed by researchers for researchers, it is built to lower the barrier to entry to UAV…
Improving Robotic Manipulation with Efficient Geometry-Aware Vision Encoder
An Dinh Vuong, Minh Nhat Vu, Ian Reid
Existing RGB-based imitation learning approaches typically employ traditional vision encoders such as ResNet or ViT, which lack explicit 3D reasoning capabilities. Recent geometry-…
Care about you: towards large-scale human-centric visual relationship detection
Bohan Zhuang, Qi Wu, Chunhua Shen +2
Visual relationship detection aims to capture interactions between pairs of objects in images. Relationships between objects and humans represent a particularly important subset of…
CalibAnyView: Beyond Single-View Camera Calibration in the Wild
Boying Li, Cheng Zhang, Weirong Chen +5
Camera calibration is fundamental to reliable geometric perception, yet classical approaches rely on dedicated targets, successful reconstruction, or dense view coverage, which cas…
SC-DepthV3: Robust Self-supervised Monocular Depth Estimation for Dynamic Scenes
Libo Sun, Jia-Wang Bian, Huangying Zhan +3
Self-supervised monocular depth estimation has shown impressive results in static scenes. It relies on the multi-view consistency assumption for training networks, however, that is…
In defense of OSVOS
Yu Liu, Yutong Dai, Anh-Dzung Doan +2
As a milestone for video object segmentation, one-shot video object segmentation (OSVOS) has achieved a large margin compared to the conventional optical-flow based methods regardi…
Oops, I Sampled it Again: Reinterpreting Confidence Intervals in Few-Shot Learning
Raphael Lafargue, Luke Smith, Franck Vermet +4
The predominant method for computing confidence intervals (CI) in few-shot learning (FSL) is based on sampling the tasks with replacement, i.e.\ allowing the same samples to appear…
Autonomy and Perception for Space Mining
Ragav Sachdeva, Ravi Hammond, James Bockman +10
Future Moon bases will likely be constructed using resources mined from the surface of the Moon. The difficulty of maintaining a human workforce on the Moon and communications lag…
Effective Training of Convolutional Neural Networks with Low-bitwidth Weights and Activations
Bohan Zhuang, Jing Liu, Mingkui Tan +3
This paper tackles the problem of training a deep convolutional neural network of both low-bitwidth weights and activations. Optimizing a low-precision network is very challenging…
A Generalized Framework for Edge-preserving and Structure-preserving Image Smoothing
Wei Liu, Pingping Zhang, Yinjie Lei +3
Image smoothing is a fundamental procedure in applications of both computer vision and graphics. The required smoothing properties can be different or even contradictive among diff…
InfiniMotion: Mamba Boosts Memory in Transformer for Arbitrary Long Motion Generation
Zeyu Zhang, Akide Liu, Qi Chen +5
Text-to-motion generation holds potential for film, gaming, and robotics, yet current methods often prioritize short motion generation, making it challenging to produce long motion…
Predicting Topological Maps for Visual Navigation in Unexplored Environments
Huangying Zhan, Hamid Rezatofighi, Ian Reid
We propose a robotic learning system for autonomous exploration and navigation in unexplored environments. We are motivated by the idea that even an unseen environment may be famil…
Rethinking Weight-Averaged Model-merging
Hu Wang, Congbo Ma, Ibrahim Almakky +3
Model merging, particularly through weight averaging, has shown surprising effectiveness in saving computations and improving model performance without any additional training. How…
GraspMamba: A Mamba-based Language-driven Grasp Detection Framework with Hierarchical Feature Learning
Huy Hoang Nguyen, An Vuong, Anh Nguyen +2
Grasp detection is a fundamental robotic task critical to the success of many industrial applications. However, current language-driven models for this task often struggle with clu…
SayPlan: Grounding Large Language Models using 3D Scene Graphs for Scalable Robot Task Planning
Krishan Rana, Jesse Haviland, Sourav Garg +3
Large language models (LLMs) have demonstrated impressive results in developing generalist planning agents for diverse tasks. However, grounding these plans in expansive, multi-flo…
World2Act: Latent Action Post-Training from World Model Dynamics
An Dinh Vuong, Tuan Van Vo, Abdullah Sohail +6
World Models (WMs) offer a promising mechanism for post-training Vision-Language-Action (VLA) policies by providing dynamics priors that improve generalization under task and scene…
Pre and Post-hoc Diagnosis and Interpretation of Malignancy from Breast DCE-MRI
Gabriel Maicas, Andrew P. Bradley, Jacinto C. Nascimento +2
We propose a new method for breast cancer screening from DCE-MRI based on a post-hoc approach that is trained using weakly annotated data (i.e., labels are available only at the im…
Deep Perm-Set Net: Learn to predict sets with unknown permutation and cardinality using deep neural networks
S. Hamid Rezatofighi, Roman Kaskman, Farbod T. Motlagh +4
Many real-world problems, e.g. object detection, have outputs that are naturally expressed as sets of entities. This creates a challenge for traditional deep neural networks which…
GaussCtrl: Multi-View Consistent Text-Driven 3D Gaussian Splatting Editing
Jing Wu, Jia-Wang Bian, Xinghui Li +4
We propose GaussCtrl, a text-driven method to edit a 3D scene reconstructed by the 3D Gaussian Splatting (3DGS). Our method first renders a collection of images by using the 3DGS a…
Unsupervised CNN for Single View Depth Estimation: Geometry to the Rescue
Ravi Garg, Vijay Kumar BG, Gustavo Carneiro +1
A significant weakness of most current deep Convolutional Neural Networks is the need to train them using vast amounts of manu- ally labelled data. In this work we propose a unsupe…
HM4: Hidden Markov Model with Memory Management for Visual Place Recognition
Anh-Dzung Doan, Yasir Latif, Tat-Jun Chin +1
Visual place recognition needs to be robust against appearance variability due to natural and man-made causes. Training data collection should thus be an ongoing process to allow c…
SG-VAE: Scene Grammar Variational Autoencoder to generate new indoor scenes
Pulak Purkait, Christopher Zach, Ian Reid
Deep generative models have been used in recent years to learn coherent latent representations in order to synthesize high-quality images. In this work, we propose a neural network…
Sensor Allocation and Online-Learning-based Path Planning for Maritime Situational Awareness Enhancement: A Multi-Agent Approach
Bach Long Nguyen, Anh-Dzung Doan, Tat-Jun Chin +5
Countries with access to large bodies of water often aim to protect their maritime transport by employing maritime surveillance systems. However, the number of available sensors (e…
MOTChallenge: A Benchmark for Single-Camera Multiple Target Tracking
Patrick Dendorfer, Aljoša Ošep, Anton Milan +5
Standardized benchmarks have been crucial in pushing the performance of computer vision algorithms, especially since the advent of deep learning. Although leaderboards should not b…
Learn 2 Rage: Experiencing The Emotional Roller Coaster That Is Reinforcement Learning
Lachlan Mares, Stefan Podgorski, Ian Reid
This work presents the experiments and solution outline for our teams winning submission in the Learn To Race Autonomous Racing Virtual Challenge 2022 hosted by AIcrowd. The object…
Model Agnostic Saliency for Weakly Supervised Lesion Detection from Breast DCE-MRI
Gabriel Maicas, Gerard Snaauw, Andrew P. Bradley +2
There is a heated debate on how to interpret the decisions provided by deep learning models (DLM), where the main approaches rely on the visualization of salient regions to interpr…
Learn to Predict Sets Using Feed-Forward Neural Networks
Hamid Rezatofighi, Tianyu Zhu, Roman Kaskman +6
This paper addresses the task of set prediction using deep feed-forward neural networks. A set is a collection of elements which is invariant under permutation and the size of a se…
Joint Learning of Social Groups, Individuals Action and Sub-group Activities in Videos
Mahsa Ehsanpour, Alireza Abedin, Fatemeh Saleh +3
The state-of-the art solutions for human activity understanding from a video stream formulate the task as a spatio-temporal problem which requires joint localization of all individ…
Deep Subspace Clustering Networks
Pan Ji, Tong Zhang, Hongdong Li +2
We present a novel deep neural network architecture for unsupervised subspace clustering. This architecture is built upon deep auto-encoders, which non-linearly map the input data…
Learning Pairwise Relationship for Multi-object Detection in Crowded Scenes
Yu Liu, Lingqiao Liu, Hamid Rezatofighi +3
As the post-processing step for object detection, non-maximum suppression (GreedyNMS) is widely used in most of the detectors for many years. It is efficient and accurate for spars…
Seeing Behind Things: Extending Semantic Segmentation to Occluded Regions
Pulak Purkait, Christopher Zach, Ian Reid
Semantic segmentation and instance level segmentation made substantial progress in recent years due to the emergence of deep neural networks (DNNs). A number of deep architectures…
Scalable Place Recognition Under Appearance Change for Autonomous Driving
Anh-Dzung Doan, Yasir Latif, Tat-Jun Chin +3
A major challenge in place recognition for autonomous driving is to be robust against appearance changes due to short-term (e.g., weather, lighting) and long-term (seasons, vegetat…
Fast Training of Triplet-based Deep Binary Embedding Networks
Bohan Zhuang, Guosheng Lin, Chunhua Shen +1
In this paper, we aim to learn a mapping (or embedding) from images to a compact binary space in which Hamming distances correspond to a ranking measure for the image retrieval tas…
Switchable Precision Neural Networks
Luis Guerra, Bohan Zhuang, Ian Reid +1
Instantaneous and on demand accuracy-efficiency trade-off has been recently explored in the context of neural networks slimming. In this paper, we propose a flexible quantization s…
Architecture Search of Dynamic Cells for Semantic Video Segmentation
Vladimir Nekrasov, Hao Chen, Chunhua Shen +1
In semantic video segmentation the goal is to acquire consistent dense semantic labelling across image frames. To this end, recent approaches have been reliant on manually arranged…
Latent Action Pretraining Through World Modeling
Bahey Tharwat, Yara Nasser, Ali Abouzeid +1
Vision-Language-Action (VLA) models have gained popularity for learning robotic manipulation tasks that follow language instructions. State-of-the-art VLAs, such as OpenVLA and $Ï…
Globally Optimal Event-Based Divergence Estimation for Ventral Landing
Sofia McLeod, Gabriele Meoni, Dario Izzo +5
Event sensing is a major component in bio-inspired flight guidance and control systems. We explore the usage of event cameras for predicting time-to-contact (TTC) with the surface…
TRiPOD: Human Trajectory and Pose Dynamics Forecasting in the Wild
Vida Adeli, Mahsa Ehsanpour, Ian Reid +4
Joint forecasting of human trajectory and pose dynamics is a fundamental building block of various applications ranging from robotics and autonomous driving to surveillance systems…
Single-view Object Shape Reconstruction Using Deep Shape Prior and Silhouette
Kejie Li, Ravi Garg, Ming Cai +1
3D shape reconstruction from a single image is a highly ill-posed problem. Modern deep learning based systems try to solve this problem by learning an end-to-end mapping from image…
How Trustworthy are Performance Evaluations for Basic Vision Tasks?
Tran Thien Dat Nguyen, Hamid Rezatofighi, Ba-Ngu Vo +3
This paper examines performance evaluation criteria for basic vision tasks involving sets of objects namely, object detection, instance-level segmentation and multi-object tracking…
Meta Learning with Differentiable Closed-form Solver for Fast Video Object Segmentation
Yu Liu, Lingqiao Liu, Haokui Zhang +2
This paper tackles the problem of video object segmentation. We are specifically concerned with the task of segmenting all pixels of a target object in all frames, given the annota…
Self-supervised Learning for Single View Depth and Surface Normal Estimation
Huangying Zhan, Chamara Saroj Weerasekera, Ravi Garg +1
In this work we present a self-supervised learning framework to simultaneously train two Convolutional Neural Networks (CNNs) to predict depth and surface normals from a single ima…
You Only Cut Once: Boosting Data Augmentation with a Single Cut
Junlin Han, Pengfei Fang, Weihao Li +5
We present You Only Cut Once (YOCO) for performing data augmentations. YOCO cuts one image into two pieces and performs data augmentations individually within each piece. Applying…
ROScopter: A Multirotor Autopilot based on ROSflight 2.0
Jacob Moore, Ian Reid, Phil Tokumaru +2
ROScopter is a lean multirotor autopilot built for researchers. ROScopter seeks to accelerate simulation and hardware testing of research code with an architecture that is both eas…
Auto-Rectify Network for Unsupervised Indoor Depth Estimation
Jia-Wang Bian, Huangying Zhan, Naiyan Wang +3
Single-View depth estimation using the CNNs trained from unlabelled videos has shown significant promise. However, excellent results have mostly been obtained in street-scene drivi…
Unsupervised Scale-consistent Depth Learning from Video
Jia-Wang Bian, Huangying Zhan, Naiyan Wang +5
We propose a monocular depth estimator SC-Depth, which requires only unlabelled videos for training and enables the scale-consistent prediction at inference time. Our contributions…
Online Multi-Target Tracking Using Recurrent Neural Networks
Anton Milan, Seyed Hamid Rezatofighi, Anthony Dick +2
We present a novel approach to online multi-target tracking based on recurrent neural networks (RNNs). Tracking multiple objects in real-world scenes involves many challenges, incl…
Social-MAE: Social Masked Autoencoder for Multi-person Motion Representation Learning
Mahsa Ehsanpour, Ian Reid, Hamid Rezatofighi
For a complete comprehension of multi-person scenes, it is essential to go beyond basic tasks like detection and tracking. Higher-level tasks, such as understanding the interaction…
Practical Robot Learning from Demonstrations using Deep End-to-End Training
Akansel Cosgun, Thomas Rowntree, Ian Reid +1
Robots need to learn behaviors in intuitive and practical ways for widespread deployment in human environments. To learn a robot behavior end-to-end, we train a variant of the ResN…
Training Compact Neural Networks with Binary Weights and Low Precision Activations
Bohan Zhuang, Chunhua Shen, Ian Reid
In this paper, we propose to train a network with binary weights and low-bitwidth activations, designed especially for mobile devices with limited power consumption. Most previous…
Exploring Context with Deep Structured models for Semantic Segmentation
Guosheng Lin, Chunhua Shen, Anton van den Hengel +1
State-of-the-art semantic image segmentation methods are mostly based on training deep convolutional neural networks (CNNs). In this work, we proffer to improve semantic segmentati…
GeoWorld: Geometric World Models
Zeyu Zhang, Danning Li, Ian Reid +1
Energy-based predictive world models provide a powerful approach for multi-step visual planning by reasoning over latent energy landscapes rather than generating pixels. However, e…
RefineNet: Multi-Path Refinement Networks for High-Resolution Semantic Segmentation
Guosheng Lin, Anton Milan, Chunhua Shen +1
Recently, very deep convolutional neural networks (CNNs) have shown outstanding performance in object recognition and have also been the first choice for dense classification probl…
A Branch-and-Bound Algorithm for Checkerboard Extraction in Camera-Laser Calibration
Alireza Khosravian, Tat-Jun Chin, Ian Reid
We address the problem of camera-to-laser-scanner calibration using a checkerboard and multiple image-laser scan pairs. Distinguishing which laser points measure the checkerboard a…
What Images are More Memorable to Machines?
Junlin Han, Huangying Zhan, Jie Hong +4
This paper studies the problem of measuring and predicting how memorable an image is to pattern recognition machines, as a path to explore machine intelligence. Firstly, we propose…
Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments
Peter Anderson, Qi Wu, Damien Teney +6
A robot that can carry out a natural-language instruction has been a dream since before the Jetsons cartoon series imagined a life of leisure mediated by a fleet of attentive robot…
MOT20: A benchmark for multi object tracking in crowded scenes
Patrick Dendorfer, Hamid Rezatofighi, Anton Milan +6
Standardized benchmarks are crucial for the majority of computer vision applications. Although leaderboards and ranking tables should not be over-claimed, benchmarks often provide…
Smart Mining for Deep Metric Learning
Ben Harwood, Vijay Kumar B G, Gustavo Carneiro +2
To solve deep metric learning problems and producing feature embeddings, current methodologies will commonly use a triplet model to minimise the relative distance between samples f…
ItTakesTwo: Leveraging Peer Representations for Semi-supervised LiDAR Semantic Segmentation
Yuyuan Liu, Yuanhong Chen, Hu Wang +3
The costly and time-consuming annotation process to produce large training sets for modelling semantic LiDAR segmentation methods has motivated the development of semi-supervised l…
Weakly Supervised Test-Time Domain Adaptation for Object Detection
Anh-Dzung Doan, Bach Long Nguyen, Terry Lim +6
Prior to deployment, an object detector is trained on a dataset compiled from a previous data collection campaign. However, the environment in which the object detector is deployed…
Training Medical Image Analysis Systems like Radiologists
Gabriel Maicas, Andrew P. Bradley, Jacinto C. Nascimento +2
The training of medical image analysis systems using machine learning approaches follows a common script: collect and annotate a large dataset, train the classifier on the training…
Visual Odometry Revisited: What Should Be Learnt?
Huangying Zhan, Chamara Saroj Weerasekera, Jiawang Bian +1
In this work we present a monocular visual odometry (VO) algorithm which leverages geometry-based methods and deep learning. Most existing VO/SLAM systems with superior performance…
Fast Neural Architecture Search of Compact Semantic Segmentation Models via Auxiliary Cells
Vladimir Nekrasov, Hao Chen, Chunhua Shen +1
Automated design of neural network architectures tailored for a specific task is an extremely promising, albeit inherently difficult, avenue to explore. While most results in this…
Indexing Multimodal Language Models for Large-scale Image Retrieval
Bahey Tharwat, Giorgos Kordopatis-Zilos, Pavel Suma +2
Multimodal Large Language Models (MLLMs) have demonstrated strong cross-modal reasoning capabilities, yet their potential for vision-only tasks remains underexplored. We investigat…
DF-VO: What Should Be Learnt for Visual Odometry?
Huangying Zhan, Chamara Saroj Weerasekera, Jia-Wang Bian +2
Multi-view geometry-based methods dominate the last few decades in monocular Visual Odometry for their superior performance, while they have been vulnerable to dynamic and low-text…