papers

Publications (191)

cs.CV2024

PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos

Meng Cao, Haoran Tang, Haoze Zhao +7

Recent advancements in video-based large language models (Video LLMs) have witnessed the emergence of diverse capabilities to reason and interpret dynamic visual content. Among the…

cs.CV2023

Residual Pattern Learning for Pixel-wise Out-of-Distribution Detection in Semantic Segmentation

Yuyuan Liu, Choubo Ding, Yu Tian +4

Semantic segmentation models classify pixels into a set of known (``in-distribution'') visual classes. When deployed in an open world, the reliability of these models depends on th…

cs.CV2020

Depth Based Semantic Scene Completion with Position Importance Aware Loss

Yu Liu, Jie Li, Xia Yuan +4

Semantic Scene Completion (SSC) refers to the task of inferring the 3D semantic segmentation of a scene while simultaneously completing the 3D shapes. We propose PALNet, a novel hy…

cs.CV2021

Rotation Coordinate Descent for Fast Globally Optimal Rotation Averaging

Álvaro Parra, Shin-Fang Chng, Tat-Jun Chin +2

Under mild conditions on the noise level of the measurements, rotation averaging satisfies strong duality, which enables global solutions to be obtained via semidefinite programmin…

cs.CV2020

Hyperspectral Classification Based on 3D Asymmetric Inception Network with Data Fusion Transfer Learning

Haokui Zhang, Yu Liu, Bei Fang +3

Hyperspectral image(HSI) classification has been improved with convolutional neural network(CNN) in very recent years. Being different from the RGB datasets, different HSI datasets…

cs.LG2020

EvidentialMix: Learning with Combined Open-set and Closed-set Noisy Labels

Ragav Sachdeva, Filipe R. Cordeiro, Vasileios Belagiannis +2

The efficacy of deep learning depends on large-scale data sets that have been carefully curated with reliable data acquisition and annotation processes. However, acquiring such lar…

cs.CV2018

MatchBench: An Evaluation of Feature Matchers

JiaWang Bian, Ruihan Yang, Yun Liu +4

Feature matching is one of the most fundamental and active research areas in computer vision. A comprehensive evaluation of feature matchers is necessary, since it would advance bo…

cs.CV2018

SceneCut: Joint Geometric and Object Segmentation for Indoor Scenes

Trung Pham, Thanh-Toan Do, Niko Sünderhauf +1

This paper presents SceneCut, a novel approach to jointly discover previously unseen objects and non-object surfaces using a single RGB-D image. SceneCut's joint reasoning over sce…

cs.CV2023

Semantic Segmentation on 3D Point Clouds with High Density Variations

Ryan Faulkner, Luke Haub, Simon Ratcliffe +2

LiDAR scanning for surveying applications acquire measurements over wide areas and long distances, which produces large-scale 3D point clouds with significant local density variati…

cs.CV2018

AffordanceNet: An End-to-End Deep Learning Approach for Object Affordance Detection

Thanh-Toan Do, Anh Nguyen, Ian Reid

We propose AffordanceNet, a new deep learning approach to simultaneously detect multiple objects and their affordances from RGB images. Our AffordanceNet has two branches: an objec…

cs.RO2017

Past, Present, and Future of Simultaneous Localization And Mapping: Towards the Robust-Perception Age

Cesar Cadena, Luca Carlone, Henry Carrillo +5

Simultaneous Localization and Mapping (SLAM)consists in the concurrent construction of a model of the environment (the map), and the estimation of the state of the robot moving wit…

cs.CV2018

Light-Weight RefineNet for Real-Time Semantic Segmentation

Vladimir Nekrasov, Chunhua Shen, Ian Reid

We consider an important task of effective and efficient semantic image segmentation. In particular, we adapt a powerful semantic segmentation architecture, called RefineNet, into…

cs.CV2015

Learning Depth from Single Monocular Images Using Deep Convolutional Neural Fields

Fayao Liu, Chunhua Shen, Guosheng Lin +1

In this article, we tackle the problem of depth estimation from single monocular images. Compared with depth estimation using multiple images such as stereo depth perception, depth…

cs.CV2018

Just-in-Time Reconstruction: Inpainting Sparse Maps using Single View Depth Predictors as Priors

Chamara Saroj Weerasekera, Thanuja Dharmasiri, Ravi Garg +2

We present ``just-in-time reconstruction" as real-time image-guided inpainting of a map with arbitrary scale and sparsity to generate a fully dense depth map for the image. In part…

cs.CV2018

Scalable Deep -Subspace Clustering

Tong Zhang, Pan Ji, Mehrtash Harandi +2

Subspace clustering algorithms are notorious for their scalability issues because building and processing large affinity matrices are demanding. In this paper, we introduce a metho…

cs.CV2025

Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling

Meng Cao, Haokun Lin, Haoyuan Li +6

Spatial reasoning, the ability to understand and interpret the 3D structure of the world, is a critical yet underdeveloped capability in Multimodal Large Language Models (MLLMs). C…

cs.CV2019

Attention-guided Network for Ghost-free High Dynamic Range Imaging

Qingsen Yan, Dong Gong, Qinfeng Shi +4

Ghosting artifacts caused by moving objects or misalignments is a key challenge in high dynamic range (HDR) imaging for dynamic scenes. Previous methods first register the input lo…

cs.RO2026

ManipArena: Comprehensive Real-world Evaluation of Reasoning-Oriented Generalist Robot Manipulation

Yu Sun, Meng Cao, Yang Ping +24

Vision-Language-Action (VLA) models and world-action models have emerged as central paradigms for general-purpose robotic intelligence, yet their empirical progress remains constra…

cs.CV2017

"Maximizing rigidity" revisited: a convex programming approach for generic 3D shape reconstruction from multiple perspective views

Pan Ji, Hongdong Li, Yuchao Dai +1

Rigid structure-from-motion (RSfM) and non-rigid structure-from-motion (NRSfM) have long been treated in the literature as separate (different) problems. Inspired by a previous wor…

cs.CV2024

Simultaneous Diffusion Sampling for Conditional LiDAR Generation

Ryan Faulkner, Luke Haub, Simon Ratcliffe +3

By enabling capturing of 3D point clouds that reflect the geometry of the immediate environment, LiDAR has emerged as a primary sensor for autonomous systems. If a LiDAR scan is to…

cs.CV2018

Binary Constrained Deep Hashing Network for Image Retrieval without Manual Annotation

Thanh-Toan Do, Tuan Hoang, Dang-Khoa Le Tan +4

Learning compact binary codes for image retrieval task using deep neural networks has attracted increasing attention recently. However, training deep hashing networks for the task…

cs.RO2025

Action Tokenizer Matters in In-Context Imitation Learning

An Dinh Vuong, Minh Nhat Vu, Dong An +1

In-context imitation learning (ICIL) is a new paradigm that enables robots to generalize from demonstrations to unseen tasks without retraining. A well-structured action representa…

cs.CV2025

Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models

Meng Cao, Pengfei Hu, Yingyao Wang +12

Recent advancements in Large Video Language Models (LVLMs) have highlighted their potential for multi-modal understanding, yet evaluating their factual grounding in videos remains…

cs.CV2026

Order from Chaos: Physical World Understanding from Glitchy Gameplay Videos

Meng Cao, Haoran Tang, Haoze Zhao +6

Understanding the physical world, including object dynamics, material properties, and causal interactions, remains a core challenge in artificial intelligence. Although recent mult…

cs.CV2015

MOTChallenge 2015: Towards a Benchmark for Multi-Target Tracking

Laura Leal-Taixé, Anton Milan, Ian Reid +2

In the recent past, the computer vision community has developed centralized benchmarks for the performance evaluation of a variety of tasks, including generic object and pedestrian…

cs.RO2025

Constraint-Aware Zero-Shot Vision-Language Navigation in Continuous Environments

Kehan Chen, Dong An, Yan Huang +5

We address the task of Vision-Language Navigation in Continuous Environments (VLN-CE) under the zero-shot setting. Zero-shot VLN-CE is particularly challenging due to the absence o…

cs.LG2019

Improved Visual Localization via Graph Smoothing

Carlos Lassance, Yasir Latif, Ravi Garg +2

Vision based localization is the problem of inferring the pose of the camera given a single image. One solution to this problem is to learn a deep neural network to infer the pose…

cs.CV2019

V2CNet: A Deep Learning Framework to Translate Videos to Commands for Robotic Manipulation

Anh Nguyen, Thanh-Toan Do, Ian Reid +2

We propose V2CNet, a new deep learning framework to automatically translate the demonstration videos to commands that can be directly used in robotic applications. Our V2CNet has t…

cs.CV2017

Joint Learning of Set Cardinality and State Distribution

S. Hamid Rezatofighi, Anton Milan, Qinfeng Shi +2

We present a novel approach for learning to predict sets using deep learning. In recent years, deep neural networks have shown remarkable results in computer vision, natural langua…

cs.CV2018

Deep-6DPose: Recovering 6D Object Pose from a Single RGB Image

Thanh-Toan Do, Ming Cai, Trung Pham +1

Detecting objects and their 6D poses from only RGB images is an important task for many robotic applications. While deep learning methods have made significant progress in visual o…

cs.CV2022

Structured Binary Neural Networks for Image Recognition

Bohan Zhuang, Chunhua Shen, Mingkui Tan +3

We propose methods to train convolutional neural networks (CNNs) with both binarized weights and activations, leading to quantized models that are specifically friendly to mobile d…

cs.CV2022

Asynchronous Optimisation for Event-based Visual Odometry

Daqi Liu, Alvaro Parra, Yasir Latif +3

Event cameras open up new possibilities for robotic perception due to their low latency and high dynamic range. On the other hand, developing effective event-based vision algorithm…

cs.CV2019

Adaptive Low-Rank Kernel Subspace Clustering

Pan Ji, Ian Reid, Ravi Garg +2

In this paper, we present a kernel subspace clustering method that can handle non-linear models. In contrast to recent kernel subspace clustering methods which use predefined kerne…

cs.CV2021

Self-supervised Mean Teacher for Semi-supervised Chest X-ray Classification

Fengbei Liu, Yu Tian, Filipe R. Cordeiro +3

The training of deep learning models generally requires a large amount of annotated data for effective convergence and generalisation. However, obtaining high-quality annotations i…

cs.CV2018

Unsupervised Learning of Monocular Depth Estimation and Visual Odometry with Deep Feature Reconstruction

Huangying Zhan, Ravi Garg, Chamara Saroj Weerasekera +3

Despite learning based methods showing promising results in single view depth estimation and visual odometry, most existing approaches treat the tasks in a supervised manner. Recen…

cs.RO2026

ROSflight 2.0: Lean ROS 2-Based Autopilot for Unmanned Aerial Vehicles

Jacob Moore, Phil Tokumaru, Ian Reid +4

ROSflight is a lean, open-source autopilot ecosystem for unmanned aerial vehicles (UAVs). Designed by researchers for researchers, it is built to lower the barrier to entry to UAV…

cs.RO2025

Improving Robotic Manipulation with Efficient Geometry-Aware Vision Encoder

An Dinh Vuong, Minh Nhat Vu, Ian Reid

Existing RGB-based imitation learning approaches typically employ traditional vision encoders such as ResNet or ViT, which lack explicit 3D reasoning capabilities. Recent geometry-…

cs.CV2017

Care about you: towards large-scale human-centric visual relationship detection

Bohan Zhuang, Qi Wu, Chunhua Shen +2

Visual relationship detection aims to capture interactions between pairs of objects in images. Relationships between objects and humans represent a particularly important subset of…

cs.CV2026

CalibAnyView: Beyond Single-View Camera Calibration in the Wild

Boying Li, Cheng Zhang, Weirong Chen +5

Camera calibration is fundamental to reliable geometric perception, yet classical approaches rely on dedicated targets, successful reconstruction, or dense view coverage, which cas…

cs.CV2023

SC-DepthV3: Robust Self-supervised Monocular Depth Estimation for Dynamic Scenes

Libo Sun, Jia-Wang Bian, Huangying Zhan +3

Self-supervised monocular depth estimation has shown impressive results in static scenes. It relies on the multi-view consistency assumption for training networks, however, that is…

cs.CV2019

In defense of OSVOS

Yu Liu, Yutong Dai, Anh-Dzung Doan +2

As a milestone for video object segmentation, one-shot video object segmentation (OSVOS) has achieved a large margin compared to the conventional optical-flow based methods regardi…

cs.LG2024

Oops, I Sampled it Again: Reinterpreting Confidence Intervals in Few-Shot Learning

Raphael Lafargue, Luke Smith, Franck Vermet +4

The predominant method for computing confidence intervals (CI) in few-shot learning (FSL) is based on sampling the tasks with replacement, i.e.\ allowing the same samples to appear…

cs.RO2022

Autonomy and Perception for Space Mining

Ragav Sachdeva, Ravi Hammond, James Bockman +10

Future Moon bases will likely be constructed using resources mined from the surface of the Moon. The difficulty of maintaining a human workforce on the Moon and communications lag…

cs.CV2021

Effective Training of Convolutional Neural Networks with Low-bitwidth Weights and Activations

Bohan Zhuang, Jing Liu, Mingkui Tan +3

This paper tackles the problem of training a deep convolutional neural network of both low-bitwidth weights and activations. Optimizing a low-precision network is very challenging…

cs.GR2019

A Generalized Framework for Edge-preserving and Structure-preserving Image Smoothing

Wei Liu, Pingping Zhang, Yinjie Lei +3

Image smoothing is a fundamental procedure in applications of both computer vision and graphics. The required smoothing properties can be different or even contradictive among diff…

cs.CV2024

InfiniMotion: Mamba Boosts Memory in Transformer for Arbitrary Long Motion Generation

Zeyu Zhang, Akide Liu, Qi Chen +5

Text-to-motion generation holds potential for film, gaming, and robotics, yet current methods often prioritize short motion generation, making it challenging to produce long motion…

cs.RO2022

Predicting Topological Maps for Visual Navigation in Unexplored Environments

Huangying Zhan, Hamid Rezatofighi, Ian Reid

We propose a robotic learning system for autonomous exploration and navigation in unexplored environments. We are motivated by the idea that even an unseen environment may be famil…

cs.LG2025

Rethinking Weight-Averaged Model-merging

Hu Wang, Congbo Ma, Ibrahim Almakky +3

Model merging, particularly through weight averaging, has shown surprising effectiveness in saving computations and improving model performance without any additional training. How…

cs.RO2024

GraspMamba: A Mamba-based Language-driven Grasp Detection Framework with Hierarchical Feature Learning

Huy Hoang Nguyen, An Vuong, Anh Nguyen +2

Grasp detection is a fundamental robotic task critical to the success of many industrial applications. However, current language-driven models for this task often struggle with clu…

cs.RO2023

SayPlan: Grounding Large Language Models using 3D Scene Graphs for Scalable Robot Task Planning

Krishan Rana, Jesse Haviland, Sourav Garg +3

Large language models (LLMs) have demonstrated impressive results in developing generalist planning agents for diverse tasks. However, grounding these plans in expansive, multi-flo…

cs.CV2026

World2Act: Latent Action Post-Training from World Model Dynamics

An Dinh Vuong, Tuan Van Vo, Abdullah Sohail +6

World Models (WMs) offer a promising mechanism for post-training Vision-Language-Action (VLA) policies by providing dynamics priors that improve generalization under task and scene…

cs.CV2019

Pre and Post-hoc Diagnosis and Interpretation of Malignancy from Breast DCE-MRI

Gabriel Maicas, Andrew P. Bradley, Jacinto C. Nascimento +2

We propose a new method for breast cancer screening from DCE-MRI based on a post-hoc approach that is trained using weakly annotated data (i.e., labels are available only at the im…

cs.CV2018

Deep Perm-Set Net: Learn to predict sets with unknown permutation and cardinality using deep neural networks

S. Hamid Rezatofighi, Roman Kaskman, Farbod T. Motlagh +4

Many real-world problems, e.g. object detection, have outputs that are naturally expressed as sets of entities. This creates a challenge for traditional deep neural networks which…

cs.CV2024

GaussCtrl: Multi-View Consistent Text-Driven 3D Gaussian Splatting Editing

Jing Wu, Jia-Wang Bian, Xinghui Li +4

We propose GaussCtrl, a text-driven method to edit a 3D scene reconstructed by the 3D Gaussian Splatting (3DGS). Our method first renders a collection of images by using the 3DGS a…

cs.CV2016

Unsupervised CNN for Single View Depth Estimation: Geometry to the Rescue

Ravi Garg, Vijay Kumar BG, Gustavo Carneiro +1

A significant weakness of most current deep Convolutional Neural Networks is the need to train them using vast amounts of manu- ally labelled data. In this work we propose a unsupe…

cs.CV2020

HM4: Hidden Markov Model with Memory Management for Visual Place Recognition

Anh-Dzung Doan, Yasir Latif, Tat-Jun Chin +1

Visual place recognition needs to be robust against appearance variability due to natural and man-made causes. Training data collection should thus be an ongoing process to allow c…

cs.CV2020

SG-VAE: Scene Grammar Variational Autoencoder to generate new indoor scenes

Pulak Purkait, Christopher Zach, Ian Reid

Deep generative models have been used in recent years to learn coherent latent representations in order to synthesize high-quality images. In this work, we propose a neural network…

cs.MA2023

Sensor Allocation and Online-Learning-based Path Planning for Maritime Situational Awareness Enhancement: A Multi-Agent Approach

Bach Long Nguyen, Anh-Dzung Doan, Tat-Jun Chin +5

Countries with access to large bodies of water often aim to protect their maritime transport by employing maritime surveillance systems. However, the number of available sensors (e…

cs.CV2020

MOTChallenge: A Benchmark for Single-Camera Multiple Target Tracking

Patrick Dendorfer, Aljoša Ošep, Anton Milan +5

Standardized benchmarks have been crucial in pushing the performance of computer vision algorithms, especially since the advent of deep learning. Although leaderboards should not b…

eess.SY2024

Learn 2 Rage: Experiencing The Emotional Roller Coaster That Is Reinforcement Learning

Lachlan Mares, Stefan Podgorski, Ian Reid

This work presents the experiments and solution outline for our teams winning submission in the Learn To Race Autonomous Racing Virtual Challenge 2022 hosted by AIcrowd. The object…

cs.CV2019

Model Agnostic Saliency for Weakly Supervised Lesion Detection from Breast DCE-MRI

Gabriel Maicas, Gerard Snaauw, Andrew P. Bradley +2

There is a heated debate on how to interpret the decisions provided by deep learning models (DLM), where the main approaches rely on the visualization of salient regions to interpr…

cs.CV2021

Learn to Predict Sets Using Feed-Forward Neural Networks

Hamid Rezatofighi, Tianyu Zhu, Roman Kaskman +6

This paper addresses the task of set prediction using deep feed-forward neural networks. A set is a collection of elements which is invariant under permutation and the size of a se…

cs.CV2020

Joint Learning of Social Groups, Individuals Action and Sub-group Activities in Videos

Mahsa Ehsanpour, Alireza Abedin, Fatemeh Saleh +3

The state-of-the art solutions for human activity understanding from a video stream formulate the task as a spatio-temporal problem which requires joint localization of all individ…

cs.CV2017

Deep Subspace Clustering Networks

Pan Ji, Tong Zhang, Hongdong Li +2

We present a novel deep neural network architecture for unsupervised subspace clustering. This architecture is built upon deep auto-encoders, which non-linearly map the input data…

cs.CV2019

Learning Pairwise Relationship for Multi-object Detection in Crowded Scenes

Yu Liu, Lingqiao Liu, Hamid Rezatofighi +3

As the post-processing step for object detection, non-maximum suppression (GreedyNMS) is widely used in most of the detectors for many years. It is efficient and accurate for spars…

cs.CV2019

Seeing Behind Things: Extending Semantic Segmentation to Occluded Regions

Pulak Purkait, Christopher Zach, Ian Reid

Semantic segmentation and instance level segmentation made substantial progress in recent years due to the emergence of deep neural networks (DNNs). A number of deep architectures…

cs.CV2019

Scalable Place Recognition Under Appearance Change for Autonomous Driving

Anh-Dzung Doan, Yasir Latif, Tat-Jun Chin +3

A major challenge in place recognition for autonomous driving is to be robust against appearance changes due to short-term (e.g., weather, lighting) and long-term (seasons, vegetat…

cs.CV2016

Fast Training of Triplet-based Deep Binary Embedding Networks

Bohan Zhuang, Guosheng Lin, Chunhua Shen +1

In this paper, we aim to learn a mapping (or embedding) from images to a compact binary space in which Hamming distances correspond to a ranking measure for the image retrieval tas…

cs.CV2020

Switchable Precision Neural Networks

Luis Guerra, Bohan Zhuang, Ian Reid +1

Instantaneous and on demand accuracy-efficiency trade-off has been recently explored in the context of neural networks slimming. In this paper, we propose a flexible quantization s…

cs.CV2019

Architecture Search of Dynamic Cells for Semantic Video Segmentation

Vladimir Nekrasov, Hao Chen, Chunhua Shen +1

In semantic video segmentation the goal is to acquire consistent dense semantic labelling across image frames. To this end, recent approaches have been reliant on manually arranged…

cs.RO2026

Latent Action Pretraining Through World Modeling

Bahey Tharwat, Yara Nasser, Ali Abouzeid +1

Vision-Language-Action (VLA) models have gained popularity for learning robotic manipulation tasks that follow language instructions. State-of-the-art VLAs, such as OpenVLA and $π…

cs.CV2022

Globally Optimal Event-Based Divergence Estimation for Ventral Landing

Sofia McLeod, Gabriele Meoni, Dario Izzo +5

Event sensing is a major component in bio-inspired flight guidance and control systems. We explore the usage of event cameras for predicting time-to-contact (TTC) with the surface…

cs.CV2021

TRiPOD: Human Trajectory and Pose Dynamics Forecasting in the Wild

Vida Adeli, Mahsa Ehsanpour, Ian Reid +4

Joint forecasting of human trajectory and pose dynamics is a fundamental building block of various applications ranging from robotics and autonomous driving to surveillance systems…

cs.CV2019

Single-view Object Shape Reconstruction Using Deep Shape Prior and Silhouette

Kejie Li, Ravi Garg, Ming Cai +1

3D shape reconstruction from a single image is a highly ill-posed problem. Modern deep learning based systems try to solve this problem by learning an end-to-end mapping from image…

cs.CV2022

How Trustworthy are Performance Evaluations for Basic Vision Tasks?

Tran Thien Dat Nguyen, Hamid Rezatofighi, Ba-Ngu Vo +3

This paper examines performance evaluation criteria for basic vision tasks involving sets of objects namely, object detection, instance-level segmentation and multi-object tracking…

cs.CV2019

Meta Learning with Differentiable Closed-form Solver for Fast Video Object Segmentation

Yu Liu, Lingqiao Liu, Haokui Zhang +2

This paper tackles the problem of video object segmentation. We are specifically concerned with the task of segmenting all pixels of a target object in all frames, given the annota…

cs.CV2019

Self-supervised Learning for Single View Depth and Surface Normal Estimation

Huangying Zhan, Chamara Saroj Weerasekera, Ravi Garg +1

In this work we present a self-supervised learning framework to simultaneously train two Convolutional Neural Networks (CNNs) to predict depth and surface normals from a single ima…

cs.CV2022

You Only Cut Once: Boosting Data Augmentation with a Single Cut

Junlin Han, Pengfei Fang, Weihao Li +5

We present You Only Cut Once (YOCO) for performing data augmentations. YOCO cuts one image into two pieces and performs data augmentations individually within each piece. Applying…

cs.RO2026

ROScopter: A Multirotor Autopilot based on ROSflight 2.0

Jacob Moore, Ian Reid, Phil Tokumaru +2

ROScopter is a lean multirotor autopilot built for researchers. ROScopter seeks to accelerate simulation and hardware testing of research code with an architecture that is both eas…

cs.CV2021

Auto-Rectify Network for Unsupervised Indoor Depth Estimation

Jia-Wang Bian, Huangying Zhan, Naiyan Wang +3

Single-View depth estimation using the CNNs trained from unlabelled videos has shown significant promise. However, excellent results have mostly been obtained in street-scene drivi…

cs.CV2021

Unsupervised Scale-consistent Depth Learning from Video

Jia-Wang Bian, Huangying Zhan, Naiyan Wang +5

We propose a monocular depth estimator SC-Depth, which requires only unlabelled videos for training and enables the scale-consistent prediction at inference time. Our contributions…

cs.CV2016

Online Multi-Target Tracking Using Recurrent Neural Networks

Anton Milan, Seyed Hamid Rezatofighi, Anthony Dick +2

We present a novel approach to online multi-target tracking based on recurrent neural networks (RNNs). Tracking multiple objects in real-world scenes involves many challenges, incl…

cs.CV2024

Social-MAE: Social Masked Autoencoder for Multi-person Motion Representation Learning

Mahsa Ehsanpour, Ian Reid, Hamid Rezatofighi

For a complete comprehension of multi-person scenes, it is essential to go beyond basic tasks like detection and tracking. Higher-level tasks, such as understanding the interaction…

cs.RO2019

Practical Robot Learning from Demonstrations using Deep End-to-End Training

Akansel Cosgun, Thomas Rowntree, Ian Reid +1

Robots need to learn behaviors in intuitive and practical ways for widespread deployment in human environments. To learn a robot behavior end-to-end, we train a variant of the ResN…

cs.CV2018

Training Compact Neural Networks with Binary Weights and Low Precision Activations

Bohan Zhuang, Chunhua Shen, Ian Reid

In this paper, we propose to train a network with binary weights and low-bitwidth activations, designed especially for mobile devices with limited power consumption. Most previous…

cs.CV2017

Exploring Context with Deep Structured models for Semantic Segmentation

Guosheng Lin, Chunhua Shen, Anton van den Hengel +1

State-of-the-art semantic image segmentation methods are mostly based on training deep convolutional neural networks (CNNs). In this work, we proffer to improve semantic segmentati…

cs.CV2026

GeoWorld: Geometric World Models

Zeyu Zhang, Danning Li, Ian Reid +1

Energy-based predictive world models provide a powerful approach for multi-step visual planning by reasoning over latent energy landscapes rather than generating pixels. However, e…

cs.CV2016

RefineNet: Multi-Path Refinement Networks for High-Resolution Semantic Segmentation

Guosheng Lin, Anton Milan, Chunhua Shen +1

Recently, very deep convolutional neural networks (CNNs) have shown outstanding performance in object recognition and have also been the first choice for dense classification probl…

cs.RO2017

A Branch-and-Bound Algorithm for Checkerboard Extraction in Camera-Laser Calibration

Alireza Khosravian, Tat-Jun Chin, Ian Reid

We address the problem of camera-to-laser-scanner calibration using a checkerboard and multiple image-laser scan pairs. Distinguishing which laser points measure the checkerboard a…

cs.CV2023

What Images are More Memorable to Machines?

Junlin Han, Huangying Zhan, Jie Hong +4

This paper studies the problem of measuring and predicting how memorable an image is to pattern recognition machines, as a path to explore machine intelligence. Firstly, we propose…

cs.CV2018

Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments

Peter Anderson, Qi Wu, Damien Teney +6

A robot that can carry out a natural-language instruction has been a dream since before the Jetsons cartoon series imagined a life of leisure mediated by a fleet of attentive robot…

cs.CV2020

MOT20: A benchmark for multi object tracking in crowded scenes

Patrick Dendorfer, Hamid Rezatofighi, Anton Milan +6

Standardized benchmarks are crucial for the majority of computer vision applications. Although leaderboards and ranking tables should not be over-claimed, benchmarks often provide…

cs.CV2017

Smart Mining for Deep Metric Learning

Ben Harwood, Vijay Kumar B G, Gustavo Carneiro +2

To solve deep metric learning problems and producing feature embeddings, current methodologies will commonly use a triplet model to minimise the relative distance between samples f…

cs.CV2025

ItTakesTwo: Leveraging Peer Representations for Semi-supervised LiDAR Semantic Segmentation

Yuyuan Liu, Yuanhong Chen, Hu Wang +3

The costly and time-consuming annotation process to produce large training sets for modelling semantic LiDAR segmentation methods has motivated the development of semi-supervised l…

cs.CV2024

Weakly Supervised Test-Time Domain Adaptation for Object Detection

Anh-Dzung Doan, Bach Long Nguyen, Terry Lim +6

Prior to deployment, an object detector is trained on a dataset compiled from a previous data collection campaign. However, the environment in which the object detector is deployed…

cs.CV2019

Training Medical Image Analysis Systems like Radiologists

Gabriel Maicas, Andrew P. Bradley, Jacinto C. Nascimento +2

The training of medical image analysis systems using machine learning approaches follows a common script: collect and annotate a large dataset, train the classifier on the training…

cs.CV2020

Visual Odometry Revisited: What Should Be Learnt?

Huangying Zhan, Chamara Saroj Weerasekera, Jiawang Bian +1

In this work we present a monocular visual odometry (VO) algorithm which leverages geometry-based methods and deep learning. Most existing VO/SLAM systems with superior performance…

cs.CV2019

Fast Neural Architecture Search of Compact Semantic Segmentation Models via Auxiliary Cells

Vladimir Nekrasov, Hao Chen, Chunhua Shen +1

Automated design of neural network architectures tailored for a specific task is an extremely promising, albeit inherently difficult, avenue to explore. While most results in this…

cs.CV2026

Indexing Multimodal Language Models for Large-scale Image Retrieval

Bahey Tharwat, Giorgos Kordopatis-Zilos, Pavel Suma +2

Multimodal Large Language Models (MLLMs) have demonstrated strong cross-modal reasoning capabilities, yet their potential for vision-only tasks remains underexplored. We investigat…

cs.CV2021

DF-VO: What Should Be Learnt for Visual Odometry?

Huangying Zhan, Chamara Saroj Weerasekera, Jia-Wang Bian +2

Multi-view geometry-based methods dominate the last few decades in monocular Visual Odometry for their superior performance, while they have been vulnerable to dynamic and low-text…