papers

Publications (122)

cs.LG2022

Learning Positional Embeddings for Coordinate-MLPs

Sameera Ramasinghe, Simon Lucey

We propose a novel method to enhance the performance of coordinate-MLPs by learning instance-specific positional embeddings. End-to-end optimization of positional embedding paramet…

cs.CV2017

Rethinking Reprojection: Closing the Loop for Pose-aware ShapeReconstruction from a Single Image

Rui Zhu, Hamed Kiani Galoogahi, Chaoyang Wang +1

An emerging problem in computer vision is the reconstruction of 3D shape and pose of an object from a single image. Hitherto, the problem has been addressed through the application…

cs.CL2026

Parameter-Efficient Fine-Tuning with Learnable Rank

Arpit Garg, Simon Lucey, Hemanth Saratchandran

Low-Rank Adaptation (LoRA) is a popular parameter-efficient fine-tuning (PEFT) method that restricts weight updates to low-rank adapters, introducing a fixed low-rank inductive bia…

cs.CV2014

Optimization Methods for Convolutional Sparse Coding

Hilton Bristow, Simon Lucey

Sparse and convolutional constraints form a natural prior for many optimization problems that arise from physical processes. Detecting motifs in speech and musical passages, super-…

cs.CV2014

Learning detectors quickly using structured covariance matrices

Jack Valmadre, Sridha Sridharan, Simon Lucey

Computer vision is increasingly becoming interested in the rapid estimation of object detectors. Canonical hard negative mining strategies are slow as they require multiple passes…

cs.CV2019

Deep Non-Rigid Structure from Motion with Missing Data

Chen Kong, Simon Lucey

Non-Rigid Structure from Motion (NRSfM) refers to the problem of reconstructing cameras and the 3D point cloud of a non-rigid object from an ensemble of images with 2D corresponden…

cs.CV2024

Convolutional Initialization for Data-Efficient Vision Transformers

Jianqiao Zheng, Xueqian Li, Simon Lucey

Training vision transformer networks on small datasets poses challenges. In contrast, convolutional neural networks (CNNs) can achieve state-of-the-art performance by leveraging th…

cs.CV2017

Image2Mesh: A Learning Framework for Single Image 3D Reconstruction

Jhony K. Pontes, Chen Kong, Sridha Sridharan +3

One challenge that remains open in 3D deep learning is how to efficiently represent 3D data to feed deep networks. Recent works have relied on volumetric or point cloud representat…

cs.CV2017

Learning Background-Aware Correlation Filters for Visual Tracking

Hamed Kiani Galoogahi, Ashton Fagg, Simon Lucey

Correlation Filters (CFs) have recently demonstrated excellent performance in terms of rapidly tracking objects under challenging photometric and geometric variations. The strength…

cs.CV2021

Neural Trajectory Fields for Dynamic Novel View Synthesis

Chaoyang Wang, Ben Eckart, Simon Lucey +1

Recent approaches to render photorealistic views from a limited set of photographs have pushed the boundaries of our interactions with pictures of static scenes. The ability to rec…

cs.LG2023

On progressive sharpening, flat minima and generalisation

Lachlan Ewen MacDonald, Jack Valmadre, Simon Lucey

We present a new approach to understanding the relationship between loss curvature and input-output model behaviour in deep learning. Specifically, we use existing empirical analys…

cs.CV2025

Enhancing Transformers Through Conditioned Embedded Tokens

Hemanth Saratchandran, Simon Lucey

Transformers have transformed modern machine learning, driving breakthroughs in computer vision, natural language processing, and robotics. At the core of their success lies the at…

cs.CV2016

The Conditional Lucas & Kanade Algorithm

Chen-Hsuan Lin, Rui Zhu, Simon Lucey

The Lucas & Kanade (LK) algorithm is the method of choice for efficient dense image and object alignment. The approach is efficient as it attempts to model the connection between a…

cs.CV2024

Multi-Body Neural Scene Flow

Kavisha Vidanapathirana, Shin-Fang Chng, Xueqian Li +1

The test-time optimization of scene flow - using a coordinate network as a neural prior - has gained popularity due to its simplicity, lack of dataset bias, and state-of-the-art pe…

cs.LG2022

How You Start Matters for Generalization

Sameera Ramasinghe, Lachlan MacDonald, Moshiur Farazi +2

Characterizing the remarkable generalization properties of over-parameterized neural networks remains an open problem. In this paper, we promote a shift of focus towards initializa…

cs.CV2019

Distill Knowledge from NRSfM for Weakly Supervised 3D Pose Learning

Chaoyang Wang, Chen Kong, Simon Lucey

We propose to learn a 3D pose estimator by distilling knowledge from Non-Rigid Structure from Motion (NRSfM). Our method uses solely 2D landmark annotations. No 3D data, multi-view…

cs.CV2018

Aligning Across Large Gaps in Time

Hunter Goforth, Simon Lucey

We present a method of temporally-invariant image registration for outdoor scenes, with invariance across time of day, across seasonal variations, and across decade-long periods, f…

cs.CV2026

3D-LFM: Lifting Foundation Model

Mosam Dabhi, Laszlo A. Jeni, Simon Lucey

The lifting of 3D structure and camera from 2D landmarks is at the cornerstone of the entire discipline of computer vision. Traditional methods have been confined to specific rigid…

cs.CV2022

Long-term Visual Map Sparsification with Heterogeneous GNN

Ming-Fang Chang, Yipu Zhao, Rajvi Shah +3

We address the problem of map sparsification for long-term visual localization. For map sparsification, a commonly employed assumption is that the pre-build map and the later captu…

cs.CV2024

SeMoLi: What Moves Together Belongs Together

Jenny Seidenschwarz, Aljoša Ošep, Francesco Ferroni +2

We tackle semi-supervised object detection based on motion cues. Recent results suggest that heuristic-based clustering methods in conjunction with object trackers can be used to p…

cs.CV2023

Re-Evaluating LiDAR Scene Flow for Autonomous Driving

Nathaniel Chodosh, Deva Ramanan, Simon Lucey

Popular benchmarks for self-supervised LiDAR scene flow (stereoKITTI, and FlyingThings3D) have unrealistic rates of dynamic motion, unrealistic correspondences, and unrealistic sam…

cs.CV2024

Invertible Neural Warp for NeRF

Shin-Fang Chng, Ravi Garg, Hemanth Saratchandran +1

This paper tackles the simultaneous optimization of pose and Neural Radiance Fields (NeRF). Departing from the conventional practice of using explicit global representations for ca…

cs.LG2020

Architectural Adversarial Robustness: The Case for Deep Pursuit

George Cazenavette, Calvin Murdock, Simon Lucey

Despite their unmatched performance, deep neural networks remain susceptible to targeted attacks by nearly imperceptible levels of adversarial noise. While the underlying cause of…

cs.CV2020

SDF-SRN: Learning Signed Distance 3D Object Reconstruction from Static Images

Chen-Hsuan Lin, Chaoyang Wang, Simon Lucey

Dense 3D object reconstruction from a single image has recently witnessed remarkable advances, but supervising neural networks with ground-truth 3D shapes is impractical due to the…

cs.CV2023

Curvature-Aware Training for Coordinate Networks

Hemanth Saratchandran, Shin-Fang Chng, Sameera Ramasinghe +2

Coordinate networks are widely used in computer vision due to their ability to represent signals as compressed, continuous entities. However, training these networks with first-ord…

cs.CV2017

Compact Model Representation for 3D Reconstruction

Jhony K. Pontes, Chen Kong, Anders Eriksson +3

3D reconstruction from 2D images is a central problem in computer vision. Recent works have been focusing on reconstruction directly from a single image. It is well known however t…

cs.LG2017

Take it in your stride: Do we need striding in CNNs?

Chen Kong, Simon Lucey

Since their inception, CNNs have utilized some type of striding operator to reduce the overlap of receptive fields and spatial dimensions. Although having clear heuristic motivatio…

cs.CV2018

Deep Convolutional Compressed Sensing for LiDAR Depth Completion

Nathaniel Chodosh, Chaoyang Wang, Simon Lucey

In this paper we consider the problem of estimating a dense depth map from a set of sparse LiDAR points. We use techniques from compressed sensing and the recently developed Altern…

cs.CV2016

Bit-Planes: Dense Subpixel Alignment of Binary Descriptors

Hatem Alismail, Brett Browning, Simon Lucey

Binary descriptors have been instrumental in the recent evolution of computationally efficient sparse image alignment algorithms. Increasingly, however, the vision community is int…

cs.CV2018

ST-GAN: Spatial Transformer Generative Adversarial Networks for Image Compositing

Chen-Hsuan Lin, Ersin Yumer, Oliver Wang +2

We address the problem of finding realistic geometric corrections to a foreground object such that it appears natural when composited into a background image. To achieve this, we p…

cs.CV2024

Fast Kernel Scene Flow

Xueqian Li, Simon Lucey

In contrast to current state-of-the-art methods, such as NSFP [25], which employ deep implicit neural functions for modeling scene flow, we present a novel approach that utilizes c…

cs.CV2022

GARF: Gaussian Activated Radiance Fields for High Fidelity Reconstruction and Pose Estimation

Shin-Fang Chng, Sameera Ramasinghe, Jamie Sherrah +1

Despite Neural Radiance Fields (NeRF) showing compelling results in photorealistic novel views synthesis of real-world scenes, most existing approaches require accurate prior camer…

cs.CV2019

Deep Interpretable Non-Rigid Structure from Motion

Chen Kong, Simon Lucey

All current non-rigid structure from motion (NRSfM) algorithms are limited with respect to: (i) the number of images, and (ii) the type of shape variability they can handle. This h…

cs.LG2025

Leaner Transformers: More Heads, Less Depth

Hemanth Saratchandran, Damien Teney, Simon Lucey

Transformers have reshaped machine learning by utilizing attention mechanisms to capture complex patterns in large datasets, leading to significant improvements in performance. Thi…

cs.LG2025

Gradient Descent as a Shrinkage Operator for Spectral Bias

Simon Lucey

We generalize the connection between activation function and spline regression/smoothing and characterize how this choice may influence spectral bias within a 1D shallow network. W…

cs.CV2014

Regression-Based Image Alignment for General Object Categories

Hilton Bristow, Simon Lucey

Gradient-descent methods have exhibited fast and reliable performance for image alignment in the facial domain, but have largely been ignored by the broader vision community. They…

cs.CV2025

Object Agnostic 3D Lifting in Space and Time

Christopher Fusco, Shin-Fang Ch'ng, Mosam Dabhi +1

We present a spatio-temporal perspective on category-agnostic 3D lifting of 2D keypoints over a temporal sequence. Our approach differs from existing state-of-the-art methods that…

cs.LG2026

The Quantization Benefits of Residual-Free Transformers

Yiping Ji, Mahalakshmi Sabanayagam, Peyman Moghadam +2

Large-scale transformer training and deployment are increasingly constrained by the transfer of activations, gradients, and optimizer states across accelerators. Low-bit quantizati…

cs.LG2026

Can Transformers Really Do It All? On the Compatibility of Inductive Biases Across Tasks

Damien Teney, Liangze Jiang, Hemanth Saratchandran +1

Transformers are remarkably versatile and their design is largely consistent across a variety of applications. But are they optimal for any given task or dataset? The answer may be…

cs.CV2021

PAUL: Procrustean Autoencoder for Unsupervised Lifting

Chaoyang Wang, Simon Lucey

Recent success in casting Non-rigid Structure from Motion (NRSfM) as an unsupervised deep learning problem has raised fundamental questions about what novelty in NRSfM prior could…

cs.CV2024

Structured Initialization for Attention in Vision Transformers

Jianqiao Zheng, Xueqian Li, Simon Lucey

The training of vision transformer (ViT) networks on small-scale datasets poses a significant challenge. By contrast, convolutional neural networks (CNNs) have an architectural ind…

cs.LG2025

Efficient Learning With Sine-Activated Low-rank Matrices

Yiping Ji, Hemanth Saratchandran, Cameron Gordon +2

Low-rank decomposition has emerged as a vital tool for enhancing parameter efficiency in neural network architectures, gaining traction across diverse applications in machine learn…

cs.CV2019

Web Stereo Video Supervision for Depth Prediction from Dynamic Scenes

Chaoyang Wang, Simon Lucey, Federico Perazzi +1

We present a fully data-driven method to compute depth from diverse monocular video sequences that contain large amounts of non-rigid objects, e.g., people. In order to learn recon…

cs.CV2025

Structured Initialization for Vision Transformers

Jianqiao Zheng, Xueqian Li, Hemanth Saratchandran +1

Convolutional Neural Networks (CNNs) inherently encode strong inductive biases, enabling effective generalization on small-scale datasets. In this paper, we propose integrating thi…

cs.CV2025

3D Gaussian Point Encoders

Jim James, Ben Wilson, Simon Lucey +1

In this work, we introduce the 3D Gaussian Point Encoder, an explicit per-point embedding built on mixtures of learned 3D Gaussians. This explicit geometric representation for 3D r…

cs.CV2026

Weight Conditioning for Smooth Optimization of Neural Networks

Hemanth Saratchandran, Thomas X. Wang, Simon Lucey

In this article, we introduce a novel normalization technique for neural network weight matrices, which we term weight conditioning. This approach aims to narrow the gap between th…

cs.CV2017

Learning Efficient Point Cloud Generation for Dense 3D Object Reconstruction

Chen-Hsuan Lin, Chen Kong, Simon Lucey

Conventional methods of 3D object generative modeling learn volumetric predictions using deep networks with 3D convolutional operations, which are direct analogies to classical 2D…

cs.LG2024

A Sampling Theory Perspective on Activations for Implicit Neural Representations

Hemanth Saratchandran, Sameera Ramasinghe, Violetta Shevchenko +2

Implicit Neural Representations (INRs) have gained popularity for encoding signals as compact, differentiable entities. While commonly using techniques like Fourier positional enco…

cs.LG2025

From Tables to Signals: Revealing Spectral Adaptivity in TabPFN

Jianqiao Zheng, Cameron Gordon, Yiping Ji +2

Task-agnostic tabular foundation models such as TabPFN have achieved impressive performance on tabular learning tasks, yet the origins of their inductive biases remain poorly under…

cs.CV2021

BARF: Bundle-Adjusting Neural Radiance Fields

Chen-Hsuan Lin, Wei-Chiu Ma, Antonio Torralba +1

Neural Radiance Fields (NeRF) have recently gained a surge of interest within the computer vision community for its power to synthesize photorealistic novel views of real-world sce…

cs.LG2023

On the effectiveness of neural priors in modeling dynamical systems

Sameera Ramasinghe, Hemanth Saratchandran, Violetta Shevchenko +1

Modelling dynamical systems is an integral component for understanding the natural world. To this end, neural networks are becoming an increasingly popular candidate owing to their…

cs.CV2021

On the Bias Against Inductive Biases

George Cazenavette, Simon Lucey

Borrowing from the transformer models that revolutionized the field of natural language processing, self-supervised feature learning for visual tasks has also seen state-of-the-art…

cs.CV2019

One Framework to Register Them All: PointNet Encoding for Point Cloud Alignment

Vinit Sarode, Xueqian Li, Hunter Goforth +5

PointNet has recently emerged as a popular representation for unstructured point cloud data, allowing application of deep learning to tasks such as object detection, segmentation a…

cs.LG2025

Cutting the Skip: Training Residual-Free Transformers

Yiping Ji, James Martens, Jianqiao Zheng +5

Transformers have achieved remarkable success across a wide range of applications, a feat often attributed to their scalability. Yet training them without skip (residual) connectio…

cs.CV2026

G3Splat: Geometrically Consistent Generalizable Gaussian Splatting

Mehdi Hosseinzadeh, Shin-Fang Chng, Yi Xu +3

3D Gaussians have become a powerful scene representation for real-time splatting and high-quality novel-view synthesis. This has motivated generalizable splatting -- methods that a…

cs.CV2025

Robust Physical Adversarial Patches Using Dynamically Optimized Clusters

Harrison Bagley, Will Meakin, Simon Lucey +2

Physical adversarial attacks on deep learning systems is concerning due to the ease of deploying such attacks, usually by placing an adversarial patch in a scene to manipulate the…

cs.CV2020

When to Use Convolutional Neural Networks for Inverse Problems

Nathaniel Chodosh, Simon Lucey

Reconstruction tasks in computer vision aim fundamentally to recover an undetermined signal from a set of noisy measurements. Examples include super-resolution, image denoising, an…

cs.CV2026

SineProject: Machine Unlearning for Stable Vision Language Alignment

Arpit Garg, Hemanth Saratchandran, Simon Lucey

Multimodal Large Language Models (MLLMs) increasingly need to forget specific knowledge such as unsafe or private information without requiring full retraining. However, existing u…

cs.CV2019

Photometric Mesh Optimization for Video-Aligned 3D Object Reconstruction

Chen-Hsuan Lin, Oliver Wang, Bryan C. Russell +4

In this paper, we address the problem of 3D object mesh reconstruction from RGB videos. Our approach combines the best of multi-view geometric and data-driven methods for 3D recons…

cs.CV2017

Deep-LK for Efficient Adaptive Object Tracking

Chaoyang Wang, Hamed Kiani Galoogahi, Chen-Hsuan Lin +1

In this paper we present a new approach for efficient regression based object tracking which we refer to as Deep- LK. Our approach is closely related to the Generic Object Tracking…

cs.LG2022

Beyond Periodicity: Towards a Unifying Framework for Activations in Coordinate-MLPs

Sameera Ramasinghe, Simon Lucey

Coordinate-MLPs are emerging as an effective tool for modeling multidimensional continuous signals, overcoming many drawbacks associated with discrete grid-based approximations. Ho…

cs.CV2020

MaskNet: A Fully-Convolutional Network to Estimate Inlier Points

Vinit Sarode, Animesh Dhagat, Rangaprasad Arun Srivatsan +3

Point clouds have grown in importance in the way computers perceive the world. From LIDAR sensors in autonomous cars and drones to the time of flight and stereo vision systems in o…

cs.LG2024

D'OH: Decoder-Only Random Hypernetworks for Implicit Neural Representations

Cameron Gordon, Lachlan Ewen MacDonald, Hemanth Saratchandran +1

Deep implicit functions have been found to be an effective tool for efficiently encoding all manner of natural signals. Their attractiveness stems from their ability to compactly r…

cs.LG2026

Stable Forgetting: Bounded Parameter-Efficient Unlearning in Foundation Models

Arpit Garg, Hemanth Saratchandran, Ravi Garg +1

Machine unlearning in foundation models (e.g., language and vision transformers) is essential for privacy and safety; however, existing approaches are unstable and unreliable. A wi…

cs.LG2026

Rethinking Attention: Polynomial Alternatives to Softmax in Transformers

Hemanth Saratchandran, Jianqiao Zheng, Yiping Ji +2

This paper questions whether the strong performance of softmax attention in transformers stems from producing a probability distribution over inputs. Instead, we argue that softmax…

cs.LG2022

On Regularizing Coordinate-MLPs

Sameera Ramasinghe, Lachlan MacDonald, Simon Lucey

We show that typical implicit regularization assumptions for deep neural networks (for regression) do not hold for coordinate-MLPs, a family of MLPs that are now ubiquitous in comp…

cs.LG2022

Reframing Neural Networks: Deep Structure in Overcomplete Representations

Calvin Murdock, George Cazenavette, Simon Lucey

In comparison to classical shallow representation learning techniques, deep neural networks have achieved superior performance in nearly every application benchmark. But despite th…

cs.LG2024

Architectural Strategies for the optimization of Physics-Informed Neural Networks

Hemanth Saratchandran, Shin-Fang Chng, Simon Lucey

Physics-informed neural networks (PINNs) offer a promising avenue for tackling both forward and inverse problems in partial differential equations (PDEs) by incorporating deep lear…

cs.CV2017

Semantic Photometric Bundle Adjustment on Natural Sequences

Rui Zhu, Chaoyang Wang, Chen-Hsuan Lin +2

The problem of obtaining dense reconstruction of an object in a natural sequence of images has been long studied in computer vision. Classically this problem has been solved throug…

cs.CV2017

Learning Depth from Monocular Videos using Direct Methods

Chaoyang Wang, Jose Miguel Buenaposada, Rui Zhu +1

The ability to predict depth from a single image - using recent advances in CNNs - is of increasing interest to the vision community. Unsupervised strategies to learning are partic…

cs.LG2025

Always Skip Attention

Yiping Ji, Hemanth Saratchandran, Peyman Moghadam +1

We highlight a curious empirical result within modern Vision Transformers (ViTs). Specifically, self-attention catastrophically fails to train unless it is used in conjunction with…

cs.CV2023

Flow supervision for Deformable NeRF

Chaoyang Wang, Lachlan Ewen MacDonald, Laszlo A. Jeni +1

In this paper we present a new method for deformable NeRF that can directly use optical flow as supervision. We overcome the major challenge with respect to the computationally ine…

cs.CV2014

Correlation Filters with Limited Boundaries

Hamed Kiani Galoogahi, Terence Sim, Simon Lucey

Correlation filters take advantage of specific properties in the Fourier domain allowing them to be estimated efficiently: O(NDlogD) in the frequency domain, versus O(D^3 + ND^2) s…

cs.LG2022

Evading the Simplicity Bias: Training a Diverse Set of Models Discovers Solutions with Superior OOD Generalization

Damien Teney, Ehsan Abbasnejad, Simon Lucey +1

Neural networks trained with SGD were recently shown to rely preferentially on linearly-predictive features and can ignore complex, equally-predictive ones. This simplicity bias ca…

cs.CV2017

Joint Max Margin and Semantic Features for Continuous Event Detection in Complex Scenes

Iman Abbasnejad, Sridha Sridharan, Simon Denman +2

In this paper the problem of complex event detection in the continuous domain (i.e. events with unknown starting and ending locations) is addressed. Existing event detection method…

cs.CV2019

Argoverse: 3D Tracking and Forecasting with Rich Maps

Ming-Fang Chang, John Lambert, Patsorn Sangkloy +8

We present Argoverse -- two datasets designed to support autonomous vehicle machine learning tasks such as 3D tracking and motion forecasting. Argoverse was collected by a fleet of…

cs.CV2023

Fast Neural Scene Flow

Xueqian Li, Jianqiao Zheng, Francesco Ferroni +2

Neural Scene Flow Prior (NSFP) is of significant interest to the vision community due to its inherent robustness to out-of-distribution (OOD) effects and its ability to deal with d…

cs.CV2020

Scene Flow from Point Clouds with or without Learning

Jhony Kaesemodel Pontes, James Hays, Simon Lucey

Scene flow is the three-dimensional (3D) motion field of a scene. It provides information about the spatial arrangement and rate of change of objects in dynamic environments. Curre…

cs.RO2016

Direct Visual Odometry using Bit-Planes

Hatem Alismail, Brett Browning, Simon Lucey

Feature descriptors, such as SIFT and ORB, are well-known for their robustness to illumination changes, which has made them popular for feature-based VSLAM\@. However, in degraded…

cs.CV2019

Deep Non-Rigid Structure from Motion

Chen Kong, Simon Lucey

Current non-rigid structure from motion (NRSfM) algorithms are mainly limited with respect to: (i) the number of images, and (ii) the type of shape variability they can handle. Thi…

cs.CV2017

Learning Policies for Adaptive Tracking with Deep Feature Cascades

Chen Huang, Simon Lucey, Deva Ramanan

Visual object tracking is a fundamental and time-critical vision task. Recent years have seen many shallow tracking methods based on real-time pixel-based correlation filters, as w…

cs.LG2026

Spectral Conditioning of Attention Improves Transformer Performance

Hemanth Saratchandran, Simon Lucey

We present a theoretical analysis of the Jacobian of an attention block within a transformer, showing that it is governed by the query, key, and value projections that define the a…

cs.CV2016

Inverse Compositional Spatial Transformer Networks

Chen-Hsuan Lin, Simon Lucey

In this paper, we establish a theoretical connection between the classical Lucas & Kanade (LK) algorithm and the emerging topic of Spatial Transformer Networks (STNs). STNs are of…

cs.CV2023

Robust Point Cloud Processing through Positional Embedding

Jianqiao Zheng, Xueqian Li, Sameera Ramasinghe +1

End-to-end trained per-point embeddings are an essential ingredient of any state-of-the-art 3D point cloud processing such as detection or alignment. Methods like PointNet, or the…

cs.CV2015

Dense Semantic Correspondence where Every Pixel is a Classifier

Hilton Bristow, Jack Valmadre, Simon Lucey

Determining dense semantic correspondences across objects and scenes is a difficult problem that underpins many higher-level computer vision algorithms. Unlike canonical dense corr…

cs.CV2021

Neural Scene Flow Prior

Xueqian Li, Jhony Kaesemodel Pontes, Simon Lucey

Before the deep learning revolution, many perception algorithms were based on runtime optimization in conjunction with a strong prior/regularization penalty. A prime example of thi…

cs.CV2021

PointNetLK Revisited

Xueqian Li, Jhony Kaesemodel Pontes, Simon Lucey

We address the generalization ability of recent learning-based point cloud registration methods. Despite their success, these approaches tend to have poor performance when applied…

cs.CV2020

Joint Pose and Shape Estimation of Vehicles from LiDAR Data

Hunter Goforth, Xiaoyan Hu, Michael Happold +1

We address the problem of estimating the pose and shape of vehicles from LiDAR scans, a common problem faced by the autonomous vehicle community. Recent work has tended to address…

cs.CV2019

PCRNet: Point Cloud Registration Network using PointNet Encoding

Vinit Sarode, Xueqian Li, Hunter Goforth +4

PointNet has recently emerged as a popular representation for unstructured point cloud data, allowing application of deep learning to tasks such as object detection, segmentation a…

cs.CV2015

Learning Temporal Alignment Uncertainty for Efficient Event Detection

Iman Abbasnejad, Sridha Sridharan, Simon Denman +2

In this paper we tackle the problem of efficient video event detection. We argue that linear detection functions should be preferred in this regard due to their scalability and eff…

cs.CV2025

Rethinking the Role of Spatial Mixing

George Cazenavette, Joel Julin, Simon Lucey

Until quite recently, the backbone of nearly every state-of-the-art computer vision model has been the 2D convolution. At its core, a 2D convolution simultaneously mixes informatio…

cs.LG2021

Rethinking Positional Encoding

Jianqiao Zheng, Sameera Ramasinghe, Simon Lucey

It is well noted that coordinate based MLPs benefit -- in terms of preserving high-frequency information -- through the encoding of coordinate positions as an array of Fourier feat…

cs.LG2020

Dataless Model Selection with the Deep Frame Potential

Calvin Murdock, Simon Lucey

Choosing a deep neural network architecture is a fundamental problem in applications that require balancing performance and parameter efficiency. Standard approaches rely on ad-hoc…

cs.CV2017

Proxy Templates for Inverse Compositional Photometric Bundle Adjustment

Christopher Ham, Simon Lucey, Surya Singh

Recent advances in 3D vision have demonstrated the strengths of photometric bundle adjustment. By directly minimizing reprojected pixel errors, instead of geometric reprojection er…

cs.CV2026

DARB-Splatting: Generalizing Splatting with Decaying Anisotropic Radial Basis Functions

Hashiru Pramuditha, Vinasirajan Viruthshaan, Vishagar Arunan +4

Splatting-based 3D reconstruction methods have gained popularity with the advent of 3D Gaussian Splatting, efficiently synthesizing high-quality novel views. These methods commonly…

cs.LG2023

On skip connections and normalisation layers in deep optimisation

Lachlan Ewen MacDonald, Jack Valmadre, Hemanth Saratchandran +1

We introduce a general theoretical framework, designed for the study of gradient optimisation of deep neural networks, that encompasses ubiquitous architecture choices including ba…

cs.CV2021

High Fidelity 3D Reconstructions with Limited Physical Views

Mosam Dabhi, Chaoyang Wang, Kunal Saluja +3

Multi-view triangulation is the gold standard for 3D reconstruction from 2D correspondences given known calibration and sufficient views. However in practice, expensive multi-view…

cs.CL2026

Mask the Target: A Plug-and-Play Regularizer Against LoRA Forgetting

Runze Xu, Arpit Garg, Hemanth Saratchandran +1

Low-Rank Adaptation (LoRA) has become one of the most widely used fine-tuning mechanisms for adapting large language models to new domains, tasks, and users. Yet adaptation perform…

cs.CV2025

Preconditioners for the Stochastic Training of Neural Fields

Shin-Fang Chng, Hemanth Saratchandran, Simon Lucey

Neural fields encode continuous multidimensional signals as neural networks, enabling diverse applications in computer vision, robotics, and geometry. While Adam is effective for s…

cs.LG2026

Memory Efficient Tabular Foundation Models

Shuting Luo, Monika Mikhail Kanaan, Cameron Gordon +2

The paper studies how to reduce the memory footprint of tabular foundation models like TabPFN using compression techniques, achieving up to 7.6× memory savings with little performa…

#tabular data#foundation models#model compression#memory efficiency