papers

Publications (32)

cs.CV2024

Zero-Shot Open-Vocabulary Tracking with Large Pre-Trained Models

Wen-Hsuan Chu, Adam W. Harley, Pavel Tokmakov +3

Object tracking is central to robot perception and scene understanding. Tracking-by-detection has long been a dominant paradigm for object tracking of specific object categories. R…

cs.CV2023

PointOdyssey: A Large-Scale Synthetic Dataset for Long-Term Point Tracking

Yang Zheng, Adam W. Harley, Bokui Shen +2

We introduce PointOdyssey, a large-scale synthetic dataset, and data generation framework, for the training and evaluation of long-term fine-grained tracking algorithms. Our goal i…

cs.CV2020

Tracking Emerges by Looking Around Static Scenes, with Neural 3D Mapping

Adam W. Harley, Shrinidhi K. Lakshmikanth, Paul Schydlo +1

We hypothesize that an agent that can look around in static scenes can learn rich visual representations applicable to 3D object tracking in complex dynamic scenes. We are motivate…

cs.CV2025

Animal Pose Labeling Using General-Purpose Point Trackers

Zhuoyang Pan, Boxiao Pan, Guandao Yang +2

Automatically estimating animal poses from videos is important for studying animal behaviors. Existing methods do not perform reliably since they are trained on datasets that are n…

cs.CV2024

PACE: A Large-Scale Dataset with Pose Annotations in Cluttered Environments

Yang You, Kai Xiong, Zhening Yang +7

We introduce PACE (Pose Annotations in Cluttered Environments), a large-scale benchmark designed to advance the development and evaluation of pose estimation methods in cluttered s…

cs.CV2025

Monocular Dynamic Gaussian Splatting: Fast, Brittle, and Scene Complexity Rules

Yiqing Liang, Mikhail Okunev, Mikaela Angelina Uy +4

Gaussian splatting methods are emerging as a popular approach for converting multi-view image data into scene representations that allow view synthesis. In particular, there is int…

cs.CV2015

Evaluation of Deep Convolutional Nets for Document Image Classification and Retrieval

Adam W. Harley, Alex Ufkes, Konstantinos G. Derpanis

This paper presents a new state-of-the-art for document image classification and retrieval, using features learned by deep convolutional neural networks (CNNs). In object and scene…

cs.CV2021

Track, Check, Repeat: An EM Approach to Unsupervised Tracking

Adam W. Harley, Yiming Zuo, Jing Wen +4

We propose an unsupervised method for detecting and tracking moving objects in 3D, in unlabelled RGB-D videos. The method begins with classic handcrafted techniques for segmenting…

cs.CV2022

Particle Video Revisited: Tracking Through Occlusions Using Point Trajectories

Adam W. Harley, Zhaoyuan Fang, Katerina Fragkiadaki

Tracking pixels in videos is typically studied as an optical flow estimation problem, where every pixel is described with a displacement vector that locates it in the next frame. E…

cs.CV2025

Generative Point Tracking with Flow Matching

Mattie Tesfaldet, Adam W. Harley, Konstantinos G. Derpanis +2

Tracking a point through a video can be a challenging task due to uncertainty arising from visual obfuscations, such as appearance changes and occlusions. Although current state-of…

cs.CV2025

AllTracker: Efficient Dense Point Tracking at High Resolution

Adam W. Harley, Yang You, Xinglong Sun +11

We introduce AllTracker: a model that estimates long-range point tracks by way of estimating the flow field between a query frame and every other frame of a video. Unlike existing…

cs.CV2017

Adversarial Inverse Graphics Networks: Learning 2D-to-3D Lifting and Image-to-Image Translation from Unpaired Supervision

Hsiao-Yu Fish Tung, Adam W. Harley, William Seto +1

Researchers have developed excellent feed-forward models that learn to map images to desired outputs, such as to the images' latent factors, or to other images, using supervised le…

cs.CV2024

Support-Set Context Matters for Bongard Problems

Nikhil Raghuraman, Adam W. Harley, Leonidas Guibas

Current machine learning methods struggle to solve Bongard problems, which are a type of IQ test that requires deriving an abstract "concept" from a set of positive and negative "s…

cs.CV2024

ODIN: A Single Model for 2D and 3D Segmentation

Ayush Jain, Pushkal Katara, Nikolaos Gkanatsios +5

State-of-the-art models on contemporary 3D segmentation benchmarks like ScanNet consume and label dataset-provided 3D point clouds, obtained through post processing of sensed multi…

cs.CV2021

Embodied Language Grounding with 3D Visual Feature Representations

Mihir Prabhudesai, Hsiao-Yu Fish Tung, Syed Ashar Javed +3

We propose associating language utterances to 3D visual abstractions of the scene they describe. The 3D visual abstractions are encoded as 3-dimensional visual feature maps. We inf…

cs.CV2017

Segmentation-Aware Convolutional Networks Using Local Attention Masks

Adam W. Harley, Konstantinos G. Derpanis, Iasonas Kokkinos

We introduce an approach to integrate segmentation information within a convolutional neural network (CNN). This counter-acts the tendency of CNNs to smooth information across regi…

cs.CV2016

Back to Basics: Unsupervised Learning of Optical Flow via Brightness Constancy and Motion Smoothness

Jason J. Yu, Adam W. Harley, Konstantinos G. Derpanis

Recently, convolutional networks (convnets) have proven useful for predicting optical flow. Much of this success is predicated on the availability of large datasets that require ex…

cs.CV2025

LookOut: Real-World Humanoid Egocentric Navigation

Boxiao Pan, Adam W. Harley, C. Karen Liu +1

The ability to predict collision-free future trajectories from egocentric observations is crucial in applications such as humanoid robotics, VR / AR, and assistive navigation. In t…

cs.CV2021

Move to See Better: Self-Improving Embodied Object Detection

Zhaoyuan Fang, Ayush Jain, Gabriel Sarch +2

Passive methods for object detection and segmentation treat images of the same scene as individual samples and do not exploit object permanence across multiple views. Generalizatio…

cs.CV2024

View-Consistent Hierarchical 3D Segmentation Using Ultrametric Feature Fields

Haodi He, Colton Stearns, Adam W. Harley +1

Large-scale vision foundation models such as Segment Anything (SAM) demonstrate impressive performance in zero-shot image segmentation at multiple levels of granularity. However, t…

cs.CV2021

CoCoNets: Continuous Contrastive 3D Scene Representations

Shamit Lal, Mihir Prabhudesai, Ishita Mediratta +2

This paper explores self-supervised learning of amodal 3D feature representations from RGB and RGB-D posed images and videos, agnostic to object and scene semantic content, and eva…

cs.CV2019

Image Disentanglement and Uncooperative Re-Entanglement for High-Fidelity Image-to-Image Translation

Adam W. Harley, Shih-En Wei, Jason Saragih +1

Cross-domain image-to-image translation should satisfy two requirements: (1) preserve the information that is common to both domains, and (2) generate convincing images covering va…

cs.CV2018

Reward Learning from Narrated Demonstrations

Hsiao-Yu Fish Tung, Adam W. Harley, Liang-Kang Huang +1

Humans effortlessly "program" one another by communicating goals and desires in natural language. In contrast, humans program robotic behaviours by indicating desired object locati…

cs.CV2024

Refining Pre-Trained Motion Models

Xinglong Sun, Adam W. Harley, Leonidas J. Guibas

Given the difficulty of manually annotating motion in video, the current best motion estimation methods are trained with synthetic data, and therefore struggle somewhat due to a tr…

cs.CV2025

PointSt3R: Point Tracking through 3D Grounded Correspondence

Rhodri Guerrier, Adam W. Harley, Dima Damen

Recent advances in foundational 3D reconstruction models, such as DUSt3R and MASt3R, have shown great potential in 2D and 3D correspondence in static scenes. In this paper, we prop…

cs.CV2022

TIDEE: Tidying Up Novel Rooms using Visuo-Semantic Commonsense Priors

Gabriel Sarch, Zhaoyuan Fang, Adam W. Harley +4

We introduce TIDEE, an embodied agent that tidies up a disordered scene based on learned commonsense object placement and room arrangement priors. TIDEE explores a home environment…

cs.CV2025

TAPIP3D: Tracking Any Point in Persistent 3D Geometry

Bowei Zhang, Lei Ke, Adam W. Harley +1

We introduce TAPIP3D, a novel approach for long-term 3D point tracking in monocular RGB and RGB-D videos. TAPIP3D represents videos as camera-stabilized spatio-temporal feature clo…

cs.CV2020

3D Object Recognition By Corresponding and Quantizing Neural 3D Scene Representations

Mihir Prabhudesai, Shamit Lal, Hsiao-Yu Fish Tung +3

We propose a system that learns to detect objects and infer their 3D poses in RGB-D images. Many existing systems can identify objects and infer 3D poses, but they heavily rely on…

cs.CV2024

EgoPoints: Advancing Point Tracking for Egocentric Videos

Ahmad Darkhalil, Rhodri Guerrier, Adam W. Harley +1

We introduce EgoPoints, a benchmark for point tracking in egocentric videos. We annotate 4.7K challenging tracks in egocentric sequences. Compared to the popular TAP-Vid-DAVIS eval…

cs.CV2020

Learning from Unlabelled Videos Using Contrastive Predictive Neural 3D Mapping

Adam W. Harley, Shrinidhi K. Lakshmikanth, Fangyu Li +3

Predictive coding theories suggest that the brain learns by predicting observations at various levels of abstraction. One of the most basic prediction tasks is view prediction: how…

cs.CV2022

Simple-BEV: What Really Matters for Multi-Sensor BEV Perception?

Adam W. Harley, Zhaoyuan Fang, Jie Li +2

Building 3D perception systems for autonomous vehicles that do not rely on high-density LiDAR is a critical research problem because of the expense of LiDAR systems compared to cam…

cs.CV2016

Learning Dense Convolutional Embeddings for Semantic Segmentation

Adam W. Harley, Konstantinos G. Derpanis, Iasonas Kokkinos

This paper proposes a new deep convolutional neural network (DCNN) architecture that learns pixel embeddings, such that pairwise distances between the embeddings can be used to inf…