Publications (32)
Zero-Shot Open-Vocabulary Tracking with Large Pre-Trained Models
Wen-Hsuan Chu, Adam W. Harley, Pavel Tokmakov +3
Object tracking is central to robot perception and scene understanding. Tracking-by-detection has long been a dominant paradigm for object tracking of specific object categories. R…
PointOdyssey: A Large-Scale Synthetic Dataset for Long-Term Point Tracking
Yang Zheng, Adam W. Harley, Bokui Shen +2
We introduce PointOdyssey, a large-scale synthetic dataset, and data generation framework, for the training and evaluation of long-term fine-grained tracking algorithms. Our goal i…
Tracking Emerges by Looking Around Static Scenes, with Neural 3D Mapping
Adam W. Harley, Shrinidhi K. Lakshmikanth, Paul Schydlo +1
We hypothesize that an agent that can look around in static scenes can learn rich visual representations applicable to 3D object tracking in complex dynamic scenes. We are motivate…
Animal Pose Labeling Using General-Purpose Point Trackers
Zhuoyang Pan, Boxiao Pan, Guandao Yang +2
Automatically estimating animal poses from videos is important for studying animal behaviors. Existing methods do not perform reliably since they are trained on datasets that are n…
PACE: A Large-Scale Dataset with Pose Annotations in Cluttered Environments
Yang You, Kai Xiong, Zhening Yang +7
We introduce PACE (Pose Annotations in Cluttered Environments), a large-scale benchmark designed to advance the development and evaluation of pose estimation methods in cluttered s…
Monocular Dynamic Gaussian Splatting: Fast, Brittle, and Scene Complexity Rules
Yiqing Liang, Mikhail Okunev, Mikaela Angelina Uy +4
Gaussian splatting methods are emerging as a popular approach for converting multi-view image data into scene representations that allow view synthesis. In particular, there is int…
Evaluation of Deep Convolutional Nets for Document Image Classification and Retrieval
Adam W. Harley, Alex Ufkes, Konstantinos G. Derpanis
This paper presents a new state-of-the-art for document image classification and retrieval, using features learned by deep convolutional neural networks (CNNs). In object and scene…
Track, Check, Repeat: An EM Approach to Unsupervised Tracking
Adam W. Harley, Yiming Zuo, Jing Wen +4
We propose an unsupervised method for detecting and tracking moving objects in 3D, in unlabelled RGB-D videos. The method begins with classic handcrafted techniques for segmenting…
Particle Video Revisited: Tracking Through Occlusions Using Point Trajectories
Adam W. Harley, Zhaoyuan Fang, Katerina Fragkiadaki
Tracking pixels in videos is typically studied as an optical flow estimation problem, where every pixel is described with a displacement vector that locates it in the next frame. E…
Generative Point Tracking with Flow Matching
Mattie Tesfaldet, Adam W. Harley, Konstantinos G. Derpanis +2
Tracking a point through a video can be a challenging task due to uncertainty arising from visual obfuscations, such as appearance changes and occlusions. Although current state-of…
AllTracker: Efficient Dense Point Tracking at High Resolution
Adam W. Harley, Yang You, Xinglong Sun +11
We introduce AllTracker: a model that estimates long-range point tracks by way of estimating the flow field between a query frame and every other frame of a video. Unlike existing…
Adversarial Inverse Graphics Networks: Learning 2D-to-3D Lifting and Image-to-Image Translation from Unpaired Supervision
Hsiao-Yu Fish Tung, Adam W. Harley, William Seto +1
Researchers have developed excellent feed-forward models that learn to map images to desired outputs, such as to the images' latent factors, or to other images, using supervised le…
Support-Set Context Matters for Bongard Problems
Nikhil Raghuraman, Adam W. Harley, Leonidas Guibas
Current machine learning methods struggle to solve Bongard problems, which are a type of IQ test that requires deriving an abstract "concept" from a set of positive and negative "s…
ODIN: A Single Model for 2D and 3D Segmentation
Ayush Jain, Pushkal Katara, Nikolaos Gkanatsios +5
State-of-the-art models on contemporary 3D segmentation benchmarks like ScanNet consume and label dataset-provided 3D point clouds, obtained through post processing of sensed multi…
Embodied Language Grounding with 3D Visual Feature Representations
Mihir Prabhudesai, Hsiao-Yu Fish Tung, Syed Ashar Javed +3
We propose associating language utterances to 3D visual abstractions of the scene they describe. The 3D visual abstractions are encoded as 3-dimensional visual feature maps. We inf…
Segmentation-Aware Convolutional Networks Using Local Attention Masks
Adam W. Harley, Konstantinos G. Derpanis, Iasonas Kokkinos
We introduce an approach to integrate segmentation information within a convolutional neural network (CNN). This counter-acts the tendency of CNNs to smooth information across regi…
Back to Basics: Unsupervised Learning of Optical Flow via Brightness Constancy and Motion Smoothness
Jason J. Yu, Adam W. Harley, Konstantinos G. Derpanis
Recently, convolutional networks (convnets) have proven useful for predicting optical flow. Much of this success is predicated on the availability of large datasets that require ex…
LookOut: Real-World Humanoid Egocentric Navigation
Boxiao Pan, Adam W. Harley, C. Karen Liu +1
The ability to predict collision-free future trajectories from egocentric observations is crucial in applications such as humanoid robotics, VR / AR, and assistive navigation. In t…
Move to See Better: Self-Improving Embodied Object Detection
Zhaoyuan Fang, Ayush Jain, Gabriel Sarch +2
Passive methods for object detection and segmentation treat images of the same scene as individual samples and do not exploit object permanence across multiple views. Generalizatio…
View-Consistent Hierarchical 3D Segmentation Using Ultrametric Feature Fields
Haodi He, Colton Stearns, Adam W. Harley +1
Large-scale vision foundation models such as Segment Anything (SAM) demonstrate impressive performance in zero-shot image segmentation at multiple levels of granularity. However, t…
CoCoNets: Continuous Contrastive 3D Scene Representations
Shamit Lal, Mihir Prabhudesai, Ishita Mediratta +2
This paper explores self-supervised learning of amodal 3D feature representations from RGB and RGB-D posed images and videos, agnostic to object and scene semantic content, and eva…
Image Disentanglement and Uncooperative Re-Entanglement for High-Fidelity Image-to-Image Translation
Adam W. Harley, Shih-En Wei, Jason Saragih +1
Cross-domain image-to-image translation should satisfy two requirements: (1) preserve the information that is common to both domains, and (2) generate convincing images covering va…
Reward Learning from Narrated Demonstrations
Hsiao-Yu Fish Tung, Adam W. Harley, Liang-Kang Huang +1
Humans effortlessly "program" one another by communicating goals and desires in natural language. In contrast, humans program robotic behaviours by indicating desired object locati…
Refining Pre-Trained Motion Models
Xinglong Sun, Adam W. Harley, Leonidas J. Guibas
Given the difficulty of manually annotating motion in video, the current best motion estimation methods are trained with synthetic data, and therefore struggle somewhat due to a tr…
PointSt3R: Point Tracking through 3D Grounded Correspondence
Rhodri Guerrier, Adam W. Harley, Dima Damen
Recent advances in foundational 3D reconstruction models, such as DUSt3R and MASt3R, have shown great potential in 2D and 3D correspondence in static scenes. In this paper, we prop…
TIDEE: Tidying Up Novel Rooms using Visuo-Semantic Commonsense Priors
Gabriel Sarch, Zhaoyuan Fang, Adam W. Harley +4
We introduce TIDEE, an embodied agent that tidies up a disordered scene based on learned commonsense object placement and room arrangement priors. TIDEE explores a home environment…
TAPIP3D: Tracking Any Point in Persistent 3D Geometry
Bowei Zhang, Lei Ke, Adam W. Harley +1
We introduce TAPIP3D, a novel approach for long-term 3D point tracking in monocular RGB and RGB-D videos. TAPIP3D represents videos as camera-stabilized spatio-temporal feature clo…
3D Object Recognition By Corresponding and Quantizing Neural 3D Scene Representations
Mihir Prabhudesai, Shamit Lal, Hsiao-Yu Fish Tung +3
We propose a system that learns to detect objects and infer their 3D poses in RGB-D images. Many existing systems can identify objects and infer 3D poses, but they heavily rely on…
EgoPoints: Advancing Point Tracking for Egocentric Videos
Ahmad Darkhalil, Rhodri Guerrier, Adam W. Harley +1
We introduce EgoPoints, a benchmark for point tracking in egocentric videos. We annotate 4.7K challenging tracks in egocentric sequences. Compared to the popular TAP-Vid-DAVIS eval…
Learning from Unlabelled Videos Using Contrastive Predictive Neural 3D Mapping
Adam W. Harley, Shrinidhi K. Lakshmikanth, Fangyu Li +3
Predictive coding theories suggest that the brain learns by predicting observations at various levels of abstraction. One of the most basic prediction tasks is view prediction: how…
Simple-BEV: What Really Matters for Multi-Sensor BEV Perception?
Adam W. Harley, Zhaoyuan Fang, Jie Li +2
Building 3D perception systems for autonomous vehicles that do not rely on high-density LiDAR is a critical research problem because of the expense of LiDAR systems compared to cam…
Learning Dense Convolutional Embeddings for Semantic Segmentation
Adam W. Harley, Konstantinos G. Derpanis, Iasonas Kokkinos
This paper proposes a new deep convolutional neural network (DCNN) architecture that learns pixel embeddings, such that pairwise distances between the embeddings can be used to inf…