papers

Publications (189)

cs.CV2026

Co-Me: Confidence-Guided Token Merging for Visual Geometric Transformers

Yutian Chen, Yuheng Qiu, Ruogu Li +4

We propose Confidence-Guided Token Merging (Co-Me), an acceleration mechanism for visual geometric transformers without retraining or finetuning the base model. Co-Me distilled a l…

cs.RO2022

TartanDrive: A Large-Scale Dataset for Learning Off-Road Dynamics Models

Samuel Triest, Matthew Sivaprakasam, Sean J. Wang +3

We present TartanDrive, a large scale dataset for learning dynamics models for off-road driving. We collected a dataset of roughly 200,000 off-road driving interactions on a modifi…

cs.RO2018

Joint Point Cloud and Image Based Localization For Efficient Inspection in Mixed Reality

Manash Pratim Das, Zhen Dong, Sebastian Scherer

This paper introduces a method of structure inspection using mixed-reality headsets to reduce the human effort in reporting accurate inspection information such as fault locations…

cs.CV2024

Geometry-Informed Distance Candidate Selection for Adaptive Lightweight Omnidirectional Stereo Vision with Fisheye Images

Conner Pulling, Je Hon Tan, Yaoyu Hu +1

Multi-view stereo omnidirectional distance estimation usually needs to build a cost volume with many hypothetical distance candidates. The cost volume building process is often com…

cs.RO2025

PIPE Planner: Pathwise Information Gain with Map Predictions for Indoor Robot Exploration

Seungjae Baek, Brady Moon, Seungchan Kim +4

Autonomous exploration in unknown environments requires estimating the information gain of an action to guide planning decisions. While prior approaches often compute information g…

cs.CV2020

ULSD: Unified Line Segment Detection across Pinhole, Fisheye, and Spherical Cameras

Hao Li, Huai Yu, Wen Yang +2

Line segment detection is essential for high-level tasks in computer vision and robotics. Currently, most stateof-the-art (SOTA) methods are dedicated to detecting straight line se…

cs.CV2019

Line-based Camera Pose Estimation in Point Cloud of Structured Environments

Huai Yu, Weikun Zhen, Wen Yang +1

Accurate registration of 2D imagery with point clouds is a key technology for image-LiDAR point cloud fusion, camera to laser scanner calibration and camera localization. Despite c…

cs.RO2025

SuperLoc: The Key to Robust LiDAR-Inertial Localization Lies in Predicting Alignment Risks

Shibo Zhao, Honghao Zhu, Yuanjun Gao +4

Map-based LiDAR localization, while widely used in autonomous systems, faces significant challenges in degraded environments due to lacking distinct geometric features. This paper…

cs.RO2017

Data-driven Planning via Imitation Learning

Sanjiban Choudhury, Mohak Bhardwaj, Sankalp Arora +4

Robot planning is the process of selecting a sequence of actions that optimize for a task specific objective. The optimal solutions to such tasks are heavily influenced by the impl…

cs.RO2024

SubT-MRS Dataset: Pushing SLAM Towards All-weather Environments

Shibo Zhao, Yuanjun Gao, Tianhao Wu +18

Simultaneous localization and mapping (SLAM) is a fundamental task for numerous applications such as autonomous navigation and exploration. Despite many SLAM datasets have been rel…

eess.SY2020

Automatic Real-time Anomaly Detection for Autonomous Aerial Vehicles

Azarakhsh Keipour, Mohammadreza Mousaei, Sebastian Scherer

The recent increase in the use of aerial vehicles raises concerns about the safety and reliability of autonomous operations. There is a growing need for methods to monitor the stat…

cs.RO2017

Adaptive Information Gathering via Imitation Learning

Sanjiban Choudhury, Ashish Kapoor, Gireeja Ranade +2

In the adaptive information gathering problem, a policy is required to select an informative sensing location using the history of measurements acquired thus far. While there is an…

cs.RO2023

Aerial Interaction with Tactile Sensing

Xiaofeng Guo, Guanqi He, Mohammadreza Mousaei +3

While autonomous Uncrewed Aerial Vehicles (UAVs) have grown rapidly, most applications only focus on passive visual tasks. Aerial interaction aims to execute tasks involving physic…

cs.RO2022

General Place Recognition Survey: Towards the Real-world Autonomy Age

Peng Yin, Shiqi Zhao, Ivan Cisneros +6

Place recognition is the fundamental module that can assist Simultaneous Localization and Mapping (SLAM) in loop-closure detection and re-localization for long-term navigation. The…

cs.LG2022

Lifelong Graph Learning

Chen Wang, Yuheng Qiu, Dasong Gao +1

Graph neural networks (GNN) are powerful models for many graph-structured tasks. Existing models often assume that the complete structure of the graph is available during training.…

cs.RO2023

UAS Simulator for Modeling, Analysis and Control in Free Flight and Physical Interaction

Azarakhsh Keipour, Mohammadreza Mousaei, Dongwei Bai +2

This paper presents the ARCAD simulator for the rapid development of Unmanned Aerial Systems (UAS), including underactuated and fully-actuated multirotors, fixed-wing aircraft, and…

cs.RO2022

AirLoop: Lifelong Loop Closure Detection

Dasong Gao, Chen Wang, Sebastian Scherer

Loop closure detection is an important building block that ensures the accuracy and robustness of simultaneous localization and mapping (SLAM) systems. Due to their generalization…

cs.RO2022

MUI-TARE: Multi-Agent Cooperative Exploration with Unknown Initial Position

Jingtian Yan, Xingqiao Lin, Zhongqiang Ren +6

Multi-agent exploration of a bounded 3D environment with unknown initial positions of agents is a challenging problem. It requires quickly exploring the environments as well as rob…

cs.RO2023

Image-based Visual Servo Control for Aerial Manipulation Using a Fully-Actuated UAV

Guanqi He, Yash Jangir, Junyi Geng +3

Using Unmanned Aerial Vehicles (UAVs) to perform high-altitude manipulation tasks beyond just passive visual application can reduce the time, cost, and risk of human workers. Prior…

cs.RO2021

Real-Time Ellipse Detection for Robotics Applications

Azarakhsh Keipour, Guilherme A. S. Pereira, Sebastian Scherer

We propose a new algorithm for real-time detection and tracking of elliptic patterns suitable for real-world robotics applications. The method fits ellipses to each contour in the…

cs.CV2025

Aug3D: Augmenting large scale outdoor datasets for Generalizable Novel View Synthesis

Aditya Rauniyar, Omar Alama, Silong Yong +2

Recent photorealistic Novel View Synthesis (NVS) advances have increasingly gained attention. However, these approaches remain constrained to small indoor scenes. While optimizatio…

cs.RO2021

VDB-EDT: An Efficient Euclidean Distance Transform Algorithm Based on VDB Data Structure

Delong Zhu, Chaoqun Wang, Wenshan Wang +3

This paper presents a fundamental algorithm, called VDB-EDT, for Euclidean distance transform (EDT) based on the VDB data structure. The algorithm executes on grid maps and generat…

cs.RO2026

Neuro-Symbolic Learning for Long-Horizon Task Planning Under Complex Logical Constraints

Qiwei Du, Zitong Zhan, Shaoshu Su +7

Task planning often suffers from severe efficiency bottlenecks when robots must reason over long-horizon action sequences under complex logical constraints, including object afford…

cs.CV2023

AirLoc: Object-based Indoor Relocalization

Aryan, Bowen Li, Sebastian Scherer +2

Indoor relocalization is vital for both robotic tasks like autonomous exploration and civil applications such as navigation with a cell phone in a shopping mall. Some previous appr…

cs.RO2025

General Place Recognition Survey: Towards Real-World Autonomy

Peng Yin, Jianhao Jiao, Shiqi Zhao +5

In the realm of robotics, the quest for achieving real-world autonomy, capable of executing large-scale and long-term operations, has positioned place recognition (PR) as a corners…

cs.RO2021

Super Odometry: IMU-centric LiDAR-Visual-Inertial Estimator for Challenging Environments

Shibo Zhao, Hengrui Zhang, Peng Wang +2

We propose Super Odometry, a high-precision multi-modal sensor fusion framework, providing a simple but effective way to fuse multiple sensors such as LiDAR, camera, and IMU sensor…

cs.CV2024

MAC-Ego3D: Multi-Agent Gaussian Consensus for Real-Time Collaborative Ego-Motion and Photorealistic 3D Reconstruction

Xiaohao Xu, Feng Xue, Shibo Zhao +3

Real-time multi-agent collaboration for ego-motion estimation and high-fidelity 3D reconstruction is vital for scalable spatial intelligence. However, traditional methods produce s…

cs.RO2023

Learning Risk-Aware Costmaps via Inverse Reinforcement Learning for Off-Road Navigation

Samuel Triest, Mateo Guaman Castro, Parv Maheshwari +3

The process of designing costmaps for off-road driving tasks is often a challenging and engineering-intensive task. Recent work in costmap design for off-road driving focuses on tr…

cs.RO2025

RayFronts: Open-Set Semantic Ray Frontiers for Online Scene Understanding and Exploration

Omar Alama, Avigyan Bhattacharya, Haoyang He +6

Open-set semantic mapping is crucial for open-world robots. Current mapping approaches either are limited by the depth range or only map beyond-range entities in constrained settin…

cs.RO2023

Off-Policy Evaluation with Online Adaptation for Robot Exploration in Challenging Environments

Yafei Hu, Junyi Geng, Chen Wang +2

Autonomous exploration has many important applications. However, classic information gain-based or frontier-based exploration only relies on the robot current state to determine th…

cs.RO2022

Challenges in Close-Proximity Safe and Seamless Operation of Manned and Unmanned Aircraft in Shared Airspace

Jay Patrikar, Joao P. A. Dantas, Sourish Ghosh +11

We propose developing an integrated system to keep autonomous unmanned aircraft safely separated and behave as expected in conjunction with manned traffic. The main goal is to achi…

cs.CV2022

TartanCalib: Iterative Wide-Angle Lens Calibration using Adaptive SubPixel Refinement of AprilTags

Bardienus P Duisterhof, Yaoyu Hu, Si Heng Teng +2

Wide-angle cameras are uniquely positioned for mobile robots, by virtue of the rich information they provide in a small, light, and cost-effective form factor. An accurate calibrat…

cs.CV2024

AirShot: Efficient Few-Shot Detection for Autonomous Exploration

Zihan Wang, Bowen Li, Chen Wang +1

Few-shot object detection has drawn increasing attention in the field of robotic exploration, where robots are required to find unseen objects with a few online provided examples.…

cs.RO2022

VTOL Failure Detection and Recovery by Utilizing Redundancy

Mohammadreza Mousaei, Azarakhsh Keipour, Junyi Geng +1

Offering vertical take-off and landing (VTOL) capabilities and the ability to travel great distances are crucial for Urban Air Mobility (UAM) vehicles. These capabilities make hybr…

cs.RO2022

Robust Modeling and Controls for Racing on the Edge

Joshua Spisak, Andrew Saba, Nayana Suvarna +5

Race cars are routinely driven to the edge of their handling limits in dynamic scenarios well above 200mph. Similar challenges are posed in autonomous racing, where a software stac…

cs.RO2018

Integrating kinematics and environment context into deep inverse reinforcement learning for predicting off-road vehicle trajectories

Yanfu Zhang, Wenshan Wang, Rogerio Bonatti +2

Predicting the motion of a mobile agent from a third-person perspective is an important component for many robotics applications, such as autonomous navigation and tracking. With a…

cs.LG2024

TartanAviation: Image, Speech, and ADS-B Trajectory Datasets for Terminal Airspace Operations

Jay Patrikar, Joao Dantas, Brady Moon +7

We introduce TartanAviation, an open-source multi-modal dataset focused on terminal-area airspace operations. TartanAviation provides a holistic view of the airport environment by…

cs.CV2020

Learning Visuomotor Policies for Aerial Navigation Using Cross-Modal Representations

Rogerio Bonatti, Ratnesh Madaan, Vibhav Vineet +2

Machines are a long way from robustly solving open-world perception-control tasks, such as first-person view (FPV) aerial navigation. While recent advances in end-to-end Machine Le…

cs.RO2025

Demonstrating ViSafe: Vision-enabled Safety for High-speed Detect and Avoid

Parv Kapoor, Ian Higgins, Nikhil Keetha +9

Assured safe-separation is essential for achieving seamless high-density operation of airborne vehicles in a shared airspace. To equip resource-constrained aerial systems with this…

cs.RO2026

Distilling Global Traversability Priors for Image-based Affordance Prediction in Off-road Environments

Matthew Sivaprakasam, Samuel Triest, Micah Nye +5

Standard methods for autonomous navigation in unstructured terrain are prone to myopic behaviors in long-horizon scenarios. The use of metric maps built from LiDAR or cameras provi…

cs.RO2022

BioSLAM: A Bio-inspired Lifelong Memory System for General Place Recognition

Peng Yin, Abulikemu Abuduweili, Shiqi Zhao +2

We present BioSLAM, a lifelong SLAM framework for learning various new appearances incrementally and maintaining accurate place recognition for previously visited areas. Unlike hum…

cs.RO2021

Unified Representation of Geometric Primitives for Graph-SLAM Optimization Using Decomposed Quadrics

Weikun Zhen, Huai Yu, Yaoyu Hu +1

In Simultaneous Localization And Mapping (SLAM) problems, high-level landmarks have the potential to build compact and informative maps compared to traditional point-based landmark…

cs.RO2022

Present and Future of SLAM in Extreme Underground Environments

Kamak Ebadi, Lukas Bernreiter, Harel Biggie +28

This paper reports on the state of the art in underground SLAM by discussing different SLAM strategies and results across six teams that participated in the three-year-long SubT co…

cs.RO2026

PRoID: Predicted Rate of Information Delivery in Multi-Robot Exploration and Relaying

Seungchan Kim, Seungjae Baek, Micah Corah +3

We address Multi-Robot Exploration and Relaying (MRER): a team of robots must explore an unknown environment and deliver acquired information to a fixed base station within a missi…

cs.CV2022

SphereVLAD++: Attention-based and Signal-enhanced Viewpoint Invariant Descriptor

Shiqi Zhao, Peng Yin, Ge Yi +1

LiDAR-based localization approach is a fundamental module for large-scale navigation tasks, such as last-mile delivery and autonomous driving, and localization robustness highly re…

cs.RO2019

LiDAR Enhanced Structure-from-Motion

Weikun Zhen, Yaoyu Hu, Huai Yu +1

Although Structure-from-Motion (SfM) as a maturing technique has been widely used in many applications, state-of-the-art SfM algorithms are still not robust enough in certain situa…

cs.CV2023

AnyLoc: Towards Universal Visual Place Recognition

Nikhil Keetha, Avneesh Mishra, Jay Karhade +4

Visual Place Recognition (VPR) is vital for robot localization. To date, the most performant VPR approaches are environment- and task-specific: while they exhibit strong performanc…

cs.CV2026

Unified Spherical Frontend: Learning Rotation-Equivariant Representations of Spherical Images from Any Camera

Mukai Yu, Mosam Dabhi, Liuyue Xie +2

Modern perception increasingly relies on fisheye, panoramic, and other wide field-of-view (FoV) cameras, yet most pipelines still apply planar CNNs designed for pinhole imagery on…

cs.RO2017

Near-Optimal Edge Evaluation in Explicit Generalized Binomial Graphs

Sanjiban Choudhury, Shervin Javdani, Siddhartha Srinivasa +1

Robotic motion-planning problems, such as a UAV flying fast in a partially-known environment or a robot arm moving around cluttered objects, require finding collision-free paths qu…

cs.CV2026

GrndCtrl: Grounding World Models via Self-Supervised Reward Alignment

Haoyang He, Jay Patrikar, Dong-Ki Kim +5

Recent advances in video world modeling have enabled large-scale generative models to simulate embodied environments with high visual fidelity, providing strong priors for predicti…

cs.RO2023

Time-Optimal Path Planning in a Constant Wind for Uncrewed Aerial Vehicles using Dubins Set Classification

Brady Moon, Sagar Sachdev, Junbin Yuan +1

Time-optimal path planning in high winds for a turning-rate constrained UAV is a challenging problem to solve and is important for deployment and field operations. Previous works h…

cs.RO2023

AutoMerge: A Framework for Map Assembling and Smoothing in City-scale Environments

Peng Yin, Haowen Lai, Shiqi Zhao +4

We present AutoMerge, a LiDAR data processing framework for assembling a large number of map segments into a complete map. Traditional large-scale map merging methods are fragile t…

cs.RO2026

SyNeT: Synthetic Negatives for Traversability Learning

Bomena Kim, Hojun Lee, Younsoo Park +3

Reliable traversability estimation is crucial for autonomous robots to navigate complex outdoor environments safely. Existing self-supervised learning frameworks primarily rely on…

cs.CV2026

AnyThermal: Towards Learning Universal Representations for Thermal Perception

Parv Maheshwari, Jay Karhade, Yogesh Chawla +8

We present AnyThermal, a thermal backbone that captures robust task-agnostic thermal features suitable for a variety of tasks such as cross-modal place recognition, thermal segment…

cs.RO2026

IA-TIGRIS: An Incremental and Adaptive Sampling-Based Planner for Online Informative Path Planning

Brady Moon, Nayana Suvarna, Andrew Jong +4

Planning paths that maximize information gain for robotic platforms has wide-ranging applications and significant potential impact. To effectively adapt to real-time data collectio…

cs.RO2024

UNRealNet: Learning Uncertainty-Aware Navigation Features from High-Fidelity Scans of Real Environments

Samuel Triest, David D. Fan, Sebastian Scherer +1

Traversability estimation in rugged, unstructured environments remains a challenging problem in field robotics. Often, the need for precise, accurate traversability estimation is i…

cs.RO2021

3D Human Reconstruction in the Wild with Collaborative Aerial Cameras

Cherie Ho, Andrew Jong, Harry Freeman +3

Aerial vehicles are revolutionizing applications that require capturing the 3D structure of dynamic targets in the wild, such as sports, medicine, and entertainment. The core chall…

cs.RO2021

Toward Efficient and Robust Multiple Camera Visual-inertial Odometry

Yao He, Huai Yu, Wen Yang +1

Efficiency and robustness are the essential criteria for the visual-inertial odometry (VIO) system. To process massive visual data, the high cost on CPU resources and computation l…

cs.LG2025

Amelia: A Large Dataset and Benchmark for Airport Surface Movement Forecasting

Ingrid Navarro, Pablo Ortega-Kral, Jay Patrikar +6

Demand for air travel is rising, straining existing aviation infrastructure. In the US, more than 90% of airport control towers are understaffed, falling short of FAA and union sta…

cs.CV2022

360FusionNeRF: Panoramic Neural Radiance Fields with Joint Guidance

Shreyas Kulkarni, Peng Yin, Sebastian Scherer

We present a method to synthesize novel views from a single panorama image based on the neural radiance field (NeRF). Prior studies in a similar setting rely on the nei…

cs.RO2019

Monocular Object and Plane SLAM in Structured Environments

Shichao Yang, Sebastian Scherer

In this paper, we present a monocular Simultaneous Localization and Mapping (SLAM) algorithm using high-level object and plane landmarks. The built map is denser, more compact and…

cs.RO2023

Design, Modeling and Control for a Tilt-rotor VTOL UAV in the Presence of Actuator Failure

Mohammadreza Mousaei, Junyi Geng, Azarakhsh Keipour +2

Enabling vertical take-off and landing while providing the ability to fly long ranges opens the door to a wide range of new real-world aircraft applications while improving many ex…

cs.CV2020

Monocular Camera Localization in Prior LiDAR Maps with 2D-3D Line Correspondences

Huai Yu, Weikun Zhen, Wen Yang +2

Light-weight camera localization in existing maps is essential for vision-based navigation. Currently, visual and visual-inertial odometry (VO\&VIO) techniques are well-developed f…

cs.CV2021

ORStereo: Occlusion-Aware Recurrent Stereo Matching for 4K-Resolution Images

Yaoyu Hu, Wenshan Wang, Huai Yu +2

Stereo reconstruction models trained on small images do not generalize well to high-resolution data. Training a model on high-resolution image size faces difficulties of data avail…

cs.RO2021

Adaptive Safety Margin Estimation for Safe Real-Time Replanning under Time-Varying Disturbance

Cherie Ho, Jay Patrikar, Rogerio Bonatti +1

Safe navigation in real-time is challenging because engineers need to work with uncertain vehicle dynamics, variable external disturbances, and imperfect controllers. A common safe…

cs.RO2019

CubeSLAM: Monocular 3D Object SLAM

Shichao Yang, Sebastian Scherer

We present a method for single image 3D cuboid object detection and multi-view object SLAM in both static and dynamic environments, and demonstrate that the two parts can improve e…

cs.RO2025

AutoODD: Agentic Audits via Bayesian Red Teaming in Black-Box Models

Rebecca Martin, Jay Patrikar, Sebastian Scherer

Specialized machine learning models, regardless of architecture and training, are susceptible to failures in deployment. With their increasing use in high risk situations, the abil…

cs.RO2026

Don't Fool Me Twice: Adapting to Adversity in the Wild with Experience-Driven Reasoning

Navin Sriram Ravie, Andrew Jong, Krrish Jain +4

In robotics, dangers and adversity modes are often embodiment-specific and relative to each agent. A frontier of autonomous mobile robotics is to enable agents to operate effective…

cs.CV2020

RGB-D SLAM in Dynamic Environments Using Point Correlations

Weichen Dai, Yu Zhang, Ping Li +2

In this paper, a simultaneous localization and mapping (SLAM) method that eliminates the influence of moving objects in dynamic environments is proposed. This method utilizes the c…

cs.RO2018

Autonomous drone cinematographer: Using artistic principles to create smooth, safe, occlusion-free trajectories for aerial filming

Rogerio Bonatti, Yanfu Zhang, Sanjiban Choudhury +2

Autonomous aerial cinematography has the potential to enable automatic capture of aesthetically pleasing videos without requiring human intervention, empowering individuals with th…

cs.RO2026

Learning Energy-Efficient Air--Ground Actuation for Hybrid Robots on Stair-Like Terrain

Jiaxing Li, Wen Tian, Xinhang Xu +3

Hybrid aerial--ground robots offer both traversability and endurance, but stair-like discontinuities create a trade-off: wheels alone often stall at edges, while flight is energy-h…

cs.CV2021

3D-SiamRPN: An End-to-End Learning Method for Real-Time 3D Single Object Tracking Using Raw Point Cloud

Zheng Fang, Sifan Zhou, Yubo Cui +1

3D single object tracking is a key issue for autonomous following robot, where the robot should robustly track and accurately localize the target for efficient following. In this p…

cs.RO2025

AirIO: Learning Inertial Odometry with Enhanced IMU Feature Observability

Yuheng Qiu, Can Xu, Yutian Chen +3

Inertial odometry (IO) using only Inertial Measurement Units (IMUs) offers a lightweight and cost-effective solution for Unmanned Aerial Vehicle (UAV) applications, yet existing le…

cs.CV2019

Improved Generalization of Heading Direction Estimation for Aerial Filming Using Semi-supervised Regression

Wenshan Wang, Aayush Ahuja, Yanfu Zhang +2

In the task of Autonomous aerial filming of a moving actor (e.g. a person or a vehicle), it is crucial to have a good heading direction estimation for the actor from the visual inp…

cs.CV2022

Attention-Enhanced Cross-modal Localization Between 360 Images and Point Clouds

Zhipeng Zhao, Huai Yu, Chenwei Lyv +2

Visual localization plays an important role for intelligent robots and autonomous driving, especially when the accuracy of GNSS is unreliable. Recently, camera localization in LiDA…

cs.RO2021

In-flight positional and energy use data set of a DJI Matrice 100 quadcopter for small package delivery

Thiago A. Rodrigues, Jay Patrikar, Arnav Choudhry +9

We autonomously direct a small quadcopter package delivery Uncrewed Aerial Vehicle (UAV) or "drone" to take off, fly a specified route, and land for a total of 209 flights while va…

cs.RO2023

Follow The Rules: Online Signal Temporal Logic Tree Search for Guided Imitation Learning in Stochastic Domains

Jasmine Jerry Aloor, Jay Patrikar, Parv Kapoor +2

Seamlessly integrating rules in Learning-from-Demonstrations (LfD) policies is a critical requirement to enable the real-world deployment of AI agents. Recently, Signal Temporal Lo…

cs.RO2025

MAC-VO: Metrics-aware Covariance for Learning-based Stereo Visual Odometry

Yuheng Qiu, Yutian Chen, Zihao Zhang +2

We propose the MAC-VO, a novel learning-based stereo VO that leverages the learned metrics-aware matching uncertainty for dual purposes: selecting keypoint and weighing the residua…

cs.RO2022

AirDOS: Dynamic SLAM benefits from Articulated Objects

Yuheng Qiu, Chen Wang, Wenshan Wang +2

Dynamic Object-aware SLAM (DOS) exploits object-level information to enable robust motion estimation in dynamic environments. Existing methods mainly focus on identifying and exclu…

cs.RO2018

Monocular and Stereo Cues for Landing Zone Evaluation for Micro UAVs

Rohit Garg, Shichao Yang, Sebastian Scherer

Autonomous and safe landing is important for unmanned aerial vehicles. We present a monocular and stereo image based method for fast and accurate landing zone evaluation for UAVs i…

cs.RO2026

UMI-on-Air: Embodiment-Aware Guidance for Embodiment-Agnostic Visuomotor Policies

Harsh Gupta, Xiaofeng Guo, Huy Ha +6

We introduce UMI-on-Air, a framework for embodiment-aware deployment of embodiment-agnostic manipulation policies. Our approach leverages diverse, unconstrained human demonstration…

eess.SY2020

ALFA: A Dataset for UAV Fault and Anomaly Detection

Azarakhsh Keipour, Mohammadreza Mousaei, Sebastian Scherer

We present a dataset of several fault types in control surfaces of a fixed-wing Unmanned Aerial Vehicle (UAV) for use in Fault Detection and Isolation (FDI) and Anomaly Detection (…

cs.RO2024

Informative Sensor Planning for a Single-Axis Gimbaled Camera on a Fixed-Wing UAV

Aditya Parandekar, Brady Moon, Nayana Suvarna +1

Uncrewed Aerial Vehicles (UAVs) are a leading choice of platforms for a variety of information-gathering applications. Sensor planning can enhance the efficiency and success of the…

cs.RO2023

PIAug -- Physics Informed Augmentation for Learning Vehicle Dynamics for Off-Road Navigation

Parv Maheshwari, Wenshan Wang, Samuel Triest +5

Modeling the precise dynamics of off-road vehicles is a complex yet essential task due to the challenging terrain they encounter and the need for optimal performance and safety. Re…

cs.RO2018

A Unified 3D Mapping Framework using a 3D or 2D LiDAR

Weikun Zhen, Sebastian Scherer

Simultaneous Localization and Mapping (SLAM) has been considered as a solved problem thanks to the progress made in the past few years. However, the great majority of LiDAR-based S…

cs.CV2021

AdaFusion: Visual-LiDAR Fusion with Adaptive Weights for Place Recognition

Haowen Lai, Peng Yin, Sebastian Scherer

Recent years have witnessed the increasing application of place recognition in various environments, such as city roads, large buildings, and a mix of indoor and outdoor places. Th…

cs.RO2025

TartanGround: A Large-Scale Dataset for Ground Robot Perception and Navigation

Manthan Patel, Fan Yang, Yuheng Qiu +4

We present TartanGround, a large-scale, multi-modal dataset to advance the perception and autonomy of ground robots operating in diverse environments. This dataset, collected in va…

cs.RO2022

Visual Servoing Approach for Autonomous UAV Landing on a Moving Vehicle

Azarakhsh Keipour, Guilherme A. S. Pereira, Rogerio Bonatti +4

Many aerial robotic applications require the ability to land on moving platforms, such as delivery trucks and marine research boats. We present a method to autonomously land an Unm…

cs.RO2025

MapEx: Indoor Structure Exploration with Probabilistic Information Gain from Global Map Predictions

Cherie Ho, Seungchan Kim, Brady Moon +6

Exploration is a critical challenge in robotics, centered on understanding unknown environments. In this work, we focus on robots exploring structured indoor environments which are…

cs.CV2025

Any4D: Unified Feed-Forward Metric 4D Reconstruction

Jay Karhade, Nikhil Keetha, Yuchen Zhang +4

We present Any4D, a scalable multi-view transformer for metric-scale, dense feed-forward 4D reconstruction. Any4D directly generates per-pixel motion and geometry predictions for N…

cs.RO2024

Learning-on-the-Drive: Self-supervised Adaptation of Visual Offroad Traversability Models

Eric Chen, Cherie Ho, Mukhtar Maulimov +2

Autonomous offroad driving is essential for applications like emergency rescue, military operations, and agriculture. Despite progress, systems struggle with high-speed vehicles ex…

cs.RO2024

SoRTS: Learned Tree Search for Long Horizon Social Robot Navigation

Ingrid Navarro, Jay Patrikar, Joao P. A. Dantas +4

The fast-growing demand for fully autonomous robots in shared spaces calls for the development of trustworthy agents that can safely and seamlessly navigate in crowded environments…

cs.RO2025

Unifying Deep Predicate Invention with Pre-trained Foundation Models

Qianwei Wang, Bowen Li, Zhanpeng Luo +6

Long-horizon robotic tasks are hard due to continuous state-action spaces and sparse feedback. Symbolic world models help by decomposing tasks into discrete predicates that capture…

cs.AI2025

LogiCity: Advancing Neuro-Symbolic AI with Abstract Urban Simulation

Bowen Li, Zhaoyu Li, Qiwei Du +10

Recent years have witnessed the rapid development of Neuro-Symbolic (NeSy) AI systems, which integrate symbolic reasoning into deep neural networks. However, most of the existing b…

cs.CV2024

SplaTAM: Splat, Track & Map 3D Gaussians for Dense RGB-D SLAM

Nikhil Keetha, Jay Karhade, Krishna Murthy Jatavallabhula +4

Dense simultaneous localization and mapping (SLAM) is crucial for robotics and augmented reality applications. However, current methods are often hampered by the non-volumetric or…

cs.CV2020

TartanVO: A Generalizable Learning-based VO

Wenshan Wang, Yaoyu Hu, Sebastian Scherer

We present the first learning-based visual odometry (VO) model, which generalizes to multiple datasets and real-world scenarios and outperforms geometry-based methods in challengin…

cs.RO2022

AirCode: A Robust Object Encoding Method

Kuan Xu, Chen Wang, Chao Chen +2

Object encoding and identification are crucial for many robotic tasks such as autonomous exploration and semantic relocalization. Existing works heavily rely on the tracking of det…

cs.CV2021

3D Segmentation Learning from Sparse Annotations and Hierarchical Descriptors

Peng Yin, Lingyun Xu, Jianmin Ji +2

One of the main obstacles to 3D semantic segmentation is the significant amount of endeavor required to generate expensive point-wise annotations for fully supervised training. To…

cs.RO2023

2D-3D Pose Tracking with Multi-View Constraints

Huai Yu, Kuangyi Chen, Wen Yang +2

Camera localization in 3D LiDAR maps has gained increasing attention due to its promising ability to handle complex scenarios, surpassing the limitations of visual-only localizatio…

cs.RO2020

A Robust Laser-Inertial Odometry and Mapping Method for Large-Scale Highway Environments

Shibo Zhao, Zheng Fang, HaoLai Li +1

In this paper, we propose a novel laser-inertial odometry and mapping method to achieve real-time, low-drift and robust pose estimation in large-scale highway environments. The pro…