Publications (189)
Co-Me: Confidence-Guided Token Merging for Visual Geometric Transformers
Yutian Chen, Yuheng Qiu, Ruogu Li +4
We propose Confidence-Guided Token Merging (Co-Me), an acceleration mechanism for visual geometric transformers without retraining or finetuning the base model. Co-Me distilled a l…
TartanDrive: A Large-Scale Dataset for Learning Off-Road Dynamics Models
Samuel Triest, Matthew Sivaprakasam, Sean J. Wang +3
We present TartanDrive, a large scale dataset for learning dynamics models for off-road driving. We collected a dataset of roughly 200,000 off-road driving interactions on a modifi…
Joint Point Cloud and Image Based Localization For Efficient Inspection in Mixed Reality
Manash Pratim Das, Zhen Dong, Sebastian Scherer
This paper introduces a method of structure inspection using mixed-reality headsets to reduce the human effort in reporting accurate inspection information such as fault locations…
Geometry-Informed Distance Candidate Selection for Adaptive Lightweight Omnidirectional Stereo Vision with Fisheye Images
Conner Pulling, Je Hon Tan, Yaoyu Hu +1
Multi-view stereo omnidirectional distance estimation usually needs to build a cost volume with many hypothetical distance candidates. The cost volume building process is often com…
PIPE Planner: Pathwise Information Gain with Map Predictions for Indoor Robot Exploration
Seungjae Baek, Brady Moon, Seungchan Kim +4
Autonomous exploration in unknown environments requires estimating the information gain of an action to guide planning decisions. While prior approaches often compute information g…
ULSD: Unified Line Segment Detection across Pinhole, Fisheye, and Spherical Cameras
Hao Li, Huai Yu, Wen Yang +2
Line segment detection is essential for high-level tasks in computer vision and robotics. Currently, most stateof-the-art (SOTA) methods are dedicated to detecting straight line se…
Line-based Camera Pose Estimation in Point Cloud of Structured Environments
Huai Yu, Weikun Zhen, Wen Yang +1
Accurate registration of 2D imagery with point clouds is a key technology for image-LiDAR point cloud fusion, camera to laser scanner calibration and camera localization. Despite c…
SuperLoc: The Key to Robust LiDAR-Inertial Localization Lies in Predicting Alignment Risks
Shibo Zhao, Honghao Zhu, Yuanjun Gao +4
Map-based LiDAR localization, while widely used in autonomous systems, faces significant challenges in degraded environments due to lacking distinct geometric features. This paper…
Data-driven Planning via Imitation Learning
Sanjiban Choudhury, Mohak Bhardwaj, Sankalp Arora +4
Robot planning is the process of selecting a sequence of actions that optimize for a task specific objective. The optimal solutions to such tasks are heavily influenced by the impl…
SubT-MRS Dataset: Pushing SLAM Towards All-weather Environments
Shibo Zhao, Yuanjun Gao, Tianhao Wu +18
Simultaneous localization and mapping (SLAM) is a fundamental task for numerous applications such as autonomous navigation and exploration. Despite many SLAM datasets have been rel…
Automatic Real-time Anomaly Detection for Autonomous Aerial Vehicles
Azarakhsh Keipour, Mohammadreza Mousaei, Sebastian Scherer
The recent increase in the use of aerial vehicles raises concerns about the safety and reliability of autonomous operations. There is a growing need for methods to monitor the stat…
Adaptive Information Gathering via Imitation Learning
Sanjiban Choudhury, Ashish Kapoor, Gireeja Ranade +2
In the adaptive information gathering problem, a policy is required to select an informative sensing location using the history of measurements acquired thus far. While there is an…
Aerial Interaction with Tactile Sensing
Xiaofeng Guo, Guanqi He, Mohammadreza Mousaei +3
While autonomous Uncrewed Aerial Vehicles (UAVs) have grown rapidly, most applications only focus on passive visual tasks. Aerial interaction aims to execute tasks involving physic…
General Place Recognition Survey: Towards the Real-world Autonomy Age
Peng Yin, Shiqi Zhao, Ivan Cisneros +6
Place recognition is the fundamental module that can assist Simultaneous Localization and Mapping (SLAM) in loop-closure detection and re-localization for long-term navigation. The…
Lifelong Graph Learning
Chen Wang, Yuheng Qiu, Dasong Gao +1
Graph neural networks (GNN) are powerful models for many graph-structured tasks. Existing models often assume that the complete structure of the graph is available during training.…
UAS Simulator for Modeling, Analysis and Control in Free Flight and Physical Interaction
Azarakhsh Keipour, Mohammadreza Mousaei, Dongwei Bai +2
This paper presents the ARCAD simulator for the rapid development of Unmanned Aerial Systems (UAS), including underactuated and fully-actuated multirotors, fixed-wing aircraft, and…
AirLoop: Lifelong Loop Closure Detection
Dasong Gao, Chen Wang, Sebastian Scherer
Loop closure detection is an important building block that ensures the accuracy and robustness of simultaneous localization and mapping (SLAM) systems. Due to their generalization…
MUI-TARE: Multi-Agent Cooperative Exploration with Unknown Initial Position
Jingtian Yan, Xingqiao Lin, Zhongqiang Ren +6
Multi-agent exploration of a bounded 3D environment with unknown initial positions of agents is a challenging problem. It requires quickly exploring the environments as well as rob…
Image-based Visual Servo Control for Aerial Manipulation Using a Fully-Actuated UAV
Guanqi He, Yash Jangir, Junyi Geng +3
Using Unmanned Aerial Vehicles (UAVs) to perform high-altitude manipulation tasks beyond just passive visual application can reduce the time, cost, and risk of human workers. Prior…
Real-Time Ellipse Detection for Robotics Applications
Azarakhsh Keipour, Guilherme A. S. Pereira, Sebastian Scherer
We propose a new algorithm for real-time detection and tracking of elliptic patterns suitable for real-world robotics applications. The method fits ellipses to each contour in the…
Aug3D: Augmenting large scale outdoor datasets for Generalizable Novel View Synthesis
Aditya Rauniyar, Omar Alama, Silong Yong +2
Recent photorealistic Novel View Synthesis (NVS) advances have increasingly gained attention. However, these approaches remain constrained to small indoor scenes. While optimizatio…
VDB-EDT: An Efficient Euclidean Distance Transform Algorithm Based on VDB Data Structure
Delong Zhu, Chaoqun Wang, Wenshan Wang +3
This paper presents a fundamental algorithm, called VDB-EDT, for Euclidean distance transform (EDT) based on the VDB data structure. The algorithm executes on grid maps and generat…
Neuro-Symbolic Learning for Long-Horizon Task Planning Under Complex Logical Constraints
Qiwei Du, Zitong Zhan, Shaoshu Su +7
Task planning often suffers from severe efficiency bottlenecks when robots must reason over long-horizon action sequences under complex logical constraints, including object afford…
AirLoc: Object-based Indoor Relocalization
Aryan, Bowen Li, Sebastian Scherer +2
Indoor relocalization is vital for both robotic tasks like autonomous exploration and civil applications such as navigation with a cell phone in a shopping mall. Some previous appr…
General Place Recognition Survey: Towards Real-World Autonomy
Peng Yin, Jianhao Jiao, Shiqi Zhao +5
In the realm of robotics, the quest for achieving real-world autonomy, capable of executing large-scale and long-term operations, has positioned place recognition (PR) as a corners…
Super Odometry: IMU-centric LiDAR-Visual-Inertial Estimator for Challenging Environments
Shibo Zhao, Hengrui Zhang, Peng Wang +2
We propose Super Odometry, a high-precision multi-modal sensor fusion framework, providing a simple but effective way to fuse multiple sensors such as LiDAR, camera, and IMU sensor…
MAC-Ego3D: Multi-Agent Gaussian Consensus for Real-Time Collaborative Ego-Motion and Photorealistic 3D Reconstruction
Xiaohao Xu, Feng Xue, Shibo Zhao +3
Real-time multi-agent collaboration for ego-motion estimation and high-fidelity 3D reconstruction is vital for scalable spatial intelligence. However, traditional methods produce s…
Learning Risk-Aware Costmaps via Inverse Reinforcement Learning for Off-Road Navigation
Samuel Triest, Mateo Guaman Castro, Parv Maheshwari +3
The process of designing costmaps for off-road driving tasks is often a challenging and engineering-intensive task. Recent work in costmap design for off-road driving focuses on tr…
RayFronts: Open-Set Semantic Ray Frontiers for Online Scene Understanding and Exploration
Omar Alama, Avigyan Bhattacharya, Haoyang He +6
Open-set semantic mapping is crucial for open-world robots. Current mapping approaches either are limited by the depth range or only map beyond-range entities in constrained settin…
Off-Policy Evaluation with Online Adaptation for Robot Exploration in Challenging Environments
Yafei Hu, Junyi Geng, Chen Wang +2
Autonomous exploration has many important applications. However, classic information gain-based or frontier-based exploration only relies on the robot current state to determine th…
Challenges in Close-Proximity Safe and Seamless Operation of Manned and Unmanned Aircraft in Shared Airspace
Jay Patrikar, Joao P. A. Dantas, Sourish Ghosh +11
We propose developing an integrated system to keep autonomous unmanned aircraft safely separated and behave as expected in conjunction with manned traffic. The main goal is to achi…
TartanCalib: Iterative Wide-Angle Lens Calibration using Adaptive SubPixel Refinement of AprilTags
Bardienus P Duisterhof, Yaoyu Hu, Si Heng Teng +2
Wide-angle cameras are uniquely positioned for mobile robots, by virtue of the rich information they provide in a small, light, and cost-effective form factor. An accurate calibrat…
AirShot: Efficient Few-Shot Detection for Autonomous Exploration
Zihan Wang, Bowen Li, Chen Wang +1
Few-shot object detection has drawn increasing attention in the field of robotic exploration, where robots are required to find unseen objects with a few online provided examples.…
VTOL Failure Detection and Recovery by Utilizing Redundancy
Mohammadreza Mousaei, Azarakhsh Keipour, Junyi Geng +1
Offering vertical take-off and landing (VTOL) capabilities and the ability to travel great distances are crucial for Urban Air Mobility (UAM) vehicles. These capabilities make hybr…
Robust Modeling and Controls for Racing on the Edge
Joshua Spisak, Andrew Saba, Nayana Suvarna +5
Race cars are routinely driven to the edge of their handling limits in dynamic scenarios well above 200mph. Similar challenges are posed in autonomous racing, where a software stac…
Integrating kinematics and environment context into deep inverse reinforcement learning for predicting off-road vehicle trajectories
Yanfu Zhang, Wenshan Wang, Rogerio Bonatti +2
Predicting the motion of a mobile agent from a third-person perspective is an important component for many robotics applications, such as autonomous navigation and tracking. With a…
TartanAviation: Image, Speech, and ADS-B Trajectory Datasets for Terminal Airspace Operations
Jay Patrikar, Joao Dantas, Brady Moon +7
We introduce TartanAviation, an open-source multi-modal dataset focused on terminal-area airspace operations. TartanAviation provides a holistic view of the airport environment by…
Learning Visuomotor Policies for Aerial Navigation Using Cross-Modal Representations
Rogerio Bonatti, Ratnesh Madaan, Vibhav Vineet +2
Machines are a long way from robustly solving open-world perception-control tasks, such as first-person view (FPV) aerial navigation. While recent advances in end-to-end Machine Le…
Demonstrating ViSafe: Vision-enabled Safety for High-speed Detect and Avoid
Parv Kapoor, Ian Higgins, Nikhil Keetha +9
Assured safe-separation is essential for achieving seamless high-density operation of airborne vehicles in a shared airspace. To equip resource-constrained aerial systems with this…
Distilling Global Traversability Priors for Image-based Affordance Prediction in Off-road Environments
Matthew Sivaprakasam, Samuel Triest, Micah Nye +5
Standard methods for autonomous navigation in unstructured terrain are prone to myopic behaviors in long-horizon scenarios. The use of metric maps built from LiDAR or cameras provi…
BioSLAM: A Bio-inspired Lifelong Memory System for General Place Recognition
Peng Yin, Abulikemu Abuduweili, Shiqi Zhao +2
We present BioSLAM, a lifelong SLAM framework for learning various new appearances incrementally and maintaining accurate place recognition for previously visited areas. Unlike hum…
Unified Representation of Geometric Primitives for Graph-SLAM Optimization Using Decomposed Quadrics
Weikun Zhen, Huai Yu, Yaoyu Hu +1
In Simultaneous Localization And Mapping (SLAM) problems, high-level landmarks have the potential to build compact and informative maps compared to traditional point-based landmark…
Present and Future of SLAM in Extreme Underground Environments
Kamak Ebadi, Lukas Bernreiter, Harel Biggie +28
This paper reports on the state of the art in underground SLAM by discussing different SLAM strategies and results across six teams that participated in the three-year-long SubT co…
PRoID: Predicted Rate of Information Delivery in Multi-Robot Exploration and Relaying
Seungchan Kim, Seungjae Baek, Micah Corah +3
We address Multi-Robot Exploration and Relaying (MRER): a team of robots must explore an unknown environment and deliver acquired information to a fixed base station within a missi…
SphereVLAD++: Attention-based and Signal-enhanced Viewpoint Invariant Descriptor
Shiqi Zhao, Peng Yin, Ge Yi +1
LiDAR-based localization approach is a fundamental module for large-scale navigation tasks, such as last-mile delivery and autonomous driving, and localization robustness highly re…
LiDAR Enhanced Structure-from-Motion
Weikun Zhen, Yaoyu Hu, Huai Yu +1
Although Structure-from-Motion (SfM) as a maturing technique has been widely used in many applications, state-of-the-art SfM algorithms are still not robust enough in certain situa…
AnyLoc: Towards Universal Visual Place Recognition
Nikhil Keetha, Avneesh Mishra, Jay Karhade +4
Visual Place Recognition (VPR) is vital for robot localization. To date, the most performant VPR approaches are environment- and task-specific: while they exhibit strong performanc…
Unified Spherical Frontend: Learning Rotation-Equivariant Representations of Spherical Images from Any Camera
Mukai Yu, Mosam Dabhi, Liuyue Xie +2
Modern perception increasingly relies on fisheye, panoramic, and other wide field-of-view (FoV) cameras, yet most pipelines still apply planar CNNs designed for pinhole imagery on…
Near-Optimal Edge Evaluation in Explicit Generalized Binomial Graphs
Sanjiban Choudhury, Shervin Javdani, Siddhartha Srinivasa +1
Robotic motion-planning problems, such as a UAV flying fast in a partially-known environment or a robot arm moving around cluttered objects, require finding collision-free paths qu…
GrndCtrl: Grounding World Models via Self-Supervised Reward Alignment
Haoyang He, Jay Patrikar, Dong-Ki Kim +5
Recent advances in video world modeling have enabled large-scale generative models to simulate embodied environments with high visual fidelity, providing strong priors for predicti…
Time-Optimal Path Planning in a Constant Wind for Uncrewed Aerial Vehicles using Dubins Set Classification
Brady Moon, Sagar Sachdev, Junbin Yuan +1
Time-optimal path planning in high winds for a turning-rate constrained UAV is a challenging problem to solve and is important for deployment and field operations. Previous works h…
AutoMerge: A Framework for Map Assembling and Smoothing in City-scale Environments
Peng Yin, Haowen Lai, Shiqi Zhao +4
We present AutoMerge, a LiDAR data processing framework for assembling a large number of map segments into a complete map. Traditional large-scale map merging methods are fragile t…
SyNeT: Synthetic Negatives for Traversability Learning
Bomena Kim, Hojun Lee, Younsoo Park +3
Reliable traversability estimation is crucial for autonomous robots to navigate complex outdoor environments safely. Existing self-supervised learning frameworks primarily rely on…
AnyThermal: Towards Learning Universal Representations for Thermal Perception
Parv Maheshwari, Jay Karhade, Yogesh Chawla +8
We present AnyThermal, a thermal backbone that captures robust task-agnostic thermal features suitable for a variety of tasks such as cross-modal place recognition, thermal segment…
IA-TIGRIS: An Incremental and Adaptive Sampling-Based Planner for Online Informative Path Planning
Brady Moon, Nayana Suvarna, Andrew Jong +4
Planning paths that maximize information gain for robotic platforms has wide-ranging applications and significant potential impact. To effectively adapt to real-time data collectio…
UNRealNet: Learning Uncertainty-Aware Navigation Features from High-Fidelity Scans of Real Environments
Samuel Triest, David D. Fan, Sebastian Scherer +1
Traversability estimation in rugged, unstructured environments remains a challenging problem in field robotics. Often, the need for precise, accurate traversability estimation is i…
3D Human Reconstruction in the Wild with Collaborative Aerial Cameras
Cherie Ho, Andrew Jong, Harry Freeman +3
Aerial vehicles are revolutionizing applications that require capturing the 3D structure of dynamic targets in the wild, such as sports, medicine, and entertainment. The core chall…
Toward Efficient and Robust Multiple Camera Visual-inertial Odometry
Yao He, Huai Yu, Wen Yang +1
Efficiency and robustness are the essential criteria for the visual-inertial odometry (VIO) system. To process massive visual data, the high cost on CPU resources and computation l…
Amelia: A Large Dataset and Benchmark for Airport Surface Movement Forecasting
Ingrid Navarro, Pablo Ortega-Kral, Jay Patrikar +6
Demand for air travel is rising, straining existing aviation infrastructure. In the US, more than 90% of airport control towers are understaffed, falling short of FAA and union sta…
360FusionNeRF: Panoramic Neural Radiance Fields with Joint Guidance
Shreyas Kulkarni, Peng Yin, Sebastian Scherer
We present a method to synthesize novel views from a single panorama image based on the neural radiance field (NeRF). Prior studies in a similar setting rely on the nei…
Monocular Object and Plane SLAM in Structured Environments
Shichao Yang, Sebastian Scherer
In this paper, we present a monocular Simultaneous Localization and Mapping (SLAM) algorithm using high-level object and plane landmarks. The built map is denser, more compact and…
Design, Modeling and Control for a Tilt-rotor VTOL UAV in the Presence of Actuator Failure
Mohammadreza Mousaei, Junyi Geng, Azarakhsh Keipour +2
Enabling vertical take-off and landing while providing the ability to fly long ranges opens the door to a wide range of new real-world aircraft applications while improving many ex…
Monocular Camera Localization in Prior LiDAR Maps with 2D-3D Line Correspondences
Huai Yu, Weikun Zhen, Wen Yang +2
Light-weight camera localization in existing maps is essential for vision-based navigation. Currently, visual and visual-inertial odometry (VO\&VIO) techniques are well-developed f…
ORStereo: Occlusion-Aware Recurrent Stereo Matching for 4K-Resolution Images
Yaoyu Hu, Wenshan Wang, Huai Yu +2
Stereo reconstruction models trained on small images do not generalize well to high-resolution data. Training a model on high-resolution image size faces difficulties of data avail…
Adaptive Safety Margin Estimation for Safe Real-Time Replanning under Time-Varying Disturbance
Cherie Ho, Jay Patrikar, Rogerio Bonatti +1
Safe navigation in real-time is challenging because engineers need to work with uncertain vehicle dynamics, variable external disturbances, and imperfect controllers. A common safe…
CubeSLAM: Monocular 3D Object SLAM
Shichao Yang, Sebastian Scherer
We present a method for single image 3D cuboid object detection and multi-view object SLAM in both static and dynamic environments, and demonstrate that the two parts can improve e…
AutoODD: Agentic Audits via Bayesian Red Teaming in Black-Box Models
Rebecca Martin, Jay Patrikar, Sebastian Scherer
Specialized machine learning models, regardless of architecture and training, are susceptible to failures in deployment. With their increasing use in high risk situations, the abil…
Don't Fool Me Twice: Adapting to Adversity in the Wild with Experience-Driven Reasoning
Navin Sriram Ravie, Andrew Jong, Krrish Jain +4
In robotics, dangers and adversity modes are often embodiment-specific and relative to each agent. A frontier of autonomous mobile robotics is to enable agents to operate effective…
RGB-D SLAM in Dynamic Environments Using Point Correlations
Weichen Dai, Yu Zhang, Ping Li +2
In this paper, a simultaneous localization and mapping (SLAM) method that eliminates the influence of moving objects in dynamic environments is proposed. This method utilizes the c…
Autonomous drone cinematographer: Using artistic principles to create smooth, safe, occlusion-free trajectories for aerial filming
Rogerio Bonatti, Yanfu Zhang, Sanjiban Choudhury +2
Autonomous aerial cinematography has the potential to enable automatic capture of aesthetically pleasing videos without requiring human intervention, empowering individuals with th…
Learning Energy-Efficient Air--Ground Actuation for Hybrid Robots on Stair-Like Terrain
Jiaxing Li, Wen Tian, Xinhang Xu +3
Hybrid aerial--ground robots offer both traversability and endurance, but stair-like discontinuities create a trade-off: wheels alone often stall at edges, while flight is energy-h…
3D-SiamRPN: An End-to-End Learning Method for Real-Time 3D Single Object Tracking Using Raw Point Cloud
Zheng Fang, Sifan Zhou, Yubo Cui +1
3D single object tracking is a key issue for autonomous following robot, where the robot should robustly track and accurately localize the target for efficient following. In this p…
AirIO: Learning Inertial Odometry with Enhanced IMU Feature Observability
Yuheng Qiu, Can Xu, Yutian Chen +3
Inertial odometry (IO) using only Inertial Measurement Units (IMUs) offers a lightweight and cost-effective solution for Unmanned Aerial Vehicle (UAV) applications, yet existing le…
Improved Generalization of Heading Direction Estimation for Aerial Filming Using Semi-supervised Regression
Wenshan Wang, Aayush Ahuja, Yanfu Zhang +2
In the task of Autonomous aerial filming of a moving actor (e.g. a person or a vehicle), it is crucial to have a good heading direction estimation for the actor from the visual inp…
Attention-Enhanced Cross-modal Localization Between 360 Images and Point Clouds
Zhipeng Zhao, Huai Yu, Chenwei Lyv +2
Visual localization plays an important role for intelligent robots and autonomous driving, especially when the accuracy of GNSS is unreliable. Recently, camera localization in LiDA…
In-flight positional and energy use data set of a DJI Matrice 100 quadcopter for small package delivery
Thiago A. Rodrigues, Jay Patrikar, Arnav Choudhry +9
We autonomously direct a small quadcopter package delivery Uncrewed Aerial Vehicle (UAV) or "drone" to take off, fly a specified route, and land for a total of 209 flights while va…
Follow The Rules: Online Signal Temporal Logic Tree Search for Guided Imitation Learning in Stochastic Domains
Jasmine Jerry Aloor, Jay Patrikar, Parv Kapoor +2
Seamlessly integrating rules in Learning-from-Demonstrations (LfD) policies is a critical requirement to enable the real-world deployment of AI agents. Recently, Signal Temporal Lo…
MAC-VO: Metrics-aware Covariance for Learning-based Stereo Visual Odometry
Yuheng Qiu, Yutian Chen, Zihao Zhang +2
We propose the MAC-VO, a novel learning-based stereo VO that leverages the learned metrics-aware matching uncertainty for dual purposes: selecting keypoint and weighing the residua…
AirDOS: Dynamic SLAM benefits from Articulated Objects
Yuheng Qiu, Chen Wang, Wenshan Wang +2
Dynamic Object-aware SLAM (DOS) exploits object-level information to enable robust motion estimation in dynamic environments. Existing methods mainly focus on identifying and exclu…
Monocular and Stereo Cues for Landing Zone Evaluation for Micro UAVs
Rohit Garg, Shichao Yang, Sebastian Scherer
Autonomous and safe landing is important for unmanned aerial vehicles. We present a monocular and stereo image based method for fast and accurate landing zone evaluation for UAVs i…
UMI-on-Air: Embodiment-Aware Guidance for Embodiment-Agnostic Visuomotor Policies
Harsh Gupta, Xiaofeng Guo, Huy Ha +6
We introduce UMI-on-Air, a framework for embodiment-aware deployment of embodiment-agnostic manipulation policies. Our approach leverages diverse, unconstrained human demonstration…
ALFA: A Dataset for UAV Fault and Anomaly Detection
Azarakhsh Keipour, Mohammadreza Mousaei, Sebastian Scherer
We present a dataset of several fault types in control surfaces of a fixed-wing Unmanned Aerial Vehicle (UAV) for use in Fault Detection and Isolation (FDI) and Anomaly Detection (…
Informative Sensor Planning for a Single-Axis Gimbaled Camera on a Fixed-Wing UAV
Aditya Parandekar, Brady Moon, Nayana Suvarna +1
Uncrewed Aerial Vehicles (UAVs) are a leading choice of platforms for a variety of information-gathering applications. Sensor planning can enhance the efficiency and success of the…
PIAug -- Physics Informed Augmentation for Learning Vehicle Dynamics for Off-Road Navigation
Parv Maheshwari, Wenshan Wang, Samuel Triest +5
Modeling the precise dynamics of off-road vehicles is a complex yet essential task due to the challenging terrain they encounter and the need for optimal performance and safety. Re…
A Unified 3D Mapping Framework using a 3D or 2D LiDAR
Weikun Zhen, Sebastian Scherer
Simultaneous Localization and Mapping (SLAM) has been considered as a solved problem thanks to the progress made in the past few years. However, the great majority of LiDAR-based S…
AdaFusion: Visual-LiDAR Fusion with Adaptive Weights for Place Recognition
Haowen Lai, Peng Yin, Sebastian Scherer
Recent years have witnessed the increasing application of place recognition in various environments, such as city roads, large buildings, and a mix of indoor and outdoor places. Th…
TartanGround: A Large-Scale Dataset for Ground Robot Perception and Navigation
Manthan Patel, Fan Yang, Yuheng Qiu +4
We present TartanGround, a large-scale, multi-modal dataset to advance the perception and autonomy of ground robots operating in diverse environments. This dataset, collected in va…
Visual Servoing Approach for Autonomous UAV Landing on a Moving Vehicle
Azarakhsh Keipour, Guilherme A. S. Pereira, Rogerio Bonatti +4
Many aerial robotic applications require the ability to land on moving platforms, such as delivery trucks and marine research boats. We present a method to autonomously land an Unm…
MapEx: Indoor Structure Exploration with Probabilistic Information Gain from Global Map Predictions
Cherie Ho, Seungchan Kim, Brady Moon +6
Exploration is a critical challenge in robotics, centered on understanding unknown environments. In this work, we focus on robots exploring structured indoor environments which are…
Any4D: Unified Feed-Forward Metric 4D Reconstruction
Jay Karhade, Nikhil Keetha, Yuchen Zhang +4
We present Any4D, a scalable multi-view transformer for metric-scale, dense feed-forward 4D reconstruction. Any4D directly generates per-pixel motion and geometry predictions for N…
Learning-on-the-Drive: Self-supervised Adaptation of Visual Offroad Traversability Models
Eric Chen, Cherie Ho, Mukhtar Maulimov +2
Autonomous offroad driving is essential for applications like emergency rescue, military operations, and agriculture. Despite progress, systems struggle with high-speed vehicles ex…
SoRTS: Learned Tree Search for Long Horizon Social Robot Navigation
Ingrid Navarro, Jay Patrikar, Joao P. A. Dantas +4
The fast-growing demand for fully autonomous robots in shared spaces calls for the development of trustworthy agents that can safely and seamlessly navigate in crowded environments…
Unifying Deep Predicate Invention with Pre-trained Foundation Models
Qianwei Wang, Bowen Li, Zhanpeng Luo +6
Long-horizon robotic tasks are hard due to continuous state-action spaces and sparse feedback. Symbolic world models help by decomposing tasks into discrete predicates that capture…
LogiCity: Advancing Neuro-Symbolic AI with Abstract Urban Simulation
Bowen Li, Zhaoyu Li, Qiwei Du +10
Recent years have witnessed the rapid development of Neuro-Symbolic (NeSy) AI systems, which integrate symbolic reasoning into deep neural networks. However, most of the existing b…
SplaTAM: Splat, Track & Map 3D Gaussians for Dense RGB-D SLAM
Nikhil Keetha, Jay Karhade, Krishna Murthy Jatavallabhula +4
Dense simultaneous localization and mapping (SLAM) is crucial for robotics and augmented reality applications. However, current methods are often hampered by the non-volumetric or…
TartanVO: A Generalizable Learning-based VO
Wenshan Wang, Yaoyu Hu, Sebastian Scherer
We present the first learning-based visual odometry (VO) model, which generalizes to multiple datasets and real-world scenarios and outperforms geometry-based methods in challengin…
AirCode: A Robust Object Encoding Method
Kuan Xu, Chen Wang, Chao Chen +2
Object encoding and identification are crucial for many robotic tasks such as autonomous exploration and semantic relocalization. Existing works heavily rely on the tracking of det…
3D Segmentation Learning from Sparse Annotations and Hierarchical Descriptors
Peng Yin, Lingyun Xu, Jianmin Ji +2
One of the main obstacles to 3D semantic segmentation is the significant amount of endeavor required to generate expensive point-wise annotations for fully supervised training. To…
2D-3D Pose Tracking with Multi-View Constraints
Huai Yu, Kuangyi Chen, Wen Yang +2
Camera localization in 3D LiDAR maps has gained increasing attention due to its promising ability to handle complex scenarios, surpassing the limitations of visual-only localizatio…
A Robust Laser-Inertial Odometry and Mapping Method for Large-Scale Highway Environments
Shibo Zhao, Zheng Fang, HaoLai Li +1
In this paper, we propose a novel laser-inertial odometry and mapping method to achieve real-time, low-drift and robust pose estimation in large-scale highway environments. The pro…