papers

Publications (98)

cs.CL2016

Learning Lexical Entries for Robotic Commands using Crowdsourcing

Junjie Hu, Jean Oh, Anatole Gershman

Robotic commands in natural language usually contain various spatial descriptions that are semantically similar but syntactically different. Mapping such syntactic variants into se…

cs.CV2025

VPOcc: Exploiting Vanishing Point for 3D Semantic Occupancy Prediction

Junsu Kim, Junhee Lee, Ukcheol Shin +2

Understanding 3D scenes semantically and spatially is crucial for the safe navigation of robots and autonomous vehicles, aiding obstacle avoidance and accurate trajectory planning.…

cs.RO2025

GRAPPA: Generalizing and Adapting Robot Policies via Online Agentic Guidance

Arthur Bucker, Pablo Ortega-Kral, Jonathan Francis +1

Robot learning approaches such as behavior cloning and reinforcement learning have shown great promise in synthesizing robot skills from human demonstrations in specific environmen…

cs.CV2020

Image Captioning with Compositional Neural Module Networks

Junjiao Tian, Jean Oh

In image captioning where fluency is an important factor in evaluation, e.g., -gram metrics, sequential models are commonly used; however, sequential models generally result in…

cs.CV2021

Anytime 3D Object Reconstruction using Multi-modal Variational Autoencoder

Hyeonwoo Yu, Jean Oh

For effective human-robot teaming, it is important for the robots to be able to share their visual perception with the human operators. In a harsh remote collaboration setting, dat…

cs.RO2024

SoRTS: Learned Tree Search for Long Horizon Social Robot Navigation

Ingrid Navarro, Jay Patrikar, Joao P. A. Dantas +4

The fast-growing demand for fully autonomous robots in shared spaces calls for the development of trustworthy agents that can safely and seamlessly navigate in crowded environments…

cs.CV2020

Trajformer: Trajectory Prediction with Local Self-Attentive Contexts for Autonomous Driving

Manoj Bhat, Jonathan Francis, Jean Oh

Effective feature-extraction is critical to models' contextual understanding, particularly for applications to robotics and autonomous driving, such as multimodal trajectory predic…

cs.RO2025

STRIVE: Structured Representation Integrating VLM Reasoning for Efficient Object Navigation

Haokun Zhu, Zongtai Li, Zhixuan Liu +4

Vision-Language Models (VLMs) have been increasingly integrated into object navigation tasks for their rich prior knowledge and strong reasoning abilities. However, applying VLMs t…

cs.CV2022

StyleCLIPDraw: Coupling Content and Style in Text-to-Drawing Synthesis

Peter Schaldenbrand, Zhixuan Liu, Jean Oh

Generating images that fit a given text description using machine learning has improved greatly with the release of technologies such as the CLIP image-text encoder model; however,…

cs.CV2026

Exposing Blindspots: Cultural Bias Evaluation in Generative Image Models

Huichan Seo, Sieun Choi, Minki Hong +8

Generative image models produce striking visuals yet often misrepresent culture. Prior work has examined cultural bias mainly in text-to-image (T2I) systems, leaving image-to-image…

cs.CV2025

Robot Synesthesia: A Sound and Emotion Guided AI Painter

Vihaan Misra, Peter Schaldenbrand, Jean Oh

If a picture paints a thousand words, sound may voice a million. While recent robotic painting and image synthesis methods have achieved progress in generating visuals from text in…

cs.RO2021

Core Challenges of Social Robot Navigation: A Survey

Christoforos Mavrogiannis, Francesca Baldini, Allan Wang +4

Robot navigation in crowded public spaces is a complex task that requires addressing a variety of engineering and human factors challenges. These challenges have motivated a great…

cs.RO2018

Social Attention: Modeling Attention in Human Crowds

Anirudh Vemula, Katharina Muelling, Jean Oh

Robots that navigate through human crowds need to be able to plan safe, efficient, and human predictable trajectories. This is a particularly challenging problem as it requires the…

cs.RO2024

Designing Anthropomorphic Soft Hands through Interaction

Pragna Mannam, Kenneth Shaw, Dominik Bauer +3

Modeling and simulating soft robot hands can aid in design iteration for complex and high degree-of-freedom (DoF) morphologies. This can be further supplemented by iterating on the…

cs.CV2023

FishRecGAN: An End to End GAN Based Network for Fisheye Rectification and Calibration

Xin Shen, Kyungdon Joo, Jean Oh

We propose an end-to-end deep learning approach to rectify fisheye images and simultaneously calibrate camera intrinsic and distortion parameters. Our method consists of two parts:…

cs.RO2026

Functional Force-Aware Retargeting from Virtual Human Demos to Soft Robot Policies

Uksang Yoo, Mengjia Zhu, Evan Pezent +8

We introduce SoftAct, a framework for teaching soft robot hands to perform human-like manipulation skills by explicitly reasoning about contact forces. Leveraging immersive virtual…

cs.CV2023

EigenTrajectory: Low-Rank Descriptors for Multi-Modal Trajectory Forecasting

Inhwan Bae, Jean Oh, Hae-Gon Jeon

Capturing high-dimensional social interactions and feasible futures is essential for predicting trajectories. To address this complex nature, several attempts have been devoted to…

cs.RO2024

SonicBoom: Contact Localization Using Array of Microphones

Moonyoung Lee, Uksang Yoo, Jean Oh +3

In cluttered environments where visual sensors encounter heavy occlusion, such as in agricultural settings, tactile signals can provide crucial spatial information for the robot to…

cs.RO2026

SysNav: Multi-Level Systematic Cooperation Enables Real-World, Cross-Embodiment Object Navigation

Haokun Zhu, Zongtai Li, Zihan Liu +8

Object navigation (ObjectNav) in real-world environments is a complex problem that requires simultaneously addressing multiple challenges, including complex spatial structure, long…

cs.RO2025

Spline-FRIDA: Towards Diverse, Humanlike Robot Painting Styles with a Sample-Efficient, Differentiable Brush Stroke Model

Lawrence Chen, Peter Schaldenbrand, Tanmay Shankar +2

A painting is more than just a picture on a wall; a painting is a process comprised of many intentional brush strokes, the shapes of which are an important component of a painting'…

cs.CV2022

Towards Real-Time Text2Video via CLIP-Guided, Pixel-Level Optimization

Peter Schaldenbrand, Zhixuan Liu, Jean Oh

We introduce an approach to generating videos based on a series of given language descriptions. Frames of the video are generated sequentially and optimized by guidance from the CL…

cs.RO2016

Path Planning in Dynamic Environments with Adaptive Dimensionality

Anirudh Vemula, Katharina Muelling, Jean Oh

Path planning in the presence of dynamic obstacles is a challenging problem due to the added time dimension in search space. In approaches that ignore the time dimension and treat…

cs.CV2026

Toward Trustworthy Portrait Editing: Evaluation of Demographic Misrepresentation in I2I Models

Huichan Seo, Minki Hong, Sieun Choi +2

Instruction-guided image-to-image (I2I) editors are increasingly used in consumer and professional visual workflows, where trustworthiness depends not only on prompt compliance but…

cs.RO2023

Core Challenges in Embodied Vision-Language Planning

Jonathan Francis, Nariaki Kitamura, Felix Labelle +3

Recent advances in the areas of Multimodal Machine Learning and Artificial Intelligence (AI) have led to the development of challenging tasks at the intersection of Computer Vision…

cs.RO2019

Following Social Groups: Socially Compliant Autonomous Navigation in Dense Crowds

Xinjie Yao, Ji Zhang, Jean Oh

In densely populated environments, socially compliant navigation is critical for autonomous robots as driving close to people is unavoidable. This manner of social navigation is ch…

cs.RO2022

RCA: Ride Comfort-Aware Visual Navigation via Self-Supervised Learning

Xinjie Yao, Ji Zhang, Jean Oh

Under shared autonomy, wheelchair users expect vehicles to provide safe and comfortable rides while following users high-level navigation plans. To find such a path, vehicles negot…

cs.CV2024

Complementary Random Masking for RGB-Thermal Semantic Segmentation

Ukcheol Shin, Kyunghyun Lee, In So Kweon +1

RGB-thermal semantic segmentation is one potential solution to achieve reliable semantic scene understanding in adverse weather and lighting conditions. However, the previous studi…

cs.RO2026

Visual Sculpting: Visually-Aligned Planning Representations for Long-Horizon Robot Clay Sculpting

Peter Schaldenbrand, Jean Oh

Clay sculpting is a nuanced, artistic task involving dexterous manipulation with long-horizon planning to achieve high-level goals. As a robotics problem, we formulate clay sculpti…

cs.CV2023

Regularizing Self-training for Unsupervised Domain Adaptation via Structural Constraints

Rajshekhar Das, Jonathan Francis, Sanket Vaibhav Mehta +3

Self-training based on pseudo-labels has emerged as a dominant approach for addressing conditional distribution shifts in unsupervised domain adaptation (UDA) for semantic segmenta…

cs.RO2020

NaviGAN: A Generative Approach for Socially Compliant Navigation

Chieh-En Tsai, Jean Oh

Robots navigating in human crowds need to optimize their paths not only for their task performance but also for their compliance to social norms. One of the key challenges in this…

cs.CV2026

ShapeShift: Text-to-Mosaic Synthesis via Semantic Phase-Field Guidance

Vihaan Misra, Peter Schaldenbrand, Jean Oh

We present ShapeShift, a method for arranging rigid objects into configurations that visually convey semantic concepts specified by natural language. While pretrained diffusion mod…

cs.RO2024

Soft Robotic Dynamic In-Hand Pen Spinning

Yunchao Yao, Uksang Yoo, Jean Oh +2

Dynamic in-hand manipulation remains a challenging task for soft robotic systems that have demonstrated advantages in safe compliant interactions but struggle with high-speed dynam…

cs.RO2024

SafeShift: Safety-Informed Distribution Shifts for Robust Trajectory Prediction in Autonomous Driving

Benjamin Stoler, Ingrid Navarro, Meghdeep Jana +3

As autonomous driving technology matures, safety and robustness of its key components, including trajectory prediction, is vital. Though real-world datasets, such as Waymo Open Mot…

cs.RO2024

RoPotter: Toward Robotic Pottery and Deformable Object Manipulation with Structural Priors

Uksang Yoo, Adam Hung, Jonathan Francis +2

Humans are capable of continuously manipulating a wide variety of deformable objects into complex shapes. This is made possible by our intuitive understanding of material propertie…

cs.LG2022

Core Challenges in Embodied Vision-Language Planning

Jonathan Francis, Nariaki Kitamura, Felix Labelle +3

Recent advances in the areas of multimodal machine learning and artificial intelligence (AI) have led to the development of challenging tasks at the intersection of Computer Vision…

cs.RO2021

Language Understanding for Field and Service Robots in a Priori Unknown Environments

Matthew R. Walter, Siddharth Patki, Andrea F. Daniele +7

Contemporary approaches to perception, planning, estimation, and control have allowed robots to operate robustly as our remote surrogates in uncertain, unstructured environments. T…

cs.RO2025

Learning Social Navigation from Positive and Negative Demonstrations and Rule-Based Specifications

Chanwoo Kim, Jihwan Yoon, Hyeonseong Kim +9

Mobile robot navigation in dynamic human environments requires policies that balance adaptability to diverse behaviors with compliance to safety constraints. We hypothesize that in…

cs.RO2025

KineSoft: Learning Proprioceptive Manipulation Policies with Soft Robot Hands

Uksang Yoo, Jonathan Francis, Jean Oh +1

Underactuated soft robot hands offer inherent safety and adaptability advantages over rigid systems, but developing dexterous manipulation skills remains challenging. While imitati…

cs.AI2025

Secure & Personalized Music-to-Video Generation via CHARCHA

Mehul Agarwal, Gauri Agarwal, Santiago Benoit +2

Music is a deeply personal experience and our aim is to enhance this with a fully-automated pipeline for personalized music video generation. Our work allows listeners to not just…

cs.CV2026

Bridging Online and Offline Handwriting via Differentiable Physical Rendering

Seonmi Park, Seunghyun Shin, Vihaan Misra +4

Realistic handwritten text generation plays an important role in numerous applications, such as font design, biometric authentication, and robotic calligraphy. Existing methods are…

cs.RO2021

Safe Autonomous Racing via Approximate Reachability on Ego-vision

Bingqing Chen, Jonathan Francis, Jean Oh +2

Racing demands each vehicle to drive at its physical limits, when any safety infraction could lead to catastrophic failure. In this work, we study the problem of safe reinforcement…

cs.RO2026

A-SLIP: Acoustic Sensing for Continuous In-hand Slip Estimation

Uksang Yoo, Yuemin Mao, Jean Oh +1

Reliable in-hand manipulation requires accurate real-time estimation of slip between a gripper and a grasped object. Existing tactile sensing approaches based on vision, capacitanc…

cs.RO2024

Design and Control Co-Optimization for Automated Design Iteration of Dexterous Anthropomorphic Soft Robotic Hands

Pragna Mannam, Xingyu Liu, Ding Zhao +2

We automate soft robotic hand design iteration by co-optimizing design and control policy for dexterous manipulation skills in simulation. Our design iteration pipeline combines ge…

cs.RO2026

VibeAct: Vibration to Actions for Contact-Rich Reactive Robot Dexterity

Yuemin Mao, Uksang Yoo, Jean Oh +2

Dexterous manipulation depends on contact events that are fast, local, and often visually occluded. Piezoelectric microphones offer a compact and high-bandwidth way to sense these…

cs.RO2026

RIO: Flexible Real-Time Robot I/O for Cross-Embodiment Robot Learning

Pablo Ortega-Kral, Eliot Xing, Arthur Bucker +13

Despite recent efforts to collect multi-task, multi-embodiment datasets, to design recipes for training Vision-Language-Action models (VLAs), and to showcase these models on differ…

cs.RO2022

FAR Planner: Fast, Attemptable Route Planner using Dynamic Visibility Update

Fan Yang, Chao Cao, Hongbiao Zhu +2

The problem of path planning in unknown environments remains a challenging problem - as the environment is gradually observed during the navigation, the underlying planner has to u…

cs.AI2025

Sample-Efficient Behavior Cloning Using General Domain Knowledge

Feiyu Zhu, Jean Oh, Reid Simmons

Behavior cloning has shown success in many sequential decision-making tasks by learning from expert demonstrations, yet they can be very sample inefficient and fail to generalize t…

cs.RO2026

SOFTMAP: Sim2Real Soft Robot Forward Modeling via Topological Mesh Alignment and Physics Prior

Ziyong Ma, Uksang Yoo, Jonathan Francis +3

While soft robot manipulators offer compelling advantages over rigid counterparts, including inherent compliance, safe human-robot interaction, and the ability to conform to comple…

cs.HC2025

Visuo-Acoustic Hand Pose and Contact Estimation

Yuemin Mao, Uksang Yoo, Yunchao Yao +5

Accurately estimating hand pose and hand-object contact events is essential for robot data-collection, immersive virtual environments, and biomechanical analysis, yet remains chall…

cs.CV2026

TokenDial: Continuous Attribute Control in Text-to-Video via Spatiotemporal Token Offsets

Zhixuan Liu, Peter Schaldenbrand, Yijun Li +5

We present TokenDial, a framework for continuous, slider-style attribute control in pretrained text-to-video generation models. While modern generators produce strong holistic vide…

cs.AI2026

InterPReT: Interactive Policy Restructuring and Training Enable Effective Imitation Learning from Laypersons

Feiyu Gavin Zhu, Jean Oh, Reid Simmons

Imitation learning has shown success in many tasks by learning from expert demonstrations. However, most existing work relies on large-scale demonstrations from technical professio…

cs.CV2023

T2FPV: Dataset and Method for Correcting First-Person View Errors in Pedestrian Trajectory Prediction

Benjamin Stoler, Meghdeep Jana, Soonmin Hwang +1

Predicting pedestrian motion is essential for developing socially-aware robots that interact in a crowded environment. While the natural visual perspective for a social interaction…

cs.RO2020

Artistic Style in Robotic Painting; a Machine Learning Approach to Learning Brushstroke from Human Artists

Ardavan Bidgoli, Manuel Ladron De Guevara, Cinnie Hsiung +2

Robotic painting has been a subject of interest among both artists and roboticists since the 1970s. Researchers and interdisciplinary artists have employed various painting techniq…

cs.RO2024

Inclusion in Assistive Haircare Robotics: Practical and Ethical Considerations in Hair Manipulation

Uksang Yoo, Nathaniel Dennler, Sarvesh Patil +2

Robot haircare systems could provide a controlled and personalized environment that is respectful of an individual's sensitivities and may offer a comfortable experience. We argue…

cs.RO2019

Explainable Semantic Mapping for First Responders

Jean Oh, Martial Hebert, Hae-Gon Jeon +3

One of the key challenges in the semantic mapping problem in postdisaster environments is how to analyze a large amount of data efficiently with minimal supervision. To address thi…

cs.RO2022

Challenges in Close-Proximity Safe and Seamless Operation of Manned and Unmanned Aircraft in Shared Airspace

Jay Patrikar, Joao P. A. Dantas, Sourish Ghosh +11

We propose developing an integrated system to keep autonomous unmanned aircraft safely separated and behave as expected in conjunction with manned traffic. The main goal is to achi…

cs.RO2022

Learn-to-Race Challenge 2022: Benchmarking Safe Learning and Cross-domain Generalisation in Autonomous Racing

Jonathan Francis, Bingqing Chen, Siddha Ganju +10

We present the results of our autonomous racing virtual challenge, based on the newly-released Learn-to-Race (L2R) simulation framework, which seeks to encourage interdisciplinary…

cs.RO2024

CoFRIDA: Self-Supervised Fine-Tuning for Human-Robot Co-Painting

Peter Schaldenbrand, Gaurav Parmar, Jun-Yan Zhu +2

Prior robot painting and drawing work, such as FRIDA, has focused on decreasing the sim-to-real gap and expanding input modalities for users, but the interaction with these systems…

cs.CV2024

SCoFT: Self-Contrastive Fine-Tuning for Equitable Image Generation

Zhixuan Liu, Peter Schaldenbrand, Beverley-Claire Okogwu +5

Accurate representation in media is known to improve the well-being of the people who consume it. Generative image models trained on large web-crawled datasets such as LAION are kn…

cs.CV2022

StyleCLIPDraw: Coupling Content and Style in Text-to-Drawing Translation

Peter Schaldenbrand, Zhixuan Liu, Jean Oh

Generating images that fit a given text description using machine learning has improved greatly with the release of technologies such as the CLIP image-text encoder model; however,…

cs.RO2023

Learned Tree Search for Long-Horizon Social Robot Navigation in Shared Airspace

Ingrid Navarro, Jay Patrikar, Joao P. A. Dantas +4

The fast-growing demand for fully autonomous aerial operations in shared spaces necessitates developing trustworthy agents that can safely and seamlessly navigate in crowded, dynam…

cs.RO2022

Predicting Like A Pilot: Dataset and Method to Predict Socially-Aware Aircraft Trajectories in Non-Towered Terminal Airspace

Jay Patrikar, Brady Moon, Jean Oh +1

Pilots operating aircraft in un-towered airspace rely on their situational awareness and prior knowledge to predict the future trajectories of other agents. These predictions are c…

cs.CL2020

Adjusting Image Attributes of Localized Regions with Low-level Dialogue

Tzu-Hsiang Lin, Alexander Rudnicky, Trung Bui +2

Natural Language Image Editing (NLIE) aims to use natural language instructions to edit images. Since novices are inexperienced with image editing techniques, their instructions ar…

cs.AI2026

Before We Trust Them: Decision-Making Failures in Navigation of Foundation Models

Jua Han, Jaeyoon Seo, Jungbin Min +4

High success rates on navigation-related tasks do not necessarily translate into reliable decision making by foundation models. To examine this gap, we evaluate current models on s…

cs.CV2021

Noticing Motion Patterns: Temporal CNN with a Novel Convolution Operator for Human Trajectory Prediction

Dapeng Zhao, Jean Oh

We propose a Convolutional Neural Network-based approach to learn, detect,and extract patterns in sequential trajectory data, known here as Social Pattern Extraction Convolution (S…

cs.CV2025

MOSAIC: Generating Consistent, Privacy-Preserving Scenes from Multiple Depth Views in Multi-Room Environments

Zhixuan Liu, Haokun Zhu, Rui Chen +4

We introduce a diffusion-based approach for generating privacy-preserving digital twins of multi-room indoor environments from depth images only. Central to our approach is a novel…

cs.RO2025

SEAL: Towards Safe Autonomous Driving via Skill-Enabled Adversary Learning for Closed-Loop Scenario Generation

Benjamin Stoler, Ingrid Navarro, Jonathan Francis +1

Verification and validation of autonomous driving (AD) systems and components is of increasing importance, as such technology increases in real-world prevalence. Safety-critical sc…

cs.RO2024

POE: Acoustic Soft Robotic Proprioception for Omnidirectional End-effectors

Uksang Yoo, Ziven Lopez, Jeffrey Ichnowski +1

Soft robotic shape estimation and proprioception are challenging because of soft robot's complex deformation behaviors and infinite degrees of freedom. A soft robot's continuously…

cs.RO2021

Autonomous Exploration Development Environment and the Planning Algorithms

Chao Cao, Hongbiao Zhu, Fan Yang +4

Autonomous Exploration Development Environment is an open-source repository released to facilitate the development of high-level planning algorithms and integration of complete aut…

cs.CV2021

Domain Adaptive Monocular Depth Estimation With Semantic Information

Fei Lu, Hyeonwoo Yu, Jean Oh

The advent of deep learning has brought an impressive advance to monocular depth estimation, e.g., supervised monocular depth estimation has been thoroughly investigated. However,…

cs.RO2017

Modeling Cooperative Navigation in Dense Human Crowds

Anirudh Vemula, Katharina Muelling, Jean Oh

For robots to be a part of our daily life, they need to be able to navigate among crowds not only safely but also in a socially compliant fashion. This is a challenging problem bec…

cs.RO2026

Interacting safely with cyclists using Hamilton-Jacobi reachability and reinforcement learning

Aarati Andrea Noronha, Jean Oh

In this paper, we present a framework for enabling autonomous vehicles to interact with cyclists in a manner that balances safety and optimality. The approach integrates Hamilton-J…

cs.CV2021

Localize, Group, and Select: Boosting Text-VQA by Scene Text Modeling

Xiaopeng Lu, Zhen Fan, Yansen Wang +2

As an important task in multimodal context understanding, Text-VQA (Visual Question Answering) aims at question answering through reading text information in images. It differentia…

cs.LG2025

Stabilizing Reinforcement Learning in Differentiable Multiphysics Simulation

Eliot Xing, Vernon Luk, Jean Oh

Recent advances in GPU-based parallel simulation have enabled practitioners to collect large amounts of data and train complex control policies using deep reinforcement learning (R…

cs.RO2022

FRIDA: A Collaborative Robot Painter with a Differentiable, Real2Sim2Real Planning Environment

Peter Schaldenbrand, James McCann, Jean Oh

Painting is an artistic process of rendering visual content that achieves the high-level communication goals of an artist that may change dynamically throughout the creative proces…

cs.RO2025

LongComp: Long-Tail Compositional Zero-Shot Generalization for Robust Trajectory Prediction

Benjamin Stoler, Jonathan Francis, Jean Oh

Methods for trajectory prediction in Autonomous Driving must contend with rare, safety-critical scenarios that make reliance on real-world data collection alone infeasible. To asse…

cs.CV2021

Anchor Distance for 3D Multi-Object Distance Estimation from 2D Single Shot

Hyeonwoo Yu, Jean Oh

Visual perception of the objects in a 3D environment is a key to successful performance in autonomous driving and simultaneous localization and mapping (SLAM). In this paper, we pr…

cs.RO2022

Knowledge-driven Scene Priors for Semantic Audio-Visual Embodied Navigation

Gyan Tatiya, Jonathan Francis, Luca Bondi +4

Generalisation to unseen contexts remains a challenge for embodied navigation agents. In the context of semantic audio-visual navigation (SAVi) tasks, the notion of generalisation…

cs.LG2025

Amelia: A Large Dataset and Benchmark for Airport Surface Movement Forecasting

Ingrid Navarro, Pablo Ortega-Kral, Jay Patrikar +6

Demand for air travel is rising, straining existing aviation infrastructure. In the US, more than 90% of airport control towers are understaffed, falling short of FAA and union sta…

cs.CL2020

A Multimodal Dialogue System for Conversational Image Editing

Tzu-Hsiang Lin, Trung Bui, Doo Soon Kim +1

In this paper, we present a multimodal dialogue system for Conversational Image Editing. We formulate our multimodal dialogue system as a Partially Observed Markov Decision Process…

cs.RO2026

Bridging Language and Action: A Survey of Language-Conditioned Robot Manipulation

Xiangtong Yao, Hongkuan Zhou, Oier Mees +12

Language-conditioned robot manipulation is an emerging field aimed at enabling seamless communication and cooperation between humans and robotic agents by teaching robots to compre…

cs.RO2025

The Foundational Pose as a Selection Mechanism for the Design of Tool-Wielding Multi-Finger Robotic Hands

Sunyu Wang, Jean Oh, Nancy S. Pollard

To wield an object means to hold and move it in a way that exploits its functions. When humans wield tools -- such as writing with a pen or cutting with scissors -- our hands would…

cs.RO2024

Towards Human-Centered Construction Robotics: A Reinforcement Learning-Driven Companion Robot for Contextually Assisting Carpentry Workers

Yuning Wu, Jiaying Wei, Jean Oh +1

In the dynamic construction industry, traditional robotic integration has primarily focused on automating specific tasks, often overlooking the complexity and variability of human…

cs.RO2022

UGV-UAV Object Geolocation in Unstructured Environments

David Guttendorf, D. W. Wilson Hamilton, Anne Harris Heckman +12

A robotic system of multiple unmanned ground vehicles (UGVs) and unmanned aerial vehicles (UAVs) has the potential for advancing autonomous object geolocation performance. Much res…

cs.CV2021

Self-supervised Learning of 3D Object Understanding by Data Association and Landmark Estimation for Image Sequence

Hyeonwoo Yu, Jean Oh

In this paper, we propose a self-supervised learningmethod for multi-object pose estimation. 3D object under-standing from 2D image is a challenging task that infers ad-ditional di…

cs.CV2024

Flow4D: Leveraging 4D Voxel Network for LiDAR Scene Flow Estimation

Jaeyeul Kim, Jungwan Woo, Ukcheol Shin +2

Understanding the motion states of the surrounding environment is critical for safe autonomous driving. These motion states can be accurately derived from scene flow, which capture…

cs.CV2024

FIReStereo: Forest InfraRed Stereo Dataset for UAS Depth Perception in Visually Degraded Environments

Devansh Dhrafani, Yifei Liu, Andrew Jong +6

Robust depth perception in visually-degraded environments is crucial for autonomous aerial systems. Thermal imaging cameras, which capture infrared radiation, are robust to visual…

cs.RO2025

Soft and Compliant Contact-Rich Hair Manipulation and Care

Uksang Yoo, Nathaniel Dennler, Eliot Xing +4

Hair care robots can help address labor shortages in elderly care while enabling those with limited mobility to maintain their hair-related identity. We present MOE-Hair, a soft ro…

cs.CV2025

Bridging Spectral-wise and Multi-spectral Depth Estimation via Geometry-guided Contrastive Learning

Ukcheol Shin, Kyunghyun Lee, Jean Oh

Deploying depth estimation networks in the real world requires high-level robustness against various adverse conditions to ensure safe and reliable autonomy. For this purpose, many…

cs.RO2022

Social-PatteRNN: Socially-Aware Trajectory Prediction Guided by Motion Patterns

Ingrid Navarro, Jean Oh

As robots across domains start collaborating with humans in shared environments, algorithms that enable them to reason over human intent are important to achieve safe interplay. In…

cs.RO2026

DYMO-Hair: Generalizable Volumetric Dynamics Modeling for Robot Hair Manipulation

Chengyang Zhao, Uksang Yoo, Arkadeep Narayan Chaudhury +4

Hair care is an essential daily activity, yet it remains inaccessible to individuals with limited mobility and challenging for autonomous robot systems due to the fine-grained phys…

cs.HC2024

How is the Pilot Doing: VTOL Pilot Workload Estimation by Multimodal Machine Learning on Psycho-physiological Signals

Jong Hoon Park, Lawrence Chen, Ian Higgins +11

Vertical take-off and landing (VTOL) aircraft do not require a prolonged runway, thus allowing them to land almost anywhere. In recent years, their flexibility has made them popula…

cs.CV2021

Content Masked Loss: Human-Like Brush Stroke Planning in a Reinforcement Learning Painting Agent

Peter Schaldenbrand, Jean Oh

The objective of most Reinforcement Learning painting agents is to minimize the loss between a target image and the paint canvas. Human painter artistry emphasizes important featur…

cs.CV2023

Towards Equitable Representation in Text-to-Image Synthesis Models with the Cross-Cultural Understanding Benchmark (CCUB) Dataset

Zhixuan Liu, Youeun Shin, Beverley-Claire Okogwu +5

It has been shown that accurate representation in media improves the well-being of the people who consume it. By contrast, inaccurate representations can negatively affect viewers…

cs.RO2025

RCG: Safety-Critical Scenario Generation for Robust Autonomous Driving via Real-World Crash Grounding

Benjamin Stoler, Juliet Yang, Jonathan Francis +1

Safety-critical scenarios are essential for training and evaluating autonomous driving (AD) systems, yet remain extremely rare in real-world driving datasets. To address this, we p…

cs.RO2022

Distribution-aware Goal Prediction and Conformant Model-based Planning for Safe Autonomous Driving

Jonathan Francis, Bingqing Chen, Weiran Yao +2

The feasibility of collecting a large amount of expert demonstrations has inspired growing research interests in learning-to-drive settings, where models learn by imitating the dri…

cs.RO2023

Follow The Rules: Online Signal Temporal Logic Tree Search for Guided Imitation Learning in Stochastic Domains

Jasmine Jerry Aloor, Jay Patrikar, Parv Kapoor +2

Seamlessly integrating rules in Learning-from-Demonstrations (LfD) policies is a critical requirement to enable the real-world deployment of AI agents. Recently, Signal Temporal Lo…

cs.RO2025

Geodesic Tracing-Based Kinematic Integration of Rolling and Sliding Contact on Manifold Meshes for Dexterous In-Hand Manipulation

Sunyu Wang, Arjun S. Lakshmipathy, Jean Oh +1

Reasoning about rolling and sliding contact, or roll-slide contact for short, is critical for dexterous manipulation tasks that involve intricate geometries. But existing works on…