Publications (52)
VEQ: a fast parametric Grad--Shafranov solver for fixed-boundary tokamak equilibria with flexible source profiles
Ruohan Zhang, Huasheng Xie, Yueyan Li +3
Veloce EQuilibrium (VEQ) is a compact parametric framework for tokamak modeling workflows that repeatedly query continuous fixed-boundary equilibria at low latency. The VEQPy imple…
Skills in Weights, Memory in Code: Hybrid Learning for Memory-Dependent Robot Manipulation
Yunhao Zhao, Zhenyang Ni, Haoyang Chen +2
Modern vision-language-action (VLA) policies have acquired broad manipulation skills, but typically generate each action chunk from the current observation or a short fixed-length…
BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation
Chengshu Li, Ruohan Zhang, Josiah Wong +32
We present BEHAVIOR-1K, a comprehensive simulation benchmark for human-centered robotics. BEHAVIOR-1K includes two components, guided and motivated by the results of an extensive s…
A Survey on Cross-domain Recommendation: Taxonomies, Methods, and Future Directions
Tianzi Zang, Yanmin Zhu, Haobing Liu +2
Traditional recommendation systems are faced with two long-standing obstacles, namely, data sparsity and cold-start problems, which promote the emergence and development of Cross-D…
Mini-BEHAVIOR: A Procedurally Generated Benchmark for Long-horizon Decision-Making in Embodied AI
Emily Jin, Jiaheng Hu, Zhuoyi Huang +4
We present Mini-BEHAVIOR, a novel benchmark for embodied AI that challenges agents to use reasoning and decision-making skills to solve complex activities that resemble everyday hu…
MimicPlay: Long-Horizon Imitation Learning by Watching Human Play
Chen Wang, Linxi Fan, Jiankai Sun +5
Imitation learning from human demonstrations is a promising paradigm for teaching robots manipulation skills in the real world. However, learning complex long-horizon tasks often r…
Recent Advances in Leveraging Human Guidance for Sequential Decision-Making Tasks
Ruohan Zhang, Faraz Torabi, Garrett Warnell +1
A longstanding goal of artificial intelligence is to create artificial agents capable of learning to perform tasks that require sequential decision making. Importantly, while it is…
TRANSIC: Sim-to-Real Policy Transfer by Learning from Online Correction
Yunfan Jiang, Chen Wang, Ruohan Zhang +2
Learning in simulation and transferring the learned policy to the real world has the potential to enable generalist robots. The key challenge of this approach is to address simulat…
Requirement-Driven Design of Whole-Body Social Tactile Sensing via Virtual Human-Robot Interaction
Dakarai Crowder, Ruohan Zhang, Alexis E. Block +1
The paper introduces a requirement‑driven framework that uses VR‑based haptic interaction data to determine where and how densely tactile sensors should be placed on a humanoid rob…
ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction
Qineng Wang, Wenlong Huang, Yu Zhou +8
Embodied cognition argues that intelligence arises from sensorimotor interaction rather than passive observation. It raises an intriguing question: do modern vision-language models…
Interaction Modeling with Multiplex Attention
Fan-Yun Sun, Isaac Kauvar, Ruohan Zhang +4
Modeling multi-agent systems requires understanding how agents interact. Such systems are often difficult to model because they can involve a variety of types of interactions that…
MoMaGen: Generating Demonstrations under Soft and Hard Constraints for Multi-Step Bimanual Mobile Manipulation
Chengshu Li, Mengdi Xu, Arpit Bahety +11
Imitation learning from large-scale, diverse human demonstrations has been shown to be effective for training robots, but collecting such data is costly and time-consuming. This ch…
UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation
Yihe Tang, Wenlong Huang, Yingke Wang +5
Understanding fine-grained object affordances is imperative for robots to manipulate objects in unstructured environments given open-ended task instructions. However, existing meth…
VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
Wenlong Huang, Chen Wang, Ruohan Zhang +3
Large language models (LLMs) are shown to possess a wealth of actionable knowledge that can be extracted for robot manipulation in the form of reasoning and planning. Despite the p…
NOIR: Neural Signal Operated Intelligent Robots for Everyday Activities
Ruohan Zhang, Sharon Lee, Minjune Hwang +11
We present Neural Signal Operated Intelligent Robots (NOIR), a general-purpose, intelligent brain-robot interface system that enables humans to command robots to perform everyday a…
DexCap: Scalable and Portable Mocap Data Collection System for Dexterous Manipulation
Chen Wang, Haochen Shi, Weizhuo Wang +3
Imitation learning from human hand motion data presents a promising avenue for imbuing robots with human-like dexterity in real-world manipulation tasks. Despite this potential, su…
MARPLE: A Benchmark for Long-Horizon Inference
Emily Jin, Zhuoyi Huang, Jan-Philipp Fränken +7
Reconstructing past events requires reasoning across long time horizons. To figure out what happened, we need to use our prior knowledge about the world and human behavior and draw…
ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation
Wenlong Huang, Chen Wang, Yunzhu Li +2
Representing robotic manipulation tasks as constraints that associate the robot and the environment is a promising way to encode desired robot behaviors. However, it remains unclea…
SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation
Nadun Ranawaka, Josiah Wong, Wei-Lin Pai +15
Training and evaluating robot policies in the real world is costly and difficult to scale. We introduce SimFoundry, a modular and automated system for zero-shot real-to-sim scene c…
Dream2Flow: Bridging Video Generation and Open-World Manipulation with 3D Object Flow
Karthik Dharmarajan, Wenlong Huang, Jiajun Wu +2
Generative video modeling has emerged as a compelling tool to zero-shot reason about plausible physical interactions for open-world manipulation. Yet, it remains a challenge to tra…
CAGE-SGG: Counterfactual Active Graph Evidence for Open-Vocabulary Scene Graph Generation
Suiyang Guang, Chenyu Liu, Ruohan Zhang +1
Open-vocabulary scene graph generation (SGG) aims to describe visual scenes with flexible and fine-grained relation phrases beyond a fixed predicate vocabulary. While recent vision…
Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
Manling Li, Shiyu Zhao, Qineng Wang +12
We aim to evaluate Large Language Models (LLMs) for embodied decision making. While a significant body of work has been leveraging LLMs for decision making in embodied environments…
SENTINEL: A Multi-Level Formal Framework for Safety Evaluation of Foundation Model-based Embodied Agents
Simon Sinong Zhan, Yao Liu, Philip Wang +13
We present SENTINEL, a framework for formally evaluating the physical safety of foundation model (FM)-based embodied agents. SENTINEL is the first to provide multi-level safety eva…
Learning Compositional Behaviors from Demonstration and Language
Weiyu Liu, Neil Nie, Ruohan Zhang +2
We introduce Behavior from Language and Demonstration (BLADE), a framework for long-horizon robotic manipulation by integrating imitation learning and model-based planning. BLADE l…
Automated Creation of Digital Cousins for Robust Policy Learning
Tianyuan Dai, Josiah Wong, Yunfan Jiang +5
Training robot policies in the real world can be unsafe, costly, and difficult to scale. Simulation serves as an inexpensive and potentially limitless source of training data, but…
AGIL: Learning Attention from Human for Visuomotor Tasks
Ruohan Zhang, Zhuode Liu, Luxin Zhang +4
When intelligent agents learn visuomotor behaviors from human demonstrations, they may benefit from knowing where the human is allocating visual attention, which can be inferred fr…
Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?
Pingyue Zhang, Zihan Huang, Yue Wang +11
Spatial embodied intelligence requires agents to act to acquire information under partial observability. While multimodal foundation models excel at passive perception, their capac…
Partial-View Object View Synthesis via Filtered Inversion
Fan-Yun Sun, Jonathan Tremblay, Valts Blukis +11
We propose Filtering Inversion (FINV), a learning framework and optimization process that predicts a renderable 3D object representation from one or few partial views. FINV address…
FruitTouch: A Perceptive Gripper for Gentle and Scalable Fruit Harvesting
Ruohan Zhang, Mohammad Amin Mirzaee, Wenzhen Yuan
The automation of fruit harvesting has gained increasing significance in response to rising labor shortages. A sensorized gripper is a key component of this process, which must be…
BEHAVIOR Vision Suite: Customizable Dataset Generation via Simulation
Yunhao Ge, Yihe Tang, Jiashu Xu +20
The systematic evaluation and understanding of computer vision models under varying conditions require large amounts of data with comprehensive and customized labels, which real-wo…
Atari-HEAD: Atari Human Eye-Tracking and Demonstration Dataset
Ruohan Zhang, Calen Walshe, Zhuode Liu +6
Large-scale public datasets have been shown to benefit research in multiple areas of modern artificial intelligence. For decision-making research that requires human data, high-qua…
Federated Cross-Training Learners for Robust Generalization under Data Heterogeneity
Zhuang Qi, Lei Meng, Ruohan Zhang +5
Federated learning benefits from cross-training strategies, which enables models to train on data from distinct sources to improve generalization capability. However, due to inhere…
Task-Driven Graph Attention for Hierarchical Relational Object Navigation
Michael Lingelbach, Chengshu Li, Minjune Hwang +6
Embodied AI agents in large scenes often need to navigate to find objects. In this work, we study a naturally emerging variant of the object navigation task, hierarchical relationa…
PneuGelSight: Soft Robotic Vision-Based Proprioception and Tactile Sensing
Ruohan Zhang, Uksang Yoo, Yichen Li +2
Soft pneumatic robot manipulators are popular in industrial and human-interactive applications due to their compliance and flexibility. However, deploying them in real-world scenar…
Modeling Dynamic Environments with Scene Graph Memory
Andrey Kurenkov, Michael Lingelbach, Tanmay Agarwal +7
Embodied AI agents that search for objects in large environments such as households often need to make efficient decisions by predicting object locations based on partial informati…
Leveraging Human Guidance for Deep Reinforcement Learning Tasks
Ruohan Zhang, Faraz Torabi, Lin Guan +2
Reinforcement learning agents can learn to solve sequential decision tasks by interacting with the environment. Human knowledge of how to solve these tasks can be incorporated usin…
EmboAlign: Aligning Video Generation with Compositional Constraints for Zero-Shot Manipulation
Gehao Zhang, Zhenyang Ni, Payal Mohapatra +3
Video generative models (VGMs) pretrained on large-scale internet data can produce temporally coherent rollout videos that capture rich object dynamics, offering a compelling found…
Widening the Pipeline in Human-Guided Reinforcement Learning with Explanation and Context-Aware Data Augmentation
Lin Guan, Mudit Verma, Sihang Guo +2
Human explanation (e.g., in terms of feature importance) has been recently used to extend the communication channel between human and agent in interactive machine learning. Under t…
StereoPolicy: Improving Robotic Manipulation Policies via Stereo Perception
Evans Han, Yunfan Jiang, Yingke Wang +6
Recent advances in robot imitation learning have produced powerful visuomotor policies that manipulate diverse objects from visual inputs. However, monocular observations lack dept…
Purely Semantic Indexing for LLM-based Generative Recommendation and Retrieval
Ruohan Zhang, Jiacheng Li, Julian McAuley +1
Semantic identifiers (IDs) have proven effective in adapting large language models for generative recommendation and retrieval. However, existing methods often suffer from semantic…
CRAFT: Designing Creative and Functional 3D Objects
Michelle Guo, Mia Tang, Hannah Cha +3
For designing a wide range of everyday objects, the design process should be aware of both the human body and the underlying semantics of the design specification. However, these t…
Machine versus Human Attention in Deep Reinforcement Learning Tasks
Sihang Guo, Ruohan Zhang, Bo Liu +4
Deep reinforcement learning (RL) algorithms are powerful tools for solving visuomotor decision tasks. However, the trained models are often difficult to interpret, because they are…
Efficiently Guiding Imitation Learning Agents with Human Gaze
Akanksha Saran, Ruohan Zhang, Elaine Schaertl Short +1
Human gaze is known to be an intention-revealing signal in human demonstrations of tasks. In this work, we use gaze cues from human demonstrators to enhance the performance of agen…
Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Embodiment Collaboration, Abby O'Neill, Abdul Rehman +291
Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, thi…
IMPASTO: Integrating Model-Based Planning with Learned Dynamics Models for Robotic Oil Painting Reproduction
Yingke Wang, Hao Li, Yifeng Zhu +6
Robotic reproduction of oil paintings using soft brushes and pigments requires force-sensitive control of deformable tools, prediction of brushstroke effects, and multi-step stroke…
TeleMoMa: A Modular and Versatile Teleoperation System for Mobile Manipulation
Shivin Dass, Wensi Ai, Yuqian Jiang +6
A critical bottleneck limiting imitation learning in robotics is the lack of data. This problem is more severe in mobile manipulation, where collecting demonstrations is harder tha…
Primitive Skill-based Robot Learning from Human Evaluative Feedback
Ayano Hiranaka, Minjune Hwang, Sharon Lee +4
Reinforcement learning (RL) algorithms face significant challenges when dealing with long-horizon robot manipulation tasks in real-world environments due to sample inefficiency and…
Behavior Uncloning: Distilling Mode Redirection into Policy Weights without Inference-Time Steering
Hao Wang, Jiuzhou Lei, Dayou Li +5
Behavior-cloned policies often learn multiple behavior modes from demonstration datasets, including modes that are unsafe or otherwise undesired at deployment. For example, a polic…
OpenLongTail: Generative Scaling of Long-Tail Driving Data
Lulin Liu, Nuo Chen, Yan Wang +15
Scaling robust driving policies is fundamentally bottlenecked by the scarcity of edge cases in curated datasets. While the real world continuously captures these critical events, s…
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models
Chen Wang, Fei Xia, Wenhao Yu +6
Learning to perform manipulation tasks from human videos is a promising approach for teaching robots. However, many manipulation tasks require changing control parameters during ta…
BEHAVIOR Robot Suite: Streamlining Real-World Whole-Body Manipulation for Everyday Household Activities
Yunfan Jiang, Ruohan Zhang, Josiah Wong +7
Real-world household tasks present significant challenges for mobile manipulation robots. An analysis of existing robotics benchmarks reveals that successful task performance hinge…
An initial attempt of combining visual selective attention with deep reinforcement learning
Liu Yuezhang, Ruohan Zhang, Dana H. Ballard
Visual attention serves as a means of feature selection mechanism in the perceptual system. Motivated by Broadbent's leaky filter model of selective attention, we evaluate how such…