papers

Publications (52)

physics.plasm-ph2026

VEQ: a fast parametric Grad--Shafranov solver for fixed-boundary tokamak equilibria with flexible source profiles

Ruohan Zhang, Huasheng Xie, Yueyan Li +3

Veloce EQuilibrium (VEQ) is a compact parametric framework for tokamak modeling workflows that repeatedly query continuous fixed-boundary equilibria at low latency. The VEQPy imple…

cs.RO2026

Skills in Weights, Memory in Code: Hybrid Learning for Memory-Dependent Robot Manipulation

Yunhao Zhao, Zhenyang Ni, Haoyang Chen +2

Modern vision-language-action (VLA) policies have acquired broad manipulation skills, but typically generate each action chunk from the current observation or a short fixed-length…

cs.RO2024

BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation

Chengshu Li, Ruohan Zhang, Josiah Wong +32

We present BEHAVIOR-1K, a comprehensive simulation benchmark for human-centered robotics. BEHAVIOR-1K includes two components, guided and motivated by the results of an extensive s…

cs.IR2022

A Survey on Cross-domain Recommendation: Taxonomies, Methods, and Future Directions

Tianzi Zang, Yanmin Zhu, Haobing Liu +2

Traditional recommendation systems are faced with two long-standing obstacles, namely, data sparsity and cold-start problems, which promote the emergence and development of Cross-D…

cs.AI2023

Mini-BEHAVIOR: A Procedurally Generated Benchmark for Long-horizon Decision-Making in Embodied AI

Emily Jin, Jiaheng Hu, Zhuoyi Huang +4

We present Mini-BEHAVIOR, a novel benchmark for embodied AI that challenges agents to use reasoning and decision-making skills to solve complex activities that resemble everyday hu…

cs.RO2023

MimicPlay: Long-Horizon Imitation Learning by Watching Human Play

Chen Wang, Linxi Fan, Jiankai Sun +5

Imitation learning from human demonstrations is a promising paradigm for teaching robots manipulation skills in the real world. However, learning complex long-horizon tasks often r…

cs.AI2021

Recent Advances in Leveraging Human Guidance for Sequential Decision-Making Tasks

Ruohan Zhang, Faraz Torabi, Garrett Warnell +1

A longstanding goal of artificial intelligence is to create artificial agents capable of learning to perform tasks that require sequential decision making. Importantly, while it is…

cs.RO2024

TRANSIC: Sim-to-Real Policy Transfer by Learning from Online Correction

Yunfan Jiang, Chen Wang, Ruohan Zhang +2

Learning in simulation and transferring the learned policy to the real world has the potential to enable generalist robots. The key challenge of this approach is to address simulat…

cs.RO2026

Requirement-Driven Design of Whole-Body Social Tactile Sensing via Virtual Human-Robot Interaction

Dakarai Crowder, Ruohan Zhang, Alexis E. Block +1

The paper introduces a requirement‑driven framework that uses VR‑based haptic interaction data to determine where and how densely tactile sensors should be placed on a humanoid rob…

#tactile sensing#human-robot interaction#virtual reality#sensor placement
cs.AI2025

ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction

Qineng Wang, Wenlong Huang, Yu Zhou +8

Embodied cognition argues that intelligence arises from sensorimotor interaction rather than passive observation. It raises an intriguing question: do modern vision-language models…

cs.LG2023

Interaction Modeling with Multiplex Attention

Fan-Yun Sun, Isaac Kauvar, Ruohan Zhang +4

Modeling multi-agent systems requires understanding how agents interact. Such systems are often difficult to model because they can involve a variety of types of interactions that…

cs.RO2026

MoMaGen: Generating Demonstrations under Soft and Hard Constraints for Multi-Step Bimanual Mobile Manipulation

Chengshu Li, Mengdi Xu, Arpit Bahety +11

Imitation learning from large-scale, diverse human demonstrations has been shown to be effective for training robots, but collecting such data is costly and time-consuming. This ch…

cs.RO2025

UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation

Yihe Tang, Wenlong Huang, Yingke Wang +5

Understanding fine-grained object affordances is imperative for robots to manipulate objects in unstructured environments given open-ended task instructions. However, existing meth…

cs.RO2023

VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models

Wenlong Huang, Chen Wang, Ruohan Zhang +3

Large language models (LLMs) are shown to possess a wealth of actionable knowledge that can be extracted for robot manipulation in the form of reasoning and planning. Despite the p…

cs.RO2023

NOIR: Neural Signal Operated Intelligent Robots for Everyday Activities

Ruohan Zhang, Sharon Lee, Minjune Hwang +11

We present Neural Signal Operated Intelligent Robots (NOIR), a general-purpose, intelligent brain-robot interface system that enables humans to command robots to perform everyday a…

cs.RO2024

DexCap: Scalable and Portable Mocap Data Collection System for Dexterous Manipulation

Chen Wang, Haochen Shi, Weizhuo Wang +3

Imitation learning from human hand motion data presents a promising avenue for imbuing robots with human-like dexterity in real-world manipulation tasks. Despite this potential, su…

cs.LG2024

MARPLE: A Benchmark for Long-Horizon Inference

Emily Jin, Zhuoyi Huang, Jan-Philipp Fränken +7

Reconstructing past events requires reasoning across long time horizons. To figure out what happened, we need to use our prior knowledge about the world and human behavior and draw…

cs.RO2024

ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation

Wenlong Huang, Chen Wang, Yunzhu Li +2

Representing robotic manipulation tasks as constraints that associate the robot and the environment is a promising way to encode desired robot behaviors. However, it remains unclea…

cs.RO2026

SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation

Nadun Ranawaka, Josiah Wong, Wei-Lin Pai +15

Training and evaluating robot policies in the real world is costly and difficult to scale. We introduce SimFoundry, a modular and automated system for zero-shot real-to-sim scene c…

cs.RO2025

Dream2Flow: Bridging Video Generation and Open-World Manipulation with 3D Object Flow

Karthik Dharmarajan, Wenlong Huang, Jiajun Wu +2

Generative video modeling has emerged as a compelling tool to zero-shot reason about plausible physical interactions for open-world manipulation. Yet, it remains a challenge to tra…

cs.CV2026

CAGE-SGG: Counterfactual Active Graph Evidence for Open-Vocabulary Scene Graph Generation

Suiyang Guang, Chenyu Liu, Ruohan Zhang +1

Open-vocabulary scene graph generation (SGG) aims to describe visual scenes with flexible and fine-grained relation phrases beyond a fixed predicate vocabulary. While recent vision…

cs.CL2025

Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making

Manling Li, Shiyu Zhao, Qineng Wang +12

We aim to evaluate Large Language Models (LLMs) for embodied decision making. While a significant body of work has been leveraging LLMs for decision making in embodied environments…

cs.AI2026

SENTINEL: A Multi-Level Formal Framework for Safety Evaluation of Foundation Model-based Embodied Agents

Simon Sinong Zhan, Yao Liu, Philip Wang +13

We present SENTINEL, a framework for formally evaluating the physical safety of foundation model (FM)-based embodied agents. SENTINEL is the first to provide multi-level safety eva…

cs.RO2025

Learning Compositional Behaviors from Demonstration and Language

Weiyu Liu, Neil Nie, Ruohan Zhang +2

We introduce Behavior from Language and Demonstration (BLADE), a framework for long-horizon robotic manipulation by integrating imitation learning and model-based planning. BLADE l…

cs.RO2024

Automated Creation of Digital Cousins for Robust Policy Learning

Tianyuan Dai, Josiah Wong, Yunfan Jiang +5

Training robot policies in the real world can be unsafe, costly, and difficult to scale. Simulation serves as an inexpensive and potentially limitless source of training data, but…

cs.CV2018

AGIL: Learning Attention from Human for Visuomotor Tasks

Ruohan Zhang, Zhuode Liu, Luxin Zhang +4

When intelligent agents learn visuomotor behaviors from human demonstrations, they may benefit from knowing where the human is allocating visual attention, which can be inferred fr…

cs.AI2026

Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?

Pingyue Zhang, Zihan Huang, Yue Wang +11

Spatial embodied intelligence requires agents to act to acquire information under partial observability. While multimodal foundation models excel at passive perception, their capac…

cs.CV2024

Partial-View Object View Synthesis via Filtered Inversion

Fan-Yun Sun, Jonathan Tremblay, Valts Blukis +11

We propose Filtering Inversion (FINV), a learning framework and optimization process that predicts a renderable 3D object representation from one or few partial views. FINV address…

cs.RO2026

FruitTouch: A Perceptive Gripper for Gentle and Scalable Fruit Harvesting

Ruohan Zhang, Mohammad Amin Mirzaee, Wenzhen Yuan

The automation of fruit harvesting has gained increasing significance in response to rising labor shortages. A sensorized gripper is a key component of this process, which must be…

cs.CV2024

BEHAVIOR Vision Suite: Customizable Dataset Generation via Simulation

Yunhao Ge, Yihe Tang, Jiashu Xu +20

The systematic evaluation and understanding of computer vision models under varying conditions require large amounts of data with comprehensive and customized labels, which real-wo…

cs.LG2019

Atari-HEAD: Atari Human Eye-Tracking and Demonstration Dataset

Ruohan Zhang, Calen Walshe, Zhuode Liu +6

Large-scale public datasets have been shown to benefit research in multiple areas of modern artificial intelligence. For decision-making research that requires human data, high-qua…

cs.AI2025

Federated Cross-Training Learners for Robust Generalization under Data Heterogeneity

Zhuang Qi, Lei Meng, Ruohan Zhang +5

Federated learning benefits from cross-training strategies, which enables models to train on data from distinct sources to improve generalization capability. However, due to inhere…

cs.AI2023

Task-Driven Graph Attention for Hierarchical Relational Object Navigation

Michael Lingelbach, Chengshu Li, Minjune Hwang +6

Embodied AI agents in large scenes often need to navigate to find objects. In this work, we study a naturally emerging variant of the object navigation task, hierarchical relationa…

cs.RO2026

PneuGelSight: Soft Robotic Vision-Based Proprioception and Tactile Sensing

Ruohan Zhang, Uksang Yoo, Yichen Li +2

Soft pneumatic robot manipulators are popular in industrial and human-interactive applications due to their compliance and flexibility. However, deploying them in real-world scenar…

cs.LG2023

Modeling Dynamic Environments with Scene Graph Memory

Andrey Kurenkov, Michael Lingelbach, Tanmay Agarwal +7

Embodied AI agents that search for objects in large environments such as households often need to make efficient decisions by predicting object locations based on partial informati…

cs.AI2019

Leveraging Human Guidance for Deep Reinforcement Learning Tasks

Ruohan Zhang, Faraz Torabi, Lin Guan +2

Reinforcement learning agents can learn to solve sequential decision tasks by interacting with the environment. Human knowledge of how to solve these tasks can be incorporated usin…

cs.RO2026

EmboAlign: Aligning Video Generation with Compositional Constraints for Zero-Shot Manipulation

Gehao Zhang, Zhenyang Ni, Payal Mohapatra +3

Video generative models (VGMs) pretrained on large-scale internet data can produce temporally coherent rollout videos that capture rich object dynamics, offering a compelling found…

cs.AI2021

Widening the Pipeline in Human-Guided Reinforcement Learning with Explanation and Context-Aware Data Augmentation

Lin Guan, Mudit Verma, Sihang Guo +2

Human explanation (e.g., in terms of feature importance) has been recently used to extend the communication channel between human and agent in interactive machine learning. Under t…

cs.RO2026

StereoPolicy: Improving Robotic Manipulation Policies via Stereo Perception

Evans Han, Yunfan Jiang, Yingke Wang +6

Recent advances in robot imitation learning have produced powerful visuomotor policies that manipulate diverse objects from visual inputs. However, monocular observations lack dept…

cs.IR2025

Purely Semantic Indexing for LLM-based Generative Recommendation and Retrieval

Ruohan Zhang, Jiacheng Li, Julian McAuley +1

Semantic identifiers (IDs) have proven effective in adapting large language models for generative recommendation and retrieval. However, existing methods often suffer from semantic…

cs.CV2025

CRAFT: Designing Creative and Functional 3D Objects

Michelle Guo, Mia Tang, Hannah Cha +3

For designing a wide range of everyday objects, the design process should be aware of both the human body and the underlying semantics of the design specification. However, these t…

cs.LG2021

Machine versus Human Attention in Deep Reinforcement Learning Tasks

Sihang Guo, Ruohan Zhang, Bo Liu +4

Deep reinforcement learning (RL) algorithms are powerful tools for solving visuomotor decision tasks. However, the trained models are often difficult to interpret, because they are…

cs.LG2021

Efficiently Guiding Imitation Learning Agents with Human Gaze

Akanksha Saran, Ruohan Zhang, Elaine Schaertl Short +1

Human gaze is known to be an intention-revealing signal in human demonstrations of tasks. In this work, we use gaze cues from human demonstrators to enhance the performance of agen…

cs.RO2025

Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Embodiment Collaboration, Abby O'Neill, Abdul Rehman +291

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, thi…

cs.RO2026

IMPASTO: Integrating Model-Based Planning with Learned Dynamics Models for Robotic Oil Painting Reproduction

Yingke Wang, Hao Li, Yifeng Zhu +6

Robotic reproduction of oil paintings using soft brushes and pigments requires force-sensitive control of deformable tools, prediction of brushstroke effects, and multi-step stroke…

cs.RO2024

TeleMoMa: A Modular and Versatile Teleoperation System for Mobile Manipulation

Shivin Dass, Wensi Ai, Yuqian Jiang +6

A critical bottleneck limiting imitation learning in robotics is the lack of data. This problem is more severe in mobile manipulation, where collecting demonstrations is harder tha…

cs.RO2023

Primitive Skill-based Robot Learning from Human Evaluative Feedback

Ayano Hiranaka, Minjune Hwang, Sharon Lee +4

Reinforcement learning (RL) algorithms face significant challenges when dealing with long-horizon robot manipulation tasks in real-world environments due to sample inefficiency and…

cs.RO2026

Behavior Uncloning: Distilling Mode Redirection into Policy Weights without Inference-Time Steering

Hao Wang, Jiuzhou Lei, Dayou Li +5

Behavior-cloned policies often learn multiple behavior modes from demonstration datasets, including modes that are unsafe or otherwise undesired at deployment. For example, a polic…

cs.CV2026

OpenLongTail: Generative Scaling of Long-Tail Driving Data

Lulin Liu, Nuo Chen, Yan Wang +15

Scaling robust driving policies is fundamentally bottlenecked by the scarcity of edge cases in curated datasets. While the real world continuously captures these critical events, s…

cs.RO2025

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models

Chen Wang, Fei Xia, Wenhao Yu +6

Learning to perform manipulation tasks from human videos is a promising approach for teaching robots. However, many manipulation tasks require changing control parameters during ta…

cs.RO2025

BEHAVIOR Robot Suite: Streamlining Real-World Whole-Body Manipulation for Everyday Household Activities

Yunfan Jiang, Ruohan Zhang, Josiah Wong +7

Real-world household tasks present significant challenges for mobile manipulation robots. An analysis of existing robotics benchmarks reveals that successful task performance hinge…

cs.LG2020

An initial attempt of combining visual selective attention with deep reinforcement learning

Liu Yuezhang, Ruohan Zhang, Dana H. Ballard

Visual attention serves as a means of feature selection mechanism in the perceptual system. Motivated by Broadbent's leaky filter model of selective attention, we evaluate how such…