Publications (32)
Dense Policy: Bidirectional Autoregressive Learning of Actions
Yue Su, Xinyu Zhan, Hongjie Fang +5
Mainstream visuomotor policies predominantly rely on generative models for holistic action prediction, while current autoregressive policies, predicting the next token or chunk, ha…
OMAD: Object Model with Articulated Deformations for Pose Estimation and Retrieval
Han Xue, Liu Liu, Wenqiang Xu +2
Articulated objects are pervasive in daily life. However, due to the intrinsic high-DoF structure, the joint states of the articulated objects are hard to be estimated. To model ar…
Prior-Guided Residual Diffusion: Calibrated and Efficient Medical Image Segmentation
Fuyou Mao, Beining Wu, Yanfeng Jiang +3
Ambiguity in medical image segmentation calls for models that capture full conditional distributions rather than a single point estimate. We present Prior-Guided Residual Diffusion…
Further analysis of weighted integral inequalities for improved exponential stability analysis of time delay neural networks systems
Yuanyuan Zhang, Han Xue, Kachong Lao +3
This work investigates the exponential stability of neural networks (NNs) systems with time delays. By considering orthogonal polynomials with weighted terms, a new weighted integr…
RFUniverse: A Multiphysics Simulation Platform for Embodied AI
Haoyuan Fu, Wenqiang Xu, Ruolin Ye +7
Multiphysics phenomena, the coupling effects involving different aspects of physics laws, are pervasive in the real world and can often be encountered when performing everyday hous…
Toward Fine-grained Facial Expression Manipulation
Jun Ling, Han Xue, Li Song +3
Facial expression manipulation aims at editing facial expression with a given condition. Previous methods edit an input image under the guidance of a discrete emotion label or abso…
Operando probing of nanocracking in CuO-derived Cu during CO electroreduction
Jiawei Wan, Ershuai Liu, Woong Choi +20
Identifying and controlling active sites in electrocatalysis remains a grand challenge due to restructuring of catalysts in the complex chemical environments during operation. Inac…
Unleashing Humanoid Reaching Potential via Real-world-Ready Skill Space
Zhikai Zhang, Chao Chen, Han Xue +6
Humans possess a large reachable space in the 3D world, enabling interaction with objects at varying heights and distances. However, realizing such large-space reaching on humanoid…
Ballistic Ejection of Microdroplets from Overpacked Interfacial Assemblies
Xuefei Wu, Gautam Bordia, Robert Streubel +11
Spontaneous emulsification, resulting from the assembly and accumulation of surfactants at liquid-liquid interfaces, is an interfacial instability where microdroplets are generated…
GenN2N: Generative NeRF2NeRF Translation
Xiangyue Liu, Han Xue, Kunming Luo +2
We present GenN2N, a unified NeRF-to-NeRF translation framework for various NeRF translation tasks such as text-driven NeRF editing, colorization, super-resolution, inpainting, etc…
Visual-Tactile Sensing for In-Hand Object Reconstruction
Wenqiang Xu, Zhenjun Yu, Han Xue +3
Tactile sensing is one of the modalities humans rely on heavily to perceive the world. Working with vision, this modality refines local geometry structure, measures deformation at…
Rethinking Camera Choice: An Empirical Study on Fisheye Camera Properties in Robotic Manipulation
Han Xue, Nan Min, Xiaotong Liu +5
The adoption of fisheye cameras in robotic manipulation, driven by their exceptionally wide Field of View (FoV), is rapidly outpacing a systematic understanding of their downstream…
Never Too Late for Force: Accelerating VLA Post-Training with Reactive Force Injection
Yi Wang, Wendi Chen, Zimo Wen +8
The paper introduces LIFT, a post‑training method that adds reactive force feedback to pretrained vision‑language‑action policies, enabling them to handle contact‑rich manipulation…
SOE: Sample-Efficient Robot Policy Self-Improvement via On-Manifold Exploration
Yang Jin, Jun Lv, Han Xue +3
Intelligent agents progress by continually refining their capabilities through actively exploring environments. Yet robot policies often lack sufficient exploration capability due…
Region-aware Adaptive Instance Normalization for Image Harmonization
Jun Ling, Han Xue, Li Song +2
Image composition plays a common but important role in photo editing. To acquire photo-realistic composite images, one must adjust the appearance and visual style of the foreground…
Dense RepPoints: Representing Visual Objects with Dense Point Sets
Ze Yang, Yinghao Xu, Han Xue +5
We present a new object representation, called Dense RepPoints, that utilizes a large set of points to describe an object at multiple levels, including both box level and pixel lev…
ImplicitRDP: An End-to-End Visual-Force Diffusion Policy with Structural Slow-Fast Learning
Wendi Chen, Han Xue, Yi Wang +6
Human-level contact-rich manipulation relies on the distinct roles of two key modalities: vision provides spatially rich but temporally slow global context, while force sensing cap…
UniFolding: Towards Sample-efficient, Scalable, and Generalizable Robotic Garment Folding
Han Xue, Yutong Li, Wenqiang Xu +3
This paper explores the development of UniFolding, a sample-efficient, scalable, and generalizable robotic system for unfolding and folding various garments. UniFolding employs the…
DeformPAM: Data-Efficient Learning for Long-horizon Deformable Object Manipulation via Preference-based Action Alignment
Wendi Chen, Han Xue, Fangyuan Zhou +2
In recent years, imitation learning has made progress in the field of robotic manipulation. However, it still faces challenges when addressing complex long-horizon tasks with defor…
Learning Athletic Humanoid Tennis Skills from Imperfect Human Motion Data
Zhikai Zhang, Haofei Lu, Yunrui Lian +12
Human athletes demonstrate versatile and highly-dynamic tennis skills to successfully conduct competitive rallies with a high-speed tennis ball. However, reproducing such behaviors…
Towards Real-World Category-level Articulation Pose Estimation
Liu Liu, Han Xue, Wenqiang Xu +2
Human life is populated with articulated objects. Current Category-level Articulation Pose Estimation (CAPE) methods are studied under the single-instance setting with a fixed kine…
Right-Side-Out: Learning Zero-Shot Sim-to-Real Garment Reversal
Chang Yu, Siyu Ma, Wenxin Du +9
Turning garments right-side out is a challenging manipulation task: it is highly dynamic, entails rapid contact changes, and is subject to severe visual occlusion. We introduce Rig…
ActiveGlasses: Learning Manipulation with Active Vision from Ego-centric Human Demonstration
Yanwen Zou, Chenyang Shi, Wenye Yu +5
Large-scale real-world robot data collection is a prerequisite for bringing robots into everyday deployment. However, existing pipelines often rely on specialized handheld devices…
In-Context Translation: Towards Unifying Image Recognition, Processing, and Generation
Han Xue, Qianru Sun, Li Song +2
We propose In-Context Translation (ICT), a general learning framework to unify visual recognition (e.g., semantic segmentation), low-level image processing (e.g., denoising), and c…
FA-RDP: A Frequency-Adaptive Reactive Diffusion Policy for Contact-Rich Manipulation
Lifeng Zhuo, Wendi Chen, Han Xue +4
The paper introduces FA-RDP, a diffusion‑based policy that adapts its inference frequency during contact‑rich manipulation, using a multi‑frequency visual‑force transformer and a m…
Freestyle Layout-to-Image Synthesis
Han Xue, Zhiwu Huang, Qianru Sun +2
Typical layout-to-image synthesis (LIS) models generate images for a closed set of semantic classes, e.g., 182 common objects in COCO-Stuff. In this work, we explore the freestyle…
Collision-Free Humanoid Traversal in Cluttered Indoor Scenes
Han Xue, Sikai Liang, Zhikai Zhang +7
We study the problem of collision-free humanoid traversal in cluttered indoor scenes, such as hurdling over objects scattered on the floor, crouching under low-hanging obstacles, o…
GarmentTracking: Category-Level Garment Pose Tracking
Han Xue, Wenqiang Xu, Jieyi Zhang +5
Garments are important to humans. A visual system that can estimate and track the complete garment pose can be useful for many downstream tasks and real-world applications. In this…
PoseTalk: Text-and-Audio-based Pose Control and Motion Refinement for One-Shot Talking Head Generation
Jun Ling, Yiwen Wang, Han Xue +2
While previous audio-driven talking head generation (THG) methods generate head poses from driving audio, the generated poses or lips cannot match the audio well or are not editabl…
RoboPocket: Improve Robot Policies Instantly with Your Phone
Junjie Fang, Wendi Chen, Han Xue +7
Scaling imitation learning is fundamentally constrained by the efficiency of data collection. While handheld interfaces have emerged as a scalable solution for in-the-wild data acq…
Track Any Motions under Any Disturbances
Zhikai Zhang, Jun Guo, Chao Chen +10
A foundational humanoid motion tracker is expected to be able to track diverse, highly dynamic, and contact-rich motions. More importantly, it needs to operate stably in real-world…
Reactive Diffusion Policy: Slow-Fast Visual-Tactile Policy Learning for Contact-Rich Manipulation
Han Xue, Jieji Ren, Wendi Chen +5
Humans can accomplish complex contact-rich tasks using vision and touch, with highly reactive capabilities such as fast response to external changes and adaptive control of contact…