papers

Publications (32)

cs.RO2025

Dense Policy: Bidirectional Autoregressive Learning of Actions

Yue Su, Xinyu Zhan, Hongjie Fang +5

Mainstream visuomotor policies predominantly rely on generative models for holistic action prediction, while current autoregressive policies, predicting the next token or chunk, ha…

cs.CV2021

OMAD: Object Model with Articulated Deformations for Pose Estimation and Retrieval

Han Xue, Liu Liu, Wenqiang Xu +2

Articulated objects are pervasive in daily life. However, due to the intrinsic high-DoF structure, the joint states of the articulated objects are hard to be estimated. To model ar…

cs.CV2026

Prior-Guided Residual Diffusion: Calibrated and Efficient Medical Image Segmentation

Fuyou Mao, Beining Wu, Yanfeng Jiang +3

Ambiguity in medical image segmentation calls for models that capture full conditional distributions rather than a single point estimate. We present Prior-Guided Residual Diffusion…

math.FA2024

Further analysis of weighted integral inequalities for improved exponential stability analysis of time delay neural networks systems

Yuanyuan Zhang, Han Xue, Kachong Lao +3

This work investigates the exponential stability of neural networks (NNs) systems with time delays. By considering orthogonal polynomials with weighted terms, a new weighted integr…

cs.RO2023

RFUniverse: A Multiphysics Simulation Platform for Embodied AI

Haoyuan Fu, Wenqiang Xu, Ruolin Ye +7

Multiphysics phenomena, the coupling effects involving different aspects of physics laws, are pervasive in the real world and can often be encountered when performing everyday hous…

cs.CV2020

Toward Fine-grained Facial Expression Manipulation

Jun Ling, Han Xue, Li Song +3

Facial expression manipulation aims at editing facial expression with a given condition. Previous methods edit an input image under the guidance of a discrete emotion label or abso…

cond-mat.mtrl-sci2024

Operando probing of nanocracking in CuO-derived Cu during CO electroreduction

Jiawei Wan, Ershuai Liu, Woong Choi +20

Identifying and controlling active sites in electrocatalysis remains a grand challenge due to restructuring of catalysts in the complex chemical environments during operation. Inac…

cs.RO2025

Unleashing Humanoid Reaching Potential via Real-world-Ready Skill Space

Zhikai Zhang, Chao Chen, Han Xue +6

Humans possess a large reachable space in the 3D world, enabling interaction with objects at varying heights and distances. However, realizing such large-space reaching on humanoid…

cond-mat.soft2023

Ballistic Ejection of Microdroplets from Overpacked Interfacial Assemblies

Xuefei Wu, Gautam Bordia, Robert Streubel +11

Spontaneous emulsification, resulting from the assembly and accumulation of surfactants at liquid-liquid interfaces, is an interfacial instability where microdroplets are generated…

cs.CV2024

GenN2N: Generative NeRF2NeRF Translation

Xiangyue Liu, Han Xue, Kunming Luo +2

We present GenN2N, a unified NeRF-to-NeRF translation framework for various NeRF translation tasks such as text-driven NeRF editing, colorization, super-resolution, inpainting, etc…

cs.CV2023

Visual-Tactile Sensing for In-Hand Object Reconstruction

Wenqiang Xu, Zhenjun Yu, Han Xue +3

Tactile sensing is one of the modalities humans rely on heavily to perceive the world. Working with vision, this modality refines local geometry structure, measures deformation at…

cs.RO2026

Rethinking Camera Choice: An Empirical Study on Fisheye Camera Properties in Robotic Manipulation

Han Xue, Nan Min, Xiaotong Liu +5

The adoption of fisheye cameras in robotic manipulation, driven by their exceptionally wide Field of View (FoV), is rapidly outpacing a systematic understanding of their downstream…

cs.RO2026

Never Too Late for Force: Accelerating VLA Post-Training with Reactive Force Injection

Yi Wang, Wendi Chen, Zimo Wen +8

The paper introduces LIFT, a post‑training method that adds reactive force feedback to pretrained vision‑language‑action policies, enabling them to handle contact‑rich manipulation…

#vision-language-action#force feedback#post-training#contact-rich manipulation
cs.RO2025

SOE: Sample-Efficient Robot Policy Self-Improvement via On-Manifold Exploration

Yang Jin, Jun Lv, Han Xue +3

Intelligent agents progress by continually refining their capabilities through actively exploring environments. Yet robot policies often lack sufficient exploration capability due…

cs.CV2021

Region-aware Adaptive Instance Normalization for Image Harmonization

Jun Ling, Han Xue, Li Song +2

Image composition plays a common but important role in photo editing. To acquire photo-realistic composite images, one must adjust the appearance and visual style of the foreground…

cs.CV2020

Dense RepPoints: Representing Visual Objects with Dense Point Sets

Ze Yang, Yinghao Xu, Han Xue +5

We present a new object representation, called Dense RepPoints, that utilizes a large set of points to describe an object at multiple levels, including both box level and pixel lev…

cs.RO2025

ImplicitRDP: An End-to-End Visual-Force Diffusion Policy with Structural Slow-Fast Learning

Wendi Chen, Han Xue, Yi Wang +6

Human-level contact-rich manipulation relies on the distinct roles of two key modalities: vision provides spatially rich but temporally slow global context, while force sensing cap…

cs.RO2023

UniFolding: Towards Sample-efficient, Scalable, and Generalizable Robotic Garment Folding

Han Xue, Yutong Li, Wenqiang Xu +3

This paper explores the development of UniFolding, a sample-efficient, scalable, and generalizable robotic system for unfolding and folding various garments. UniFolding employs the…

cs.RO2025

DeformPAM: Data-Efficient Learning for Long-horizon Deformable Object Manipulation via Preference-based Action Alignment

Wendi Chen, Han Xue, Fangyuan Zhou +2

In recent years, imitation learning has made progress in the field of robotic manipulation. However, it still faces challenges when addressing complex long-horizon tasks with defor…

cs.RO2026

Learning Athletic Humanoid Tennis Skills from Imperfect Human Motion Data

Zhikai Zhang, Haofei Lu, Yunrui Lian +12

Human athletes demonstrate versatile and highly-dynamic tennis skills to successfully conduct competitive rallies with a high-speed tennis ball. However, reproducing such behaviors…

cs.CV2021

Towards Real-World Category-level Articulation Pose Estimation

Liu Liu, Han Xue, Wenqiang Xu +2

Human life is populated with articulated objects. Current Category-level Articulation Pose Estimation (CAPE) methods are studied under the single-instance setting with a fixed kine…

cs.RO2026

Right-Side-Out: Learning Zero-Shot Sim-to-Real Garment Reversal

Chang Yu, Siyu Ma, Wenxin Du +9

Turning garments right-side out is a challenging manipulation task: it is highly dynamic, entails rapid contact changes, and is subject to severe visual occlusion. We introduce Rig…

cs.RO2026

ActiveGlasses: Learning Manipulation with Active Vision from Ego-centric Human Demonstration

Yanwen Zou, Chenyang Shi, Wenye Yu +5

Large-scale real-world robot data collection is a prerequisite for bringing robots into everyday deployment. However, existing pipelines often rely on specialized handheld devices…

cs.CV2024

In-Context Translation: Towards Unifying Image Recognition, Processing, and Generation

Han Xue, Qianru Sun, Li Song +2

We propose In-Context Translation (ICT), a general learning framework to unify visual recognition (e.g., semantic segmentation), low-level image processing (e.g., denoising), and c…

cs.RO2026

FA-RDP: A Frequency-Adaptive Reactive Diffusion Policy for Contact-Rich Manipulation

Lifeng Zhuo, Wendi Chen, Han Xue +4

The paper introduces FA-RDP, a diffusion‑based policy that adapts its inference frequency during contact‑rich manipulation, using a multi‑frequency visual‑force transformer and a m…

#contact-rich manipulation#diffusion policies#frequency adaptation#multimodality
cs.CV2023

Freestyle Layout-to-Image Synthesis

Han Xue, Zhiwu Huang, Qianru Sun +2

Typical layout-to-image synthesis (LIS) models generate images for a closed set of semantic classes, e.g., 182 common objects in COCO-Stuff. In this work, we explore the freestyle…

cs.RO2026

Collision-Free Humanoid Traversal in Cluttered Indoor Scenes

Han Xue, Sikai Liang, Zhikai Zhang +7

We study the problem of collision-free humanoid traversal in cluttered indoor scenes, such as hurdling over objects scattered on the floor, crouching under low-hanging obstacles, o…

cs.CV2025

GarmentTracking: Category-Level Garment Pose Tracking

Han Xue, Wenqiang Xu, Jieyi Zhang +5

Garments are important to humans. A visual system that can estimate and track the complete garment pose can be useful for many downstream tasks and real-world applications. In this…

cs.CV2024

PoseTalk: Text-and-Audio-based Pose Control and Motion Refinement for One-Shot Talking Head Generation

Jun Ling, Yiwen Wang, Han Xue +2

While previous audio-driven talking head generation (THG) methods generate head poses from driving audio, the generated poses or lips cannot match the audio well or are not editabl…

cs.RO2026

RoboPocket: Improve Robot Policies Instantly with Your Phone

Junjie Fang, Wendi Chen, Han Xue +7

Scaling imitation learning is fundamentally constrained by the efficiency of data collection. While handheld interfaces have emerged as a scalable solution for in-the-wild data acq…

cs.RO2025

Track Any Motions under Any Disturbances

Zhikai Zhang, Jun Guo, Chao Chen +10

A foundational humanoid motion tracker is expected to be able to track diverse, highly dynamic, and contact-rich motions. More importantly, it needs to operate stably in real-world…

cs.RO2025

Reactive Diffusion Policy: Slow-Fast Visual-Tactile Policy Learning for Contact-Rich Manipulation

Han Xue, Jieji Ren, Wendi Chen +5

Humans can accomplish complex contact-rich tasks using vision and touch, with highly reactive capabilities such as fast response to external changes and adaptive control of contact…