papers

Publications (54)

cs.RO2026

asRoBallet: Closing the Sim2Real Gap via Friction-Aware Reinforcement Learning for Underactuated Spherical Dynamics

Fang Wan, Guangyi Huang, Tianyu Wu +5

We introduce asRoBallet, to the best of our knowledge, the first end-to-end reinforcement learning (RL) locomotion policy deployed on a humanoid ballbot hardware platform. Historic…

cs.CV2022

Integrally Migrating Pre-trained Transformer Encoder-decoders for Visual Object Detection

Feng Liu, Xiaosong Zhang, Zhiliang Peng +4

Modern object detectors have taken the advantages of backbone networks pre-trained on large scale datasets. Except for the backbone networks, however, other components such as the…

cs.RO2026

A Dataset and Benchmark for Robotic Cloth Unfolding Grasp Selection: The ICRA 2024 Cloth Competition

Victor-Louis De Gusseme, Thomas Lips, Remko Proesmans +59

Robotic cloth manipulation suffers from a lack of standardized benchmarks and shared datasets for evaluating and comparing different approaches. To address this, we created a bench…

cs.CV2019

Utilizing the Instability in Weakly Supervised Object Detection

Yan Gao, Boxiao Liu, Nan Guo +4

Weakly supervised object detection (WSOD) focuses on training object detector with only image-level annotations, and is challenging due to the gap between the supervision and the o…

cs.RO2020

Design of an Optoelectronically Innervated Gripper for Rigid-Soft Interactive Grasping

Linhan Yang, Xudong Han, Weijie Guo +4

Over the past few decades, efforts have been made towards robust robotic grasping, and therefore dexterous manipulation. The soft gripper has shown their potential in robust graspi…

stat.AP2024

Utilising high-dimensional data in randomised clinical trials: a review of methods and practice

Svetlana Cherlin, Theophile Bigirumurame, Michael J Grayling +6

Introduction: Even in effectively conducted randomised trials, the probability of a successful study remains relatively low. With recent advances in the next-generation sequencing…

cs.RO2023

Underwater Intention Recognition using Head Motion and Throat Vibration for Supernumerary Robotic Assistance

Yuqin Guo, Rongzheng Zhang, Wanghongjie Qiu +3

This study presents a multi-modal mechanism for recognizing human intentions while diving underwater, aiming to achieve natural human-robot interactions through an underwater super…

cs.RO2020

Hybrid Actuator Design for a Gait Augmentation Wearable

Fang Wan, Zheng Wang, Brooke Franchuk +3

We describe a fluidic actuator design that replaces the sealed chamber of a hydraulic cylinder using a soft actuator to provide compliant linear compression with a large force ($\g…

cs.RO2025

MagiClaw: A Dual-Use, Vision-Based Soft Gripper for Bridging the Human Demonstration to Robotic Deployment Gap

Tianyu Wu, Xudong Han, Haoran Sun +4

The transfer of manipulation skills from human demonstration to robotic execution is often hindered by a "domain gap" in sensing and morphology. This paper introduces MagiClaw, a v…

cs.RO2023

Jigsaw-based Benchmarking for Learning Robotic Manipulation

Xiaobo Liu, Fang Wan, Sheng Ge +3

Benchmarking provides experimental evidence of the scientific baseline to enhance the progression of fundamental research, which is also applicable to robotics. In this paper, we p…

cs.RO2024

Evolutionary Morphology Towards Overconstrained Locomotion via Large-Scale, Multi-Terrain Deep Reinforcement Learning

Yenan Chen, Chuye Zhang, Pengxi Gu +10

While the animals' Fin-to-Limb evolution has been well-researched in biology, such morphological transformation remains under-adopted in the modern design of advanced robotic limbs…

cs.RO2025

One-DoF Robotic Design of Overconstrained Limbs with Energy-Efficient, Self-Collision-Free Motion

Yuping Gu, Bangchao Huang, Haoran Sun +6

While it is expected to build robotic limbs with multiple degrees of freedom (DoF) inspired by nature, a single DoF design remains fundamental, providing benefits that include, but…

cs.CV2021

Multiple instance active learning for object detection

Tianning Yuan, Fang Wan, Mengying Fu +4

Despite the substantial progress of active learning for image recognition, there still lacks an instance-level active learning method specified for object detection. In this paper,…

cs.RO2021

Learning-based Optoelectronically Innervated Tactile Finger for Rigid-Soft Interactive Grasping

Linhan Yang, Xudong Han, Weijie Guo +3

This paper presents a novel design of a soft tactile finger with omni-directional adaptation using multi-channel optical fibers for rigid-soft interactive grasping. Machine learnin…

cs.CV2026

Vorch-Omni: Multi-Task Orchestration of Sight and Sound

Vorch Team, Xiaoyu Chen, Yang Ding +25

Recent advances in generative video modeling have enabled diverse generation, reference-based synthesis, extension, and editing, but existing approaches often rely on fragmented ta…

cs.AI2017

Logical Learning Through a Hybrid Neural Network with Auxiliary Inputs

Fang Wan, Chaoyang Song

The human reasoning process is seldom a one-way process from an input leading to an output. Instead, it often involves a systematic deduction by ruling out other possible outcomes…

cs.CV2019

Min-Entropy Latent Model for Weakly Supervised Object Detection

Fang Wan, Pengxu Wei, Zhenjun Han +2

Weakly supervised object detection is a challenging task when provided with image category supervision but required to learn, at the same time, object locations and object detector…

cs.CV2019

SIXray : A Large-scale Security Inspection X-ray Benchmark for Prohibited Item Discovery in Overlapping Images

Caijing Miao, Lingxi Xie, Fang Wan +4

In this paper, we present a large-scale dataset and establish a baseline for prohibited item discovery in Security Inspection X-ray images. Our dataset, named SIXray, consists of 1…

stat.ME2026

Data-driven controlled subgroup selection in clinical trials

Manuel M. Müller, Björn Bornkamp, Frank Bretz +7

Subgroup selection in clinical trials is essential for identifying patient groups that react differently to a treatment, thereby enabling personalised medicine. In particular, subg…

cs.RO2020

Robotic Cane as a Soft SuperLimb for Elderly Sit-to-Stand Assistance

Xia Wu, Haiyuan Liu, Ziqi Liu +6

Many researchers have identified robotics as a potential solution to the aging population faced by many developed and developing countries. If so, how should we address the cogniti…

cs.LG2025

Can Data-Driven Dynamics Reveal Hidden Physics? There Is A Need for Interpretable Neural Operators

Wenhan Gao, Jian Luo, Fang Wan +4

Recently, neural operators have emerged as powerful tools for learning mappings between function spaces, enabling data-driven simulations of complex dynamics. Despite their success…

cs.RO2023

Active Surface with Passive Omni-Directional Adaptation of Soft Polyhedral Fingers for In-Hand Manipulation

Sen Li, Fang Wan, Chaoyang Song

Track systems effectively distribute loads, augmenting traction and maneuverability on unstable terrains, leveraging their expansive contact areas. This tracked locomotion capabili…

cs.RO2020

Rigid-Soft Interactive Learning for Robust Grasping

Linhan Yang, Fang Wan, Haokun Wang +4

Inspired by widely used soft fingers on grasping, we propose a method of rigid-soft interactive learning, aiming at reducing the time of data collection. In this paper, we classify…

cs.RO2024

Proprioceptive Learning with Soft Polyhedral Networks

Xiaobo Liu, Xudong Han, Wei Hong +2

Proprioception is the "sixth sense" that detects limb postures with motor neurons. It requires a natural integration between the musculoskeletal systems and sensory receptors, whic…

stat.ME2022

Confidence Sets for a level set in linear regression

Fang Wan, Wei Liu, Frank Bretz

Regression modeling is the workhorse of statistics and there is a vast literature on estimation of the regression function. It is realized in recent years that in regression analys…

cs.CV2024

Ray Denoising: Depth-aware Hard Negative Sampling for Multi-view 3D Object Detection

Feng Liu, Tengteng Huang, Qianjing Zhang +5

Multi-view 3D object detection systems often struggle with generating precise predictions due to the challenges in estimating depth from images, increasing redundant and incorrect…

cs.RO2024

One Fling to Goal: Environment-aware Dynamics for Goal-conditioned Fabric Flinging

Linhan Yang, Lei Yang, Haoran Sun +5

Fabric manipulation dynamically is commonly seen in manufacturing and domestic settings. While dynamically manipulating a fabric piece to reach a target state is highly efficient,…

cs.CL2025

Geometric-Mean Policy Optimization

Yuzhong Zhao, Yue Liu, Junpeng Liu +9

Group Relative Policy Optimization (GRPO) has significantly enhanced the reasoning capability of large language models by optimizing the arithmetic mean of token-level rewards. Unf…

cs.RO2024

Proprioceptive State Estimation for Amphibious Tactile Sensing

Ning Guo, Xudong Han, Shuqiao Zhong +5

This paper presents a novel vision-based proprioception approach for a soft robotic finger that can estimate and reconstruct tactile interactions in both terrestrial and aquatic en…

cs.RO2020

Reconfigurable Design for Omni-adaptive Grasp Learning

Fang Wan, Haokun Wang, Jiyuan Wu +3

The engineering design of robotic grippers presents an ample design space for optimization towards robust grasping. In this paper, we adopt the reconfigurable design of the robotic…

cs.RO2020

DeepClaw: A Robotic Hardware Benchmarking Platform for Learning Object Manipulation

Fang Wan, Haokun Wang, Xiaobo Liu +2

We present DeepClaw as a reconfigurable benchmark of robotic hardware and task hierarchy for robot learning. The DeepClaw benchmark aims at a mechatronics perspective of the robot…

cs.RO2023

Autoencoding a Soft Touch to Learn Grasping from On-land to Underwater

Ning Guo, Xudong Han, Xiaobo Liu +6

Robots play a critical role as the physical agent of human operators in exploring the ocean. However, it remains challenging to grasp objects reliably while fully submerging under…

cs.CV2025

DynRefer: Delving into Region-level Multimodal Tasks via Dynamic Resolution

Yuzhong Zhao, Feng Liu, Yue Liu +4

One fundamental task of multimodal models is to translate referred image regions to human preferred language descriptions. Existing methods, however, ignore the resolution adaptabi…

cs.CV2025

Thinking with Images via Self-Calling Agent

Wenxi Yang, Yuzhong Zhao, Fang Wan +1

Thinking-with-images paradigms have showcased remarkable visual reasoning capability by integrating visual information as dynamic elements into the Chain-of-Thought (CoT). However,…

cs.CV2021

TS-CAM: Token Semantic Coupled Attention Map for Weakly Supervised Object Localization

Wei Gao, Fang Wan, Xingjia Pan +5

Weakly supervised object localization (WSOL) is a challenging problem when given image category labels but requires to learn object localization models. Optimizing a convolutional…

cs.RO2020

A Lobster-inspired Robotic Glove for Hand Rehabilitation

Yaohui Chen, Sing Le, Qiao Chu Tan +3

This paper presents preliminary results of the design, development, and evaluation of a hand rehabilitation glove fabricated using lobster-inspired hybrid design with rigid and sof…

cs.CV2024

Close the Sim2real Gap via Physically-based Structured Light Synthetic Data Simulation

Kaixin Bai, Lei Zhang, Zhaopeng Chen +2

Despite the substantial progress in deep learning, its adoption in industrial robotics projects remains limited, primarily due to challenges in data acquisition and labeling. Previ…

cs.RO2020

Scalable Tactile Sensing for an Omni-adaptive Soft Robot Finger

Zeyi Yang, Sheng Ge, Fang Wan +2

Robotic fingers made of soft material and compliant structures usually lead to superior adaptation when interacting with the unstructured physical environment. In this paper, we pr…

cs.RO2020

A Reconfigurable Hybrid Actuator with Rigid and Soft Components

Yaohui Chen, Sing Le, Qiao Chu Tan +3

Classical rigid-bodied robotic systems are presented with proven success in theoretical development and industrial applications, are recently challenged by the emergence of soft ro…

cs.LG2026

Uncertainty-Calibrated Diffusion for Reliable 3D Molecular Graph Generation

Fang Wan, Jingxiang Qu, Yi Liu

Bayesian inference provides a principled framework for modeling epistemic uncertainty in neural networks by treating predictions as distributions rather than deterministic values.…

cs.RO2024

On Flange-based 3D Hand-Eye Calibration for Soft Robotic Tactile Welding

Xudong Han, Ning Guo, Yu Jie +3

This paper investigates the direct application of standardized designs on the robot for conducting robot hand-eye calibration by employing 3D scanners with collaborative robots. Th…

cs.CV2023

Generative Prompt Model for Weakly Supervised Object Localization

Yuzhong Zhao, Qixiang Ye, Weijia Wu +2

Weakly supervised object localization (WSOL) remains challenging when learning object localization models from image category labels. Conventional methods that discriminatively tra…

cs.RO2023

Describing Robots from Design to Learning: Towards an Interactive Lifecycle Representation of Robots

Nuofan Qiu, Fang Wan, Chaoyang Song

The robot development process is divided into several stages, which create barriers to the exchange of information between these different stages. We advocate for an interactive li…

cs.RO2024

Overconstrained Locomotion

Haoran Sun, Bangchao Huang, Zishang Zhang +11

This paper studies the design, control, and learning of a novel robotic limb that produces overconstrained locomotion by employing the Bennett linkage for motion generation, capabl…

cs.CV2024

ControlCap: Controllable Region-level Captioning

Yuzhong Zhao, Yue Liu, Zonghao Guo +4

Region-level captioning is challenged by the caption degeneration issue, which refers to that pre-trained multimodal models tend to predict the most frequent captions but miss the…

cs.CV2020

Weakly-Supervised Action Localization with Expectation-Maximization Multi-Instance Learning

Zhekun Luo, Devin Guillory, Baifeng Shi +4

Weakly-supervised action localization requires training a model to localize the action segments in the video given only video level action label. It can be solved under the Multipl…

cs.CV2025

Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model

Feng Liu, Shiwei Zhang, Xiaofeng Wang +6

As a fundamental backbone for video generation, diffusion models are challenged by low inference speed due to the sequential nature of denoising. Previous methods speed up the mode…

cs.CV2024

Correspondence-Guided SfM-Free 3D Gaussian Splatting for NVS

Wei Sun, Xiaosong Zhang, Fang Wan +4

Novel View Synthesis (NVS) without Structure-from-Motion (SfM) pre-processed camera poses--referred to as SfM-free methods--is crucial for promoting rapid response capabilities and…

cs.CV2020

Domain Contrast for Domain Adaptive Object Detection

Feng Liu, Xiaoxong Zhang, Fang Wan +2

We present Domain Contrast (DC), a simple yet effective approach inspired by contrastive learning for training domain adaptive detectors. DC is deduced from the error bound minimiz…

cs.RO2026

Multi-Layered Reasoning from a Single Viewpoint for Learning See-Through Grasping

Fang Wan, Chaoyang Song

Sensory substitution enables biological systems to perceive stimuli that are typically perceived by another organ, which is inspirational for physical agents. Multimodal perception…

cs.CV2024

Evaluation of Text-to-Video Generation Models: A Dynamics Perspective

Mingxiang Liao, Hannan Lu, Xinyu Zhang +6

Comprehensive and constructive evaluation protocols play an important role in the development of sophisticated text-to-video (T2V) generation models. Existing evaluation protocols…

stat.ME2018

Subgroup analysis of treatment effects for misclassified biomarkers with time-to-event data

Fang Wan, Andrew C. Titman, Thomas F. Jaki

Analysing subgroups defined by biomarkers is of increasing importance in clinical research. In some situations the biomarker is subject to misclassification error, meaning the true…

cs.CV2019

FreeAnchor: Learning to Match Anchors for Visual Object Detection

Xiaosong Zhang, Fang Wan, Chang Liu +2

Modern CNN-based object detectors assign anchors for ground-truth objects under the restriction of object-anchor Intersection-over-Unit (IoU). In this study, we propose a learning-…

cs.CV2019

C-MIL: Continuation Multiple Instance Learning for Weakly Supervised Object Detection

Fang Wan, Chang Liu, Wei Ke +3

Weakly supervised object detection (WSOD) is a challenging task when provided with image category supervision but required to simultaneously learn object locations and object detec…