papers

Publications (51)

cs.RO2024

LVDiffusor: Distilling Functional Rearrangement Priors from Large Models into Diffusor

Yiming Zeng, Mingdong Wu, Long Yang +4

Object rearrangement, a fundamental challenge in robotics, demands versatile strategies to handle diverse objects, configurations, and functional needs. To achieve this, the AI rob…

cs.CV2026

Learning Reference-Guided Exposure Correction with Hybrid Illumination Characteristics

Hao Ren, Zetong Bi, Zhaoliang Wan +1

We present HICNet, a reference-guided exposure correction framework. A lightweight, content-agnostic encoder distills each image into a compact illumination embedding capturing reg…

cond-mat.mtrl-sci2024

Pressure tunable magnetic skyrmion phase in Co8Zn8Mn4 single crystals

Zhun Li, Xinrun Mi, Xinming Wang +14

In a magnetic skyrmion phase, magnetic moments form vortex-like topological textures which are of both fundamental and industrial interests. In -Mn-type Co-Zn-Mn alloys, chrial…

cs.RO2026

RLMM-Flow: A Flow-based Mobile Manipulation Framework with Latent-Space Reinforcement Learning

Shuhang Wang, Ziming Li, Hui Cheng

The paper introduces RLMM-Flow, a framework that first learns a flow-based generative policy from expert demonstrations for mobile manipulation and then improves it with latent-spa…

#mobile manipulation#flow-based policies#reinforcement learning#latent space
cs.CV2024

OmniGS: Fast Radiance Field Reconstruction using Omnidirectional Gaussian Splatting

Longwei Li, Huajian Huang, Sai-Kit Yeung +1

Photorealistic reconstruction relying on 3D Gaussian Splatting has shown promising potential in various domains. However, the current 3D Gaussian Splatting system only supports rad…

cs.CV2024

FedDiv: Collaborative Noise Filtering for Federated Learning with Noisy Labels

Jichang Li, Guanbin Li, Hui Cheng +2

Federated learning with noisy labels (F-LNL) aims at seeking an optimal server model via collaborative distributed learning by aggregating multiple client models trained with local…

cs.LG2026

STaT: Resolving Shape Distortion in Non-Stationary Time Series via Tri-Modal Synergy

Hui Cheng, Jinsheng Guo, Zhenhao Weng +2

Recent research in time series forecasting frequently investigates the integration of textual and visual modalities with numerical models to better navigate non-stationary environm…

cs.CV2017

Recurrent 3D Pose Sequence Machines

Mude Lin, Liang Lin, Xiaodan Liang +2

3D human articulated pose recovery from monocular image sequences is very challenging due to the diverse appearances, viewpoints, occlusions, and also the human 3D pose is inherent…

cs.RO2021

Decentralized Global Connectivity Maintenance for Multi-Robot Navigation: A Reinforcement Learning Approach

Minghao Li, Yingrui Jie, Yang Kong +1

The problem of multi-robot navigation of connectivity maintenance is challenging in multi-robot applications. This work investigates how to navigate a multi-robot team in unknown e…

cs.RO2025

RAPID Hand Prototype: Design of an Affordable, Fully-Actuated Biomimetic Hand for Dexterous Teleoperation

Zhaoliang Wan, Zida Zhou, Zetong Bi +3

This paper addresses the scarcity of affordable, fully-actuated five-fingered hands for dexterous teleoperation, which is crucial for collecting large-scale real-robot data within…

eess.SY2021

Cluster Synchronization of Coupled Systems with Nonidentical Linear Dynamics

Zhongchang Liu, Wing Shing Wong, Hui Cheng

This paper considers the cluster synchronization problem of generic linear dynamical systems whose system models are distinct in different clusters. These nonidentical linear model…

cs.RO2026

X-NavDP: Generalizing Navigation Diffusion Policy to Novel Behavior and Embodiments with Group Q-score Reweighted Matching

Tianyu Yang, Yiming Zeng, Wenzhe Cai +5

The paper introduces X-NavDP, a diffusion-based visual navigation policy that is fine‑tuned with a novel Group Q-score Reweighted Matching (GQRM) reinforcement learning framework t…

#visual navigation#diffusion policies#reinforcement learning#cross‑embodiment
cs.CV2023

Improving Knowledge Distillation via Transferring Learning Ability

Long Liu, Tong Li, Hui Cheng

Existing knowledge distillation methods generally use a teacher-student approach, where the student network solely learns from a well-trained teacher. However, this approach overlo…

cs.CV2022

Divide and Contrast: Source-free Domain Adaptation via Adaptive Contrastive Learning

Ziyi Zhang, Weikai Chen, Hui Cheng +4

We investigate a practical domain adaptation task, called source-free domain adaptation (SFUDA), where the source-pretrained model is adapted to the target domain without access to…

cs.RO2025

Prior Does Matter: Visual Navigation via Denoising Diffusion Bridge Models

Hao Ren, Yiming Zeng, Zetong Bi +3

Recent advancements in diffusion-based imitation learning, which show impressive performance in modeling multimodal distributions and training stability, have led to substantial pr…

cs.RO2021

Safe Learning-based Tracking Control for Quadrotors under Wind Disturbances

Lei Zheng, Rui Yang, Jiesen Pan +1

Enforcing safety on precise trajectory tracking is critical for aerial robotics subject to wind disturbances. In this paper, we present a learning-based safety-preserving cascaded…

cs.CV2019

Instance-Aware Representation Learning and Association for Online Multi-Person Tracking

Hefeng Wu, Yafei Hu, Keze Wang +3

Multi-Person Tracking (MPT) is often addressed within the detection-to-association paradigm. In such approaches, human detections are first extracted in every frame and person traj…

cs.CV2022

Less is More: Adaptive Curriculum Learning for Thyroid Nodule Diagnosis

Haifan Gong, Hui Cheng, Yifan Xie +4

Thyroid nodule classification aims at determining whether the nodule is benign or malignant based on a given ultrasound image. However, the label obtained by the cytological biopsy…

cs.CV2024

GIC: Gaussian-Informed Continuum for Physical Property Identification and Simulation

Junhao Cai, Yuji Yang, Weihao Yuan +5

This paper studies the problem of estimating physical properties (system identification) through visual observations. To facilitate geometry-aware guidance in physical property est…

cs.RO2022

Safe Learning-based Gradient-free Model Predictive Control Based on Cross-entropy Method

Lei Zheng, Rui Yang, Zhixuan Wu +2

In this paper, a safe and learning-based control framework for model predictive control (MPC) is proposed to optimize nonlinear systems with a non-differentiable objective function…

cs.CV2016

LSTM-CF: Unifying Context Modeling and Fusion with LSTMs for RGB-D Scene Labeling

Zhen Li, Yukang Gan, Xiaodan Liang +3

Semantic labeling of RGB-D scenes is crucial to many intelligent applications including perceptual robotics. It generates pixelwise and fine-grained label maps from simultaneously…

cs.RO2019

MetaGrasp: Data Efficient Grasping by Affordance Interpreter Network

Junhao Cai, Hui Cheng, Zhanpeng Zhang +1

Data-driven approach for grasping shows significant advance recently. But these approaches usually require much training data. To increase the efficiency of grasping data collectio…

cs.CV2024

360Loc: A Dataset and Benchmark for Omnidirectional Visual Localization with Cross-device Queries

Huajian Huang, Changkun Liu, Yipeng Zhu +3

Portable 360 cameras are becoming a cheap and efficient tool to establish large visual databases. By capturing omnidirectional views of a scene, these cameras could expedit…

cs.RO2019

End-to-end Decentralized Multi-robot Navigation in Unknown Complex Environments via Deep Reinforcement Learning

Juntong Lin, Xuyun Yang, Peiwei Zheng +1

In this paper, a novel deep reinforcement learning (DRL)-based method is proposed to navigate the robot team through unknown complex environments, where the geometric centroid of t…

cs.NI2016

Vehicular Communications for 5G Cooperative Small Cell Networks

Xiaohu Ge, Hui Cheng, Guoqiang Mao +2

The cooperative transmission is an effective approach for vehicular communications to improve the wireless transmission capacity and reliability in the fifth generation (5G) small…

cs.RO2022

Learning-based Predictive Path Following Control for Nonlinear Systems Under Uncertain Disturbances

Rui Yang, Lei Zheng, Jiesen Pan +1

Accurate path following is challenging for autonomous robots operating in uncertain environments. Adaptive and predictive control strategies are crucial for a nonlinear robotic sys…

cs.RO2022

Safe Learning-Based Feedback Linearization Tracking Control for Nonlinear System with Event-Triggered Model Update

Zhixuan Wu, Rui Yang, Lei Zheng +1

Learning-based methods are powerful in handling complex scenarios. However, it is still challenging to use learning-based methods under uncertain environments while stability, safe…

cs.RO2020

Learning-Based Safety-Stability-Driven Control for Safety-Critical Systems under Model Uncertainties

Lei Zheng, Jiesen Pan, Rui Yang +2

Safety and tracking stability are crucial for safety-critical systems such as self-driving cars, autonomous mobile robots, industrial manipulators. To efficiently control safety-cr…

cs.LG2026

DHRCL:Training Code LLMs with Dense Hierarchical Rewards and Curriculum Learning

Shuhang Wang, Ziming Li, Hui Cheng

The paper introduces DHRCL, a reinforcement‑learning framework for code‑focused large language models that uses a hierarchy of dense rewards (syntax, execution, unit‑test pass, and…

#code generation#reinforcement learning#curriculum learning#dense rewards
cs.RO2023

H2-Mapping: Real-time Dense Mapping Using Hierarchical Hybrid Representation

Chenxing Jiang, Hanwen Zhang, Peize Liu +4

Constructing a high-quality dense map in real-time is essential for robotics, AR/VR, and digital twins applications. As Neural Radiance Field (NeRF) greatly improves the mapping pe…

cs.CV2017

Knowledge-Guided Recurrent Neural Network Learning for Task-Oriented Action Prediction

Liang Lin, Lili Huang, Tianshui Chen +2

This paper aims at task-oriented action prediction, i.e., predicting a sequence of actions towards accomplishing a specific task under a certain scene, which is a new problem in co…

cs.RO2026

Fisher-Preserving Guidance: Training-Free Manifold Constraints for Safe Diffusion Control

Hao Ren, Zetong Bi, Yiming Zeng +5

Diffusion models are effective for waypoint prediction in visual navigation, but standard sampling and test time guidance can produce unreliable or inefficient trajectories when up…

cs.RO2025

OPG-Policy: Occluded Push-Grasp Policy Learning with Amodal Segmentation

Hao Ding, Yiming Zeng, Zhaoliang Wan +1

Goal-oriented grasping in dense clutter, a fundamental challenge in robotics, demands an adaptive policy to handle occluded target objects and diverse configurations. Previous meth…

cs.RO2024

Star-Searcher: A Complete and Efficient Aerial System for Autonomous Target Search in Complex Unknown Environments

Yiming Luo, Zixuan Zhuang, Neng Pan +5

This paper tackles the challenge of autonomous target search using unmanned aerial vehicles (UAVs) in complex unknown environments. To fill the gap in systematic approaches for thi…

cs.CV2015

Depth Extraction from Videos Using Geometric Context and Occlusion Boundaries

S. Hussain Raza, Omar Javed, Aveek Das +3

We present an algorithm to estimate depth in dynamic video scenes. We propose to learn and infer depth in videos from appearance, motion, occlusion boundaries, and geometric contex…

cs.RO2025

Unidirectional-Road-Network-Based Global Path Planning for Cleaning Robots in Semi-Structured Environments

Yong Li, Hui Cheng

Practical global path planning is critical for commercializing cleaning robots working in semi-structured environments. In the literature, global path planning methods for free spa…

cs.RO2025

APP: A* Post-Processing Algorithm for Robots with Bidirectional Shortcut and Path Perturbation

Yong Li, Hui Cheng

Paths generated by A* and other graph-search-based planners are widely used in the robotic field. Due to the restricted node-expansion directions, the resulting paths are usually n…

cs.CV2018

Deep Reasoning with Knowledge Graph for Social Relationship Understanding

Zhouxia Wang, Tianshui Chen, Jimmy Ren +3

Social relationships (e.g., friends, couple etc.) form the basis of the social network in our daily life. Automatically interpreting such relationships bears a great potential for…

cs.CV2025

SC-OmniGS: Self-Calibrating Omnidirectional Gaussian Splatting

Huajian Huang, Yingshu Chen, Longwei Li +4

360-degree cameras streamline data collection for radiance field 3D reconstruction by capturing comprehensive scene data. However, traditional radiance field methods do not address…

cs.CV2026

STRNet: Visual Navigation with Spatio-Temporal Representation through Dynamic Graph Aggregation

Hao Ren, Zetong Bi, Yiming Zeng +3

Visual navigation requires the robot to reach a specified goal such as an image, based on a sequence of first-person visual observations. While recent learning-based approaches hav…

cs.RO2025

GET: Goal-directed Exploration and Targeting for Large-Scale Unknown Environments

Lanxiang Zheng, Ruidong Mei, Mingxin Wei +2

Object search in large-scale, unstructured environments remains a fundamental challenge in robotics, particularly in dynamic or expansive settings such as outdoor autonomous explor…

cs.RO2022

Volumetric-based Contact Point Detection for 7-DoF Grasping

Junhao Cai, Jingcheng Su, Zida Zhou +3

In this paper, we propose a novel grasp pipeline based on contact point detection on the truncated signed distance function (TSDF) volume to achieve closed-loop 7-degree-of-freedom…

cs.CV2016

Human Pose Estimation from Depth Images via Inference Embedded Multi-task Learning

Keze Wang, Shengfu Zhai, Hui Cheng +2

Human pose estimation (i.e., locating the body parts / joints of a person) is a fundamental problem in human-computer interaction and multimedia applications. Significant progress…

cs.RO2025

Bootstrapping Imitation Learning for Long-horizon Manipulation via Hierarchical Data Collection Space

Jinrong Yang, Kexun Chen, Zhuoling Li +9

Imitation learning (IL) with human demonstrations is a promising method for robotic manipulation tasks. While minimal demonstrations enable robotic action execution, achieving high…

cs.RO2022

VG-Swarm: A Vision-based Gene Regulation Network for UAVs Swarm Behavior Emergence

Yuwei Cai, Huanlin Li, Zhun Fan +6

Unmanned Aerial Vehicles (UAVs) dynamic encirclement is an emerging field with great potential. Researchers often get inspiration from biological systems, either from macro-world l…

cs.RO2025

VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception

Zhaoliang Wan, Yonggen Ling, Senlin Yi +9

This paper addresses the scarcity of large-scale datasets for accurate object-in-hand pose estimation, which is crucial for robotic in-hand manipulation within the ``Perception-Pla…

cs.RO2025

RAPID Hand: A Robust, Affordable, Perception-Integrated, Dexterous Manipulation Platform for Generalist Robot Autonomy

Zhaoliang Wan, Zetong Bi, Zida Zhou +7

This paper addresses the scarcity of low-cost but high-dexterity platforms for collecting real-world multi-fingered robot manipulation data towards generalist robot autonomy. To ac…

cs.RO2025

TOPP-DWR: Time-Optimal Path Parameterization of Differential-Driven Wheeled Robots Considering Piecewise-Constant Angular Velocity Constraints

Yong Li, Yujun Huang, Yi Chen +1

Differential-driven wheeled robots (DWR) represent the quintessential type of mobile robots and find extensive appli- cations across the robotic field. Most high-performance contro…

cs.CV2024

Photo-SLAM: Real-time Simultaneous Localization and Photorealistic Mapping for Monocular, Stereo, and RGB-D Cameras

Huajian Huang, Longwei Li, Hui Cheng +1

The integration of neural rendering and the SLAM system recently showed promising results in joint localization and photorealistic view reconstruction. However, existing methods, f…

cs.CV2015

Zero-Shot Event Detection by Multimodal Distributional Semantic Embedding of Videos

Mohamed Elhoseiny, Jingen Liu, Hui Cheng +2

We propose a new zero-shot Event Detection method by Multi-modal Distributional Semantic embedding of videos. Our model embeds object and action concepts as well as other available…

cs.RO2025

NaviDiffusor: Cost-Guided Diffusion Model for Visual Navigation

Yiming Zeng, Hao Ren, Shuhang Wang +2

Visual navigation, a fundamental challenge in mobile robotics, demands versatile policies to handle diverse environments. Classical methods leverage geometric solutions to minimize…