Publications (51)
LVDiffusor: Distilling Functional Rearrangement Priors from Large Models into Diffusor
Yiming Zeng, Mingdong Wu, Long Yang +4
Object rearrangement, a fundamental challenge in robotics, demands versatile strategies to handle diverse objects, configurations, and functional needs. To achieve this, the AI rob…
Learning Reference-Guided Exposure Correction with Hybrid Illumination Characteristics
Hao Ren, Zetong Bi, Zhaoliang Wan +1
We present HICNet, a reference-guided exposure correction framework. A lightweight, content-agnostic encoder distills each image into a compact illumination embedding capturing reg…
Pressure tunable magnetic skyrmion phase in Co8Zn8Mn4 single crystals
Zhun Li, Xinrun Mi, Xinming Wang +14
In a magnetic skyrmion phase, magnetic moments form vortex-like topological textures which are of both fundamental and industrial interests. In -Mn-type Co-Zn-Mn alloys, chrial…
RLMM-Flow: A Flow-based Mobile Manipulation Framework with Latent-Space Reinforcement Learning
Shuhang Wang, Ziming Li, Hui Cheng
The paper introduces RLMM-Flow, a framework that first learns a flow-based generative policy from expert demonstrations for mobile manipulation and then improves it with latent-spa…
OmniGS: Fast Radiance Field Reconstruction using Omnidirectional Gaussian Splatting
Longwei Li, Huajian Huang, Sai-Kit Yeung +1
Photorealistic reconstruction relying on 3D Gaussian Splatting has shown promising potential in various domains. However, the current 3D Gaussian Splatting system only supports rad…
FedDiv: Collaborative Noise Filtering for Federated Learning with Noisy Labels
Jichang Li, Guanbin Li, Hui Cheng +2
Federated learning with noisy labels (F-LNL) aims at seeking an optimal server model via collaborative distributed learning by aggregating multiple client models trained with local…
STaT: Resolving Shape Distortion in Non-Stationary Time Series via Tri-Modal Synergy
Hui Cheng, Jinsheng Guo, Zhenhao Weng +2
Recent research in time series forecasting frequently investigates the integration of textual and visual modalities with numerical models to better navigate non-stationary environm…
Recurrent 3D Pose Sequence Machines
Mude Lin, Liang Lin, Xiaodan Liang +2
3D human articulated pose recovery from monocular image sequences is very challenging due to the diverse appearances, viewpoints, occlusions, and also the human 3D pose is inherent…
Decentralized Global Connectivity Maintenance for Multi-Robot Navigation: A Reinforcement Learning Approach
Minghao Li, Yingrui Jie, Yang Kong +1
The problem of multi-robot navigation of connectivity maintenance is challenging in multi-robot applications. This work investigates how to navigate a multi-robot team in unknown e…
RAPID Hand Prototype: Design of an Affordable, Fully-Actuated Biomimetic Hand for Dexterous Teleoperation
Zhaoliang Wan, Zida Zhou, Zetong Bi +3
This paper addresses the scarcity of affordable, fully-actuated five-fingered hands for dexterous teleoperation, which is crucial for collecting large-scale real-robot data within…
Cluster Synchronization of Coupled Systems with Nonidentical Linear Dynamics
Zhongchang Liu, Wing Shing Wong, Hui Cheng
This paper considers the cluster synchronization problem of generic linear dynamical systems whose system models are distinct in different clusters. These nonidentical linear model…
X-NavDP: Generalizing Navigation Diffusion Policy to Novel Behavior and Embodiments with Group Q-score Reweighted Matching
Tianyu Yang, Yiming Zeng, Wenzhe Cai +5
The paper introduces X-NavDP, a diffusion-based visual navigation policy that is fine‑tuned with a novel Group Q-score Reweighted Matching (GQRM) reinforcement learning framework t…
Improving Knowledge Distillation via Transferring Learning Ability
Long Liu, Tong Li, Hui Cheng
Existing knowledge distillation methods generally use a teacher-student approach, where the student network solely learns from a well-trained teacher. However, this approach overlo…
Divide and Contrast: Source-free Domain Adaptation via Adaptive Contrastive Learning
Ziyi Zhang, Weikai Chen, Hui Cheng +4
We investigate a practical domain adaptation task, called source-free domain adaptation (SFUDA), where the source-pretrained model is adapted to the target domain without access to…
Prior Does Matter: Visual Navigation via Denoising Diffusion Bridge Models
Hao Ren, Yiming Zeng, Zetong Bi +3
Recent advancements in diffusion-based imitation learning, which show impressive performance in modeling multimodal distributions and training stability, have led to substantial pr…
Safe Learning-based Tracking Control for Quadrotors under Wind Disturbances
Lei Zheng, Rui Yang, Jiesen Pan +1
Enforcing safety on precise trajectory tracking is critical for aerial robotics subject to wind disturbances. In this paper, we present a learning-based safety-preserving cascaded…
Instance-Aware Representation Learning and Association for Online Multi-Person Tracking
Hefeng Wu, Yafei Hu, Keze Wang +3
Multi-Person Tracking (MPT) is often addressed within the detection-to-association paradigm. In such approaches, human detections are first extracted in every frame and person traj…
Less is More: Adaptive Curriculum Learning for Thyroid Nodule Diagnosis
Haifan Gong, Hui Cheng, Yifan Xie +4
Thyroid nodule classification aims at determining whether the nodule is benign or malignant based on a given ultrasound image. However, the label obtained by the cytological biopsy…
GIC: Gaussian-Informed Continuum for Physical Property Identification and Simulation
Junhao Cai, Yuji Yang, Weihao Yuan +5
This paper studies the problem of estimating physical properties (system identification) through visual observations. To facilitate geometry-aware guidance in physical property est…
Safe Learning-based Gradient-free Model Predictive Control Based on Cross-entropy Method
Lei Zheng, Rui Yang, Zhixuan Wu +2
In this paper, a safe and learning-based control framework for model predictive control (MPC) is proposed to optimize nonlinear systems with a non-differentiable objective function…
LSTM-CF: Unifying Context Modeling and Fusion with LSTMs for RGB-D Scene Labeling
Zhen Li, Yukang Gan, Xiaodan Liang +3
Semantic labeling of RGB-D scenes is crucial to many intelligent applications including perceptual robotics. It generates pixelwise and fine-grained label maps from simultaneously…
MetaGrasp: Data Efficient Grasping by Affordance Interpreter Network
Junhao Cai, Hui Cheng, Zhanpeng Zhang +1
Data-driven approach for grasping shows significant advance recently. But these approaches usually require much training data. To increase the efficiency of grasping data collectio…
360Loc: A Dataset and Benchmark for Omnidirectional Visual Localization with Cross-device Queries
Huajian Huang, Changkun Liu, Yipeng Zhu +3
Portable 360 cameras are becoming a cheap and efficient tool to establish large visual databases. By capturing omnidirectional views of a scene, these cameras could expedit…
End-to-end Decentralized Multi-robot Navigation in Unknown Complex Environments via Deep Reinforcement Learning
Juntong Lin, Xuyun Yang, Peiwei Zheng +1
In this paper, a novel deep reinforcement learning (DRL)-based method is proposed to navigate the robot team through unknown complex environments, where the geometric centroid of t…
Vehicular Communications for 5G Cooperative Small Cell Networks
Xiaohu Ge, Hui Cheng, Guoqiang Mao +2
The cooperative transmission is an effective approach for vehicular communications to improve the wireless transmission capacity and reliability in the fifth generation (5G) small…
Learning-based Predictive Path Following Control for Nonlinear Systems Under Uncertain Disturbances
Rui Yang, Lei Zheng, Jiesen Pan +1
Accurate path following is challenging for autonomous robots operating in uncertain environments. Adaptive and predictive control strategies are crucial for a nonlinear robotic sys…
Safe Learning-Based Feedback Linearization Tracking Control for Nonlinear System with Event-Triggered Model Update
Zhixuan Wu, Rui Yang, Lei Zheng +1
Learning-based methods are powerful in handling complex scenarios. However, it is still challenging to use learning-based methods under uncertain environments while stability, safe…
Learning-Based Safety-Stability-Driven Control for Safety-Critical Systems under Model Uncertainties
Lei Zheng, Jiesen Pan, Rui Yang +2
Safety and tracking stability are crucial for safety-critical systems such as self-driving cars, autonomous mobile robots, industrial manipulators. To efficiently control safety-cr…
DHRCL:Training Code LLMs with Dense Hierarchical Rewards and Curriculum Learning
Shuhang Wang, Ziming Li, Hui Cheng
The paper introduces DHRCL, a reinforcement‑learning framework for code‑focused large language models that uses a hierarchy of dense rewards (syntax, execution, unit‑test pass, and…
H2-Mapping: Real-time Dense Mapping Using Hierarchical Hybrid Representation
Chenxing Jiang, Hanwen Zhang, Peize Liu +4
Constructing a high-quality dense map in real-time is essential for robotics, AR/VR, and digital twins applications. As Neural Radiance Field (NeRF) greatly improves the mapping pe…
Knowledge-Guided Recurrent Neural Network Learning for Task-Oriented Action Prediction
Liang Lin, Lili Huang, Tianshui Chen +2
This paper aims at task-oriented action prediction, i.e., predicting a sequence of actions towards accomplishing a specific task under a certain scene, which is a new problem in co…
Fisher-Preserving Guidance: Training-Free Manifold Constraints for Safe Diffusion Control
Hao Ren, Zetong Bi, Yiming Zeng +5
Diffusion models are effective for waypoint prediction in visual navigation, but standard sampling and test time guidance can produce unreliable or inefficient trajectories when up…
OPG-Policy: Occluded Push-Grasp Policy Learning with Amodal Segmentation
Hao Ding, Yiming Zeng, Zhaoliang Wan +1
Goal-oriented grasping in dense clutter, a fundamental challenge in robotics, demands an adaptive policy to handle occluded target objects and diverse configurations. Previous meth…
Star-Searcher: A Complete and Efficient Aerial System for Autonomous Target Search in Complex Unknown Environments
Yiming Luo, Zixuan Zhuang, Neng Pan +5
This paper tackles the challenge of autonomous target search using unmanned aerial vehicles (UAVs) in complex unknown environments. To fill the gap in systematic approaches for thi…
Depth Extraction from Videos Using Geometric Context and Occlusion Boundaries
S. Hussain Raza, Omar Javed, Aveek Das +3
We present an algorithm to estimate depth in dynamic video scenes. We propose to learn and infer depth in videos from appearance, motion, occlusion boundaries, and geometric contex…
Unidirectional-Road-Network-Based Global Path Planning for Cleaning Robots in Semi-Structured Environments
Yong Li, Hui Cheng
Practical global path planning is critical for commercializing cleaning robots working in semi-structured environments. In the literature, global path planning methods for free spa…
APP: A* Post-Processing Algorithm for Robots with Bidirectional Shortcut and Path Perturbation
Yong Li, Hui Cheng
Paths generated by A* and other graph-search-based planners are widely used in the robotic field. Due to the restricted node-expansion directions, the resulting paths are usually n…
Deep Reasoning with Knowledge Graph for Social Relationship Understanding
Zhouxia Wang, Tianshui Chen, Jimmy Ren +3
Social relationships (e.g., friends, couple etc.) form the basis of the social network in our daily life. Automatically interpreting such relationships bears a great potential for…
SC-OmniGS: Self-Calibrating Omnidirectional Gaussian Splatting
Huajian Huang, Yingshu Chen, Longwei Li +4
360-degree cameras streamline data collection for radiance field 3D reconstruction by capturing comprehensive scene data. However, traditional radiance field methods do not address…
STRNet: Visual Navigation with Spatio-Temporal Representation through Dynamic Graph Aggregation
Hao Ren, Zetong Bi, Yiming Zeng +3
Visual navigation requires the robot to reach a specified goal such as an image, based on a sequence of first-person visual observations. While recent learning-based approaches hav…
GET: Goal-directed Exploration and Targeting for Large-Scale Unknown Environments
Lanxiang Zheng, Ruidong Mei, Mingxin Wei +2
Object search in large-scale, unstructured environments remains a fundamental challenge in robotics, particularly in dynamic or expansive settings such as outdoor autonomous explor…
Volumetric-based Contact Point Detection for 7-DoF Grasping
Junhao Cai, Jingcheng Su, Zida Zhou +3
In this paper, we propose a novel grasp pipeline based on contact point detection on the truncated signed distance function (TSDF) volume to achieve closed-loop 7-degree-of-freedom…
Human Pose Estimation from Depth Images via Inference Embedded Multi-task Learning
Keze Wang, Shengfu Zhai, Hui Cheng +2
Human pose estimation (i.e., locating the body parts / joints of a person) is a fundamental problem in human-computer interaction and multimedia applications. Significant progress…
Bootstrapping Imitation Learning for Long-horizon Manipulation via Hierarchical Data Collection Space
Jinrong Yang, Kexun Chen, Zhuoling Li +9
Imitation learning (IL) with human demonstrations is a promising method for robotic manipulation tasks. While minimal demonstrations enable robotic action execution, achieving high…
VG-Swarm: A Vision-based Gene Regulation Network for UAVs Swarm Behavior Emergence
Yuwei Cai, Huanlin Li, Zhun Fan +6
Unmanned Aerial Vehicles (UAVs) dynamic encirclement is an emerging field with great potential. Researchers often get inspiration from biological systems, either from macro-world l…
VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception
Zhaoliang Wan, Yonggen Ling, Senlin Yi +9
This paper addresses the scarcity of large-scale datasets for accurate object-in-hand pose estimation, which is crucial for robotic in-hand manipulation within the ``Perception-Pla…
RAPID Hand: A Robust, Affordable, Perception-Integrated, Dexterous Manipulation Platform for Generalist Robot Autonomy
Zhaoliang Wan, Zetong Bi, Zida Zhou +7
This paper addresses the scarcity of low-cost but high-dexterity platforms for collecting real-world multi-fingered robot manipulation data towards generalist robot autonomy. To ac…
TOPP-DWR: Time-Optimal Path Parameterization of Differential-Driven Wheeled Robots Considering Piecewise-Constant Angular Velocity Constraints
Yong Li, Yujun Huang, Yi Chen +1
Differential-driven wheeled robots (DWR) represent the quintessential type of mobile robots and find extensive appli- cations across the robotic field. Most high-performance contro…
Photo-SLAM: Real-time Simultaneous Localization and Photorealistic Mapping for Monocular, Stereo, and RGB-D Cameras
Huajian Huang, Longwei Li, Hui Cheng +1
The integration of neural rendering and the SLAM system recently showed promising results in joint localization and photorealistic view reconstruction. However, existing methods, f…
Zero-Shot Event Detection by Multimodal Distributional Semantic Embedding of Videos
Mohamed Elhoseiny, Jingen Liu, Hui Cheng +2
We propose a new zero-shot Event Detection method by Multi-modal Distributional Semantic embedding of videos. Our model embeds object and action concepts as well as other available…
NaviDiffusor: Cost-Guided Diffusion Model for Visual Navigation
Yiming Zeng, Hao Ren, Shuhang Wang +2
Visual navigation, a fundamental challenge in mobile robotics, demands versatile policies to handle diverse environments. Classical methods leverage geometric solutions to minimize…