papers

Publications (37)

cs.RO2026

Direct Contact-Tolerant Motion Planning With Vision Language Models

He Li, Jian Sun, Chengyang Li +4

Navigation in cluttered environments often requires robots to tolerate contact with movable or deformable objects to maintain efficiency. Existing contact-tolerant motion planning…

cs.RO2026

HPTune: Hierarchical Proactive Tuning for Collision-Free Model Predictive Control

Wei Zuo, Chengyang Li, Yikun Wang +5

Parameter tuning is a powerful approach to enhance adaptability in model predictive control (MPC) motion planners. However, existing methods typically operate in a myopic fashion t…

cs.CL2021

Knowledge Distillation with Noisy Labels for Natural Language Understanding

Shivendra Bhardwaj, Abbas Ghaddar, Ahmad Rashid +5

Knowledge Distillation (KD) is extensively used to compress and deploy large pre-trained language models on edge devices for real-world applications. However, one neglected area of…

cs.RO2026

IR-SIM: A Lightweight Skill-Native Simulator for Navigation, Learning, and Benchmarking

Ruihua Han, Shuai Wang, Chengyang Li +8

Simulation plays a key role in automated robotics research supported by large language models (LLMs). However, existing simulators often require custom code or complex interfaces,…

cs.CV2018

Multispectral Pedestrian Detection via Simultaneous Detection and Segmentation

Chengyang Li, Dan Song, Ruofeng Tong +1

Multispectral pedestrian detection has attracted increasing attention from the research community due to its crucial competence for many around-the-clock applications (e.g., video…

cs.RO2023

An Efficient Spatial-Temporal Trajectory Planner for Autonomous Vehicles in Unstructured Environments

Zhichao Han, Yuwei Wu, Tong Li +8

As a core part of autonomous driving systems, motion planning has received extensive attention from academia and industry. However, real-time trajectory planning capable of spatial…

cs.CV2026

GenDet: Painting Colored Bounding Boxes on Images via Diffusion Model for Object Detection

Chen Min, Chengyang Li, Fanjie Kong +3

This paper presents GenDet, a novel framework that redefines object detection as an image generation task. In contrast to traditional approaches, GenDet adopts a pioneering approac…

cs.CV2026

Point Cloud Registration for Fusion between SPECT MPI and CTA Images

Ni Yao, Xiangyu Liu, Shaojie Tang +9

Clinical fusion of Single Photon Emission Computed Tomography Myocardial Perfusion Imaging (SPECT MPI) and Computed Tomography Angiography (CTA) remains limited by cross-modality m…

cs.CV2026

Remote Sensing Image Dehazing: A Systematic Review of Progress, Challenges, and Prospects

Heng Zhou, Xiaoxiong Liu, Zhenxi Zhang +6

Remote sensing images (RSIs) are frequently degraded by haze, fog, and thin clouds, which obscure surface reflectance and hinder downstream applications. This study presents the fi…

cs.CV2026

FaceRefiner: High-Fidelity Facial Texture Refinement with Differentiable Rendering-based Style Transfer

Chengyang Li, Baoping Cheng, Yao Cheng +5

Recent facial texture generation methods prefer to use deep networks to synthesize image content and then fill in the UV map, thus generating a compelling full texture from a singl…

cs.CV2018

Illumination-aware Faster R-CNN for Robust Multispectral Pedestrian Detection

Chengyang Li, Dan Song, Ruofeng Tong +1

Multispectral images of color-thermal pairs have shown more effective than a single color channel for pedestrian detection, especially under challenging illumination conditions. Ho…

cs.RO2026

DPNet: Doppler LiDAR Motion Planning for Highly-Dynamic Environments

Wei Zuo, Zeyi Ren, Chengyang Li +7

Existing motion planning methods often struggle with rapid-motion obstacles due to an insufficient understanding of environmental changes. To address this, we propose integrating m…

cs.CV2023

EMEF: Ensemble Multi-Exposure Image Fusion

Renshuai Liu, Chengyang Li, Haitao Cao +3

Although remarkable progress has been made in recent years, current multi-exposure image fusion (MEF) research is still bounded by the lack of real ground truth, objective evaluati…

cs.RO2026

Memory-Native Non-Terrestrial Networks for Embodied Intelligence

Chengyang Li, Yikun Wang, Jiahui He +6

Non-terrestrial networks (NTN) provide ubiquitous connectivity for embodied intelligence (EI), enabling robots in wilderness to leverage cloud resources or report critical informat…

cs.CV2025

Myocardial Region-guided Feature Aggregation Net for Automatic Coronary artery Segmentation and Stenosis Assessment using Coronary Computed Tomography Angiography

Ni Yao, Xiangyu Liu, Danyang Sun +7

Coronary artery disease (CAD) remains a leading cause of mortality worldwide, requiring accurate segmentation and stenosis detection using Coronary Computed Tomography angiography…

cs.CV2024

Multi-scale Frequency Enhancement Network for Blind Image Deblurring

Yawen Xiang, Heng Zhou, Chengyang Li +2

Image deblurring is an essential image preprocessing technique, aiming to recover clear and detailed images form blurry ones. However, existing algorithms often fail to effectively…

cs.RO2022

Federated Deep Learning Meets Autonomous Vehicle Perception: Design and Verification

Shuai Wang, Chengyang Li, Derrick Wing Kwan Ng +4

Realizing human-like perception is a challenge in open driving scenarios due to corner cases and visual occlusions. To gather knowledge of rare and occluded instances, federated le…

cs.CV2022

PixelGame: Infrared small target segmentation as a Nash equilibrium

Heng Zhou, Chunna Tian, Zhenxi Zhang +3

A key challenge of infrared small target segmentation (ISTS) is to balance false negative pixels (FNs) and false positive pixels (FPs). Traditional methods combine FNs and FPs into…

cs.RO2026

NeuPAN: Direct Point Robot Navigation with End-to-End Model-based Learning

Ruihua Han, Shuai Wang, Shuaijun Wang +8

Navigating a nonholonomic robot in a cluttered, unknown environment requires accurate perception and precise motion control for real-time collision avoidance. This paper presents N…

cs.CR2022

IOLLVM: enhance version of OLLVM

Chengyang Li, Tianbo Huang, Xiarun Chen +2

Code obfuscation increases the difficulty of understanding programs, improves software security, and, in particular, OLLVM offers the possibility of cross-platform code obfuscation…

cs.RO2026

Semantic Anchoring for Robotic Action Representations

Yuan Xu, Youheng Shi, Chengyang Li +2

The paper studies how fine‑tuning vision‑language‑action models for robots can degrade the semantic structure of their action representations, and proposes a plug‑and‑play anchorin…

#action representation#vision-language models#semantic anchoring#robot learning
cs.CV2024

SMGDiff: Soccer Motion Generation using diffusion probabilistic models

Hongdi Yang, Chengyang Li, Zhenxuan Wu +5

Soccer is a globally renowned sport with significant applications in video games and VR/AR. However, generating realistic soccer motions remains challenging due to the intricate in…

cs.CV2022

Enhancing and Dissecting Crowd Counting By Synthetic Data

Yi Hou, Chengyang Li, Yuheng Lu +4

In this article, we propose a simulated crowd counting dataset CrowdX, which has a large scale, accurate labeling, parameterized realization, and high fidelity. The experimental re…

cs.CV2024

OmniColor: A Global Camera Pose Optimization Approach of LiDAR-360Camera Fusion for Colorizing Point Clouds

Bonan Liu, Guoyang Zhao, Jianhao Jiao +6

A Colored point cloud, as a simple and efficient 3D representation, has many advantages in various fields, including robotic navigation and scene reconstruction. This representatio…

cs.RO2026

Memory Centric Power Allocation for Multi-Agent Embodied Question Answering

Chengyang Li, Shuai Wang, Kejiang Ye +5

This paper considers multi-agent embodied question answering (MA-EQA), which aims to query robot teams on what they have seen over a long horizon. In contrast to existing edge reso…

cs.RO2026

GazeVLA: Learning Human Intention for Robotic Manipulation

Chengyang Li, Kaiyi Xiong, Yuan Xu +3

Embodied foundation models have achieved significant breakthroughs in robotic manipulation, yet they still depend heavily on large-scale robot demonstrations. Although recent works…

cs.CV2023

You Do Not Need Additional Priors in Camouflage Object Detection

Yuchen Dong, Heng Zhou, Chengyang Li +3

Camouflage object detection (COD) poses a significant challenge due to the high resemblance between camouflaged objects and their surroundings. Although current deep learning metho…

cs.CV2024

Learn2Talk: 3D Talking Face Learns from 2D Talking Face

Yixiang Zhuang, Baoping Cheng, Yao Cheng +6

Speech-driven facial animation methods usually contain two main classes, 3D and 2D talking face, both of which attract considerable research attention in recent years. However, to…

cs.CV2024

Deep learning in motion deblurring: current status, benchmarks and future prospects

Yawen Xiang, Heng Zhou, Chengyang Li +3

Motion deblurring is one of the fundamental problems of computer vision and has received continuous attention. The variability in blur, both within and across images, imposes limit…

cs.CV2022

Position-Aware Relation Learning for RGB-Thermal Salient Object Detection

Heng Zhou, Chunna Tian, Zhenxi Zhang +4

RGB-Thermal salient object detection (SOD) combines two spectra to segment visually conspicuous regions in images. Most existing methods use boundary maps to learn the sharp bounda…

cs.CV2026

GaussianSwap: Animatable Video Face Swapping with 3D Gaussian Splatting

Xuan Cheng, Jiahao Rao, Chengyang Li +3

We introduce GaussianSwap, a novel video face swapping framework that constructs a 3D Gaussian Splatting based face avatar from a target video while transferring identity from a so…

cs.RO2025

Federated Split Learning for Resource-Constrained Robots in Industrial IoT: Framework Comparison, Optimization Strategies, and Future Directions

Wanli Ni, Hui Tian, Shuai Wang +3

Federated split learning (FedSL) has emerged as a promising paradigm for enabling collaborative intelligence in industrial Internet of Things (IoT) systems, particularly in smart f…

cs.CV2022

BBA-net: A bi-branch attention network for crowd counting

Yi Hou, Chengyang Li, Fan Yang +5

In the field of crowd counting, the current mainstream CNN-based regression methods simply extract the density information of pedestrians without finding the position of each perso…

cs.RO2022

Adaptive Environment Modeling Based Reinforcement Learning for Collision Avoidance in Complex Scenes

Shuaijun Wang, Rui Gao, Ruihua Han +3

The major challenges of collision avoidance for robot navigation in crowded scenes lie in accurate environment modeling, fast perceptions, and trustworthy motion planning policies.…

cs.RO2026

Agentic Self-Evolutionary Replanning for Embodied Navigation

Guoliang Li, Ruihua Han, Chengyang Li +5

Failure is inevitable for embodied navigation in complex environments. To enhance the resilience, replanning (RP) is a viable option, where the robot is allowed to fail, but is cap…

cs.CV2025

C-DGPA: Class-Centric Dual-Alignment Generative Prompt Adaptation

Chao Li, Dasha Hu, Chengyang Li +2

Unsupervised Domain Adaptation transfers knowledge from a labeled source domain to an unlabeled target domain. Directly deploying Vision-Language Models (VLMs) with prompt tuning i…

cs.RO2023

Decentralized Planning for Car-Like Robotic Swarm in Cluttered Environments

Changjia Ma, Zhichao Han, Tingrui Zhang +5

Robot swarm is a hot spot in robotic research community. In this paper, we propose a decentralized framework for car-like robotic swarm which is capable of real-time planning in cl…