papers

Publications (55)

cs.CV2026

Noise-Robust Box-Supervised Infrared Small Target Detection via Physics-Inspired Soft Label Optimization

Xizhe Zhang, Fan Shi, Mianzhao Wang +3

Infrared small target detection (IRSTD) commonly relies on pixel-level mask supervision. Such annotations, however, are costly and inherently uncertain because infrared targets hav…

cs.RO2020

Versatile Multilinked Aerial Robot with Tilting Propellers: Design, Modeling, Control and State Estimation for Autonomous Flight and Manipulation

Moju Zhao, Tomoki Anzai, Fan Shi +6

Multilinked aerial robot is one of the state-of-the-art works in aerial robotics, which demonstrates the deformability benefiting both maneuvering and manipulation. However, the pe…

cs.RO2020

Circus ANYmal: A Quadruped Learning Dexterous Manipulation with Its Limbs

Fan Shi, Timon Homberger, Joonho Lee +6

Quadrupedal robots are skillful at locomotion tasks while lacking manipulation skills, not to mention dexterous manipulation abilities. Inspired by the animal behavior and the dual…

cs.CV2025

A Causal Adjustment Module for Debiasing Scene Graph Generation

Li Liu, Shuzhou Sun, Shuaifeng Zhi +4

While recent debiasing methods for Scene Graph Generation (SGG) have shown impressive performance, these efforts often attribute model bias solely to the long-tail distribution of…

cs.CV2026

AvatarForcing: One-Step Streaming Talking Avatars via Local-Future Sliding-Window Denoising

Liyuan Cui, Wentao Hu, Wenyuan Zhang +3

Real-time talking avatar generation requires low latency and minute-level temporal stability. Autoregressive (AR) forcing enables streaming inference but suffers from exposure bias…

cs.RO2024

Residual Policy Learning for Perceptive Quadruped Control Using Differentiable Simulation

Jing Yuan Luo, Yunlong Song, Victor Klemm +3

First-order Policy Gradient (FoPG) algorithms such as Backpropagation through Time and Analytical Policy Gradients leverage local simulation physics to accelerate policy search, si…

cs.LG2025

Dynamic Graph-Like Learning with Contrastive Clustering on Temporally-Factored Ship Motion Data for Imbalanced Sea State Estimation in Autonomous Vessel

Kexin Wang, Mengna Liu, Xu Cheng +3

Accurate sea state estimation is crucial for the real-time control and future state prediction of autonomous vessels. However, traditional methods struggle with challenges such as…

cond-mat.mes-hall2020

Terahertz response of gadolinium gallium garnet (GGG) and gadolinium scandium gallium garnet (SGGG)

Mohsen Sabbaghi, George W. Hanson, Michael Weinert +2

We report the magneto-optical response of Gadolinium Gallium Garnet (GGG) and Gadolinium Scandium Gallium Garnet (SGGG) at frequencies ranging from to $1 \, \…

cs.RO2026

Safe Navigation in Unknown and Cluttered Environments via Direction-Aware Convex Free-Region Generation

Zhicheng Song, Yongjian Li, Kai Chen +3

Convex free regions provide a structured and optimization-friendly representation of collision-free space for robot navigation in unknown and cluttered environments. However, exist…

cs.RO2026

FlowHOI: Flow-based Semantics-Grounded Generation of Hand-Object Interactions for Dexterous Robot Manipulation

Huajian Zeng, Lingyun Chen, Jiaqi Yang +4

Recent vision-language-action (VLA) models can generate plausible end-effector motions, yet they often fail in long-horizon, contact-rich tasks because the underlying hand-object i…

cs.RO2024

HumanMimic: Learning Natural Locomotion and Transitions for Humanoid Robot via Wasserstein Adversarial Imitation

Annan Tang, Takuma Hiraoka, Naoki Hiraoka +5

Transferring human motion skills to humanoid robots remains a significant challenge. In this study, we introduce a Wasserstein adversarial imitation learning system, allowing human…

cs.CV2025

SCSegamba: Lightweight Structure-Aware Vision Mamba for Crack Segmentation in Structures

Hui Liu, Chen Jia, Fan Shi +2

Pixel-level segmentation of structural cracks across various scenarios remains a considerable challenge. Current methods encounter challenges in effectively modeling crack morpholo…

cs.CV2023

RefineNet: Enhancing Text-to-Image Conversion with High-Resolution and Detail Accuracy through Hierarchical Transformers and Progressive Refinement

Fan Shi

In this research, we introduce RefineNet, a novel architecture designed to address resolution limitations in text-to-image conversion systems. We explore the challenges of generati…

cs.CV2023

High-order Spatial Interactions Enhanced Lightweight Model for Optical Remote Sensing Image-based Small Ship Detection

Yifan Yin, Xu Cheng, Fan Shi +3

Accurate and reliable optical remote sensing image-based small-ship detection is crucial for maritime surveillance systems, but existing methods often struggle with balancing detec…

cs.AI2023

Raven's Progressive Matrices Completion with Latent Gaussian Process Priors

Fan Shi, Bin Li, Xiangyang Xue

Abstract reasoning ability is fundamental to human intelligence. It enables humans to uncover relations among abstract concepts and further deduce implicit rules from the relations…

cs.RO2024

Rethinking Robustness Assessment: Adversarial Attacks on Learning-based Quadrupedal Locomotion Controllers

Fan Shi, Chong Zhang, Takahiro Miki +3

Legged locomotion has recently achieved remarkable success with the progress of machine learning techniques, especially deep reinforcement learning (RL). Controllers employing neur…

cs.RO2026

Fast and Safe Trajectory Optimization for Mobile Manipulators With Neural Configuration Space Distance Field

Yulin Li, Zhiyuan Song, Yiming Li +8

Mobile manipulators promise agile, long-horizon behavior by coordinating base and arm motion, yet whole-body trajectory optimization in cluttered, confined spaces remains difficult…

cs.RO2026

Soft Responsive Materials Enhance Humanoid Safety

Chunzheng Wang, Yiyuan Zhang, Annan Tang +12

Humanoid robots are envisioned as general-purpose platforms in human-centered environments, yet their deployment is limited by vulnerability to falls and the risks posed by rigid m…

cs.GR2026

Fast and Reliable Gradients for Deformables Across Frictional Contact Regimes

Ziqiu Zeng, Gang Yang, Zhenhao Huang +5

Differentiable simulation establishes the mathematical foundation for solving challenging inverse problems in computer graphics and robotics, such as physical system identification…

cs.CV2026

An Angular-Temporal Interaction Network for Light Field Object Tracking in Low-Light Scenes

Mianzhao Wang, Fan Shi, Xu Cheng +2

High-quality 4D light field representation with efficient angular feature modeling is crucial for scene perception, as it can provide discriminative spatial-angular cues to identif…

cs.CV2025

CoLa: Chinese Character Decomposition with Compositional Latent Components

Fan Shi, Haiyang Yu, Bin Li +1

Humans can decompose Chinese characters into compositional components and recombine them to recognize unseen characters. This reflects two cognitive principles: Compositionality, t…

cs.RO2026

RARM: Confidence-Gated Progress Reward Modeling for RL in Manipulation

Pengzhi Yang, Xinyu Wang, Pengyu Jing +7

Reinforcement learning for robot manipulation is often bottlenecked by reward design, especially in long-horizon tasks: sparse success rewards provide weak supervision, while hand-…

cs.CV2026

LFM: Leveraging Foundation Models for Source-Free Universal Domain Adaptation

Jing Li, Pan Liu, Meng Zhao +7

Source-free universal domain adaptation (SF-UniDA) adapts a pre-trained source model to an unlabeled target domain under both covariate and label shifts, without access to source d…

cs.RO2025

Learning to Walk in Costume: Adversarial Motion Priors for Aesthetically Constrained Humanoids

Arturo Flores Alvarez, Fatemeh Zargarbashi, Havel Liu +7

We present a Reinforcement Learning (RL)-based locomotion system for Cosmo, a custom-built humanoid robot designed for entertainment applications. Unlike traditional humanoids, ent…

cs.LG2024

An End-to-End Model for Time Series Classification In the Presence of Missing Values

Pengshuai Yao, Mengna Liu, Xu Cheng +4

Time series classification with missing data is a prevalent issue in time series analysis, as temporal data often contain missing values in practical applications. The traditional…

cs.SD2023

Revolutionizing Personalized Voice Synthesis: The Journey towards Emotional and Individual Authenticity with DIVSE (Dynamic Individual Voice Synthesis Engine)

Fan Shi

This comprehensive paper delves into the forefront of personalized voice synthesis within artificial intelligence (AI), spotlighting the Dynamic Individual Voice Synthesis Engine (…

cs.CL2025

Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning

Haiyang Yu, Yuchuan Wu, Fan Shi +18

Chinese ancient documents, invaluable carriers of millennia of Chinese history and culture, hold rich knowledge across diverse fields but face challenges in digitization and unders…

cs.RO2026

FLASH: Fast Learning via GPU-Accelerated Simulation for High-Fidelity Deformable Manipulation in Minutes

Siyuan Luo, Bingyang Zhou, Chong Zhang +9

Simulation frameworks such as Isaac Sim have enabled scalable robot learning for locomotion and rigid-body manipulation; however, contact-rich simulation remains a major bottleneck…

cs.LG2025

Prototype-based Heterogeneous Federated Learning for Blade Icing Detection in Wind Turbines with Class Imbalanced Data

Lele Qi, Mengna Liu, Xu Cheng +3

Wind farms, typically in high-latitude regions, face a high risk of blade icing. Traditional centralized training methods raise serious privacy concerns. To enhance data privacy in…

cs.GR2025

Fast But Accurate: A Real-Time Hyperelastic Simulator with Robust Frictional Contact

Ziqiu Zeng, Siyuan Luo, Fan Shi +1

We present a GPU-friendly framework for real-time implicit simulation of elastic material in the presence of frictional contacts. The integration of hyperelasticity, non-interpenet…

cs.AI2024

Towards Generative Abstract Reasoning: Completing Raven's Progressive Matrix via Rule Abstraction and Selection

Fan Shi, Bin Li, Xiangyang Xue

Endowing machines with abstract reasoning ability has been a long-term research topic in artificial intelligence. Raven's Progressive Matrix (RPM) is widely used to probe abstract…

cs.CV2026

Staircase Cascaded Fusion of Lightweight Local Pattern Recognition and Long-Range Dependencies for Structural Crack Segmentation

Hui Liu, Chen Jia, Fan Shi +4

Accurately segmenting structural cracks at the pixel level remains a major hurdle, as existing methods fail to integrate local textures with pixel dependencies, often leading to fr…

cs.CV2025

Beyond Task-Specific Reasoning: A Unified Conditional Generative Framework for Abstract Visual Reasoning

Fan Shi, Bin Li, Xiangyang Xue

Abstract visual reasoning (AVR) enables humans to quickly discover and generalize abstract rules to new scenarios. Designing intelligent systems with human-like AVR abilities has b…

cs.CV2025

GGTalker: Talking Head Systhesis with Generalizable Gaussian Priors and Identity-Specific Adaptation

Wentao Hu, Shunkai Li, Ziqiao Peng +6

Creating high-quality, generalizable speech-driven 3D talking heads remains a persistent challenge. Previous methods achieve satisfactory results for fixed viewpoints and small-sca…

cs.RO2026

Enabling Robust Cloth Manipulation via Inference-Time Simulator-in-the-Loop Refinement

Xin Liu, Yulin Li, Ziming Li +7

Simulator-in-the-loop optimization offers a promising inference-time mechanism for robot manipulation. It uses a physical simulator as a backend rollout engine to evaluate candidat…

cs.CV2023

Compositional Law Parsing with Latent Random Functions

Fan Shi, Bin Li, Xiangyang Xue

Human cognition has compositionality. We understand a scene by decomposing the scene into different concepts (e.g., shape and position of an object) and learning the respective law…

cs.RO2025

GO-Flock: Goal-Oriented Flocking in 3D Unknown Environments with Depth Maps

Yan Rui Tan, Wenqi Liu, Wai Lun Leong +4

Artificial Potential Field (APF) methods are widely used for reactive flocking control, but they often suffer from challenges such as deadlocks and local minima, especially in the…

cs.CV2025

LIDAR: Lightweight Adaptive Cue-Aware Fusion Vision Mamba for Multimodal Segmentation of Structural Cracks

Hui Liu, Chen Jia, Fan Shi +4

Achieving pixel-level segmentation with low computational cost using multimodal data remains a key challenge in crack segmentation tasks. Existing methods lack the capability for a…

cs.CR2026

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing

Jiaren Peng, Zeqin Li, Chang You +17

The rapid advancement of Large Language Models (LLMs) has created new opportunities for Automated Penetration Testing (AutoPT), spawning numerous frameworks aimed at achieving end-…

cs.CR2025

Beyond Jailbreak: Unveiling Risks in LLM Applications Arising from Blurred Capability Boundaries

Yunyi Zhang, Shibo Cui, Baojun Liu +4

LLM applications (i.e., LLM apps) leverage the powerful capabilities of LLMs to provide users with customized services, revolutionizing traditional application development. While t…

cs.AI2026

Using large language models for embodied planning introduces systematic safety risks

Tao Zhang, Kaixian Qu, Zhibin Li +4

Large language models are increasingly used as planners for robotic systems, yet how safely they plan remains an open question. To evaluate safe planning systematically, we introdu…

cs.CV2026

Compass: Degradation-Simulated Reciprocal Learning with Lightweight Needle RWKV for Multimodal Crack Segmentation under Missing Modalities

Hui Liu, Chen Jia, Fan Shi +3

In multimodal crack segmentation for industrial facilities, the key challenge is preventing missing modalities from degrading pixel-level performance while maintaining low computat…

cs.RO2025

REBot: Reflexive Evasion Robot for Instantaneous Dynamic Obstacle Avoidance

Zihao Xu, Ce Hao, Chunzheng Wang +3

Dynamic obstacle avoidance (DOA) is critical for quadrupedal robots operating in environments with moving obstacles or humans. Existing approaches typically rely on navigation-base…

cs.AI2023

Abstracting Concept-Changing Rules for Solving Raven's Progressive Matrix Problems

Fan Shi, Bin Li, Xiangyang Xue

The abstract visual reasoning ability in human intelligence benefits discovering underlying rules in the novel environment. Raven's Progressive Matrix (RPM) is a classic test to re…

cs.CV2021

ViDA-MAN: Visual Dialog with Digital Humans

Tong Shen, Jiawei Zuo, Fan Shi +7

We demonstrate ViDA-MAN, a digital-human agent for multi-modal interaction, which offers realtime audio-visual responses to instant speech inquiries. Compared to traditional text o…

cs.RO2026

Few-Shot Neural Differentiable Simulator: Real-to-Sim Rigid-Contact Modeling

Zhenhao Huang, Siyuan Luo, Bingyang Zhou +3

Accurate physics simulation is essential for robotic learning and control, yet analytical simulators often fail to capture complex contact dynamics, while learning-based simulators…

cs.RO2026

Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents

Guanxiong Chen, Qianjun Xia, Jiawei Peng +21

Real-to-sim conversion for robotic interaction with objects remains labor-intensive because it requires more than visual reconstruction: a streamlined real2sim process must recover…

cs.RO2026

SAC-Loco: Safe and Adjustable Compliant Quadrupedal Locomotion

Aoqian Zhang, Zixuan Zhuang, Chunzheng Wang +3

Quadruped robots are designed to achieve agile and robust locomotion by drawing inspiration from legged animals. However, most existing control methods for quadruped robots lack a…

cs.RO2025

Learning Quiet Walking for a Small Home Robot

Ryo Watanabe, Takahiro Miki, Fan Shi +5

As home robotics gains traction, robots are increasingly integrated into households, offering companionship and assistance. Quadruped robots, particularly those resembling dogs, ha…

cs.CV2026

SCRWKV: Ultra-Compact Structure-Calibrated Vision-RWKV for Topological Crack Segmentation

Hanxu Zhang, Chen Jia, Hui Liu +3

Achieving pixel-level accurate segmentation of structural cracks across diverse scenarios remains a formidable challenge. Existing methods face significant bottlenecks in balancing…

cs.CV2026

Kling-MotionControl Technical Report

Kling Team, Jialu Chen, Yikang Ding +21

Character animation aims to generate lifelike videos by transferring motion dynamics from a driving video to a reference image. Recent strides in generative models have paved the w…

cs.RO2023

Sensing and Navigation of Aerial Robot for Measuring Tree Location and Size in Forest Environment

Tomoki Anzai, Moju Zhao, Fan Shi +2

This paper shows the achievement of a sensing and navigation system of aerial robot for measuring location and size of trees in a forest environment autonomously. Although forestry…

cs.GR2026

DiffPhD: A Unified Differentiable Solver for Projective Heterogeneous Materials in Elastodynamics with Contact-Rich GPU-Acceleration

Shih-Yu Lai, Sung-Han Tien, Jui-I Huang +9

Differentiable simulation of soft bodies is a foundation for system identification, trajectory optimization, and Real2Sim transfer. Yet, existing methods such as the differentiable…

physics.optics2016

Generation of femtosecond optical vortex beams in all-fiber mode-locked fiber laser using mode selective coupler

Teng Wang, Feng Wang, Fan Shi +4

We experimentally demonstrated a high-order optical vortex pulsed laser based on a mode selective all-fiber fused coupler composed of a single-mode fiber (SMF) and a few-mode fiber…

cs.RO2026

Real-IKEA: Physical Fidelity is the Prerequisite for Robust Manipulation

Kunqi Xu, Zhenhao Huang, Siyuan Luo +2

Robotic manipulation robustness often founders on the physics gap between simplified simulations and the resistance-laden real world. In this work, we emphasize that physical reali…