Publications (55)
Noise-Robust Box-Supervised Infrared Small Target Detection via Physics-Inspired Soft Label Optimization
Xizhe Zhang, Fan Shi, Mianzhao Wang +3
Infrared small target detection (IRSTD) commonly relies on pixel-level mask supervision. Such annotations, however, are costly and inherently uncertain because infrared targets hav…
Versatile Multilinked Aerial Robot with Tilting Propellers: Design, Modeling, Control and State Estimation for Autonomous Flight and Manipulation
Moju Zhao, Tomoki Anzai, Fan Shi +6
Multilinked aerial robot is one of the state-of-the-art works in aerial robotics, which demonstrates the deformability benefiting both maneuvering and manipulation. However, the pe…
Circus ANYmal: A Quadruped Learning Dexterous Manipulation with Its Limbs
Fan Shi, Timon Homberger, Joonho Lee +6
Quadrupedal robots are skillful at locomotion tasks while lacking manipulation skills, not to mention dexterous manipulation abilities. Inspired by the animal behavior and the dual…
A Causal Adjustment Module for Debiasing Scene Graph Generation
Li Liu, Shuzhou Sun, Shuaifeng Zhi +4
While recent debiasing methods for Scene Graph Generation (SGG) have shown impressive performance, these efforts often attribute model bias solely to the long-tail distribution of…
AvatarForcing: One-Step Streaming Talking Avatars via Local-Future Sliding-Window Denoising
Liyuan Cui, Wentao Hu, Wenyuan Zhang +3
Real-time talking avatar generation requires low latency and minute-level temporal stability. Autoregressive (AR) forcing enables streaming inference but suffers from exposure bias…
Residual Policy Learning for Perceptive Quadruped Control Using Differentiable Simulation
Jing Yuan Luo, Yunlong Song, Victor Klemm +3
First-order Policy Gradient (FoPG) algorithms such as Backpropagation through Time and Analytical Policy Gradients leverage local simulation physics to accelerate policy search, si…
Dynamic Graph-Like Learning with Contrastive Clustering on Temporally-Factored Ship Motion Data for Imbalanced Sea State Estimation in Autonomous Vessel
Kexin Wang, Mengna Liu, Xu Cheng +3
Accurate sea state estimation is crucial for the real-time control and future state prediction of autonomous vessels. However, traditional methods struggle with challenges such as…
Terahertz response of gadolinium gallium garnet (GGG) and gadolinium scandium gallium garnet (SGGG)
Mohsen Sabbaghi, George W. Hanson, Michael Weinert +2
We report the magneto-optical response of Gadolinium Gallium Garnet (GGG) and Gadolinium Scandium Gallium Garnet (SGGG) at frequencies ranging from to $1 \, \…
Safe Navigation in Unknown and Cluttered Environments via Direction-Aware Convex Free-Region Generation
Zhicheng Song, Yongjian Li, Kai Chen +3
Convex free regions provide a structured and optimization-friendly representation of collision-free space for robot navigation in unknown and cluttered environments. However, exist…
FlowHOI: Flow-based Semantics-Grounded Generation of Hand-Object Interactions for Dexterous Robot Manipulation
Huajian Zeng, Lingyun Chen, Jiaqi Yang +4
Recent vision-language-action (VLA) models can generate plausible end-effector motions, yet they often fail in long-horizon, contact-rich tasks because the underlying hand-object i…
HumanMimic: Learning Natural Locomotion and Transitions for Humanoid Robot via Wasserstein Adversarial Imitation
Annan Tang, Takuma Hiraoka, Naoki Hiraoka +5
Transferring human motion skills to humanoid robots remains a significant challenge. In this study, we introduce a Wasserstein adversarial imitation learning system, allowing human…
SCSegamba: Lightweight Structure-Aware Vision Mamba for Crack Segmentation in Structures
Hui Liu, Chen Jia, Fan Shi +2
Pixel-level segmentation of structural cracks across various scenarios remains a considerable challenge. Current methods encounter challenges in effectively modeling crack morpholo…
RefineNet: Enhancing Text-to-Image Conversion with High-Resolution and Detail Accuracy through Hierarchical Transformers and Progressive Refinement
Fan Shi
In this research, we introduce RefineNet, a novel architecture designed to address resolution limitations in text-to-image conversion systems. We explore the challenges of generati…
High-order Spatial Interactions Enhanced Lightweight Model for Optical Remote Sensing Image-based Small Ship Detection
Yifan Yin, Xu Cheng, Fan Shi +3
Accurate and reliable optical remote sensing image-based small-ship detection is crucial for maritime surveillance systems, but existing methods often struggle with balancing detec…
Raven's Progressive Matrices Completion with Latent Gaussian Process Priors
Fan Shi, Bin Li, Xiangyang Xue
Abstract reasoning ability is fundamental to human intelligence. It enables humans to uncover relations among abstract concepts and further deduce implicit rules from the relations…
Rethinking Robustness Assessment: Adversarial Attacks on Learning-based Quadrupedal Locomotion Controllers
Fan Shi, Chong Zhang, Takahiro Miki +3
Legged locomotion has recently achieved remarkable success with the progress of machine learning techniques, especially deep reinforcement learning (RL). Controllers employing neur…
Fast and Safe Trajectory Optimization for Mobile Manipulators With Neural Configuration Space Distance Field
Yulin Li, Zhiyuan Song, Yiming Li +8
Mobile manipulators promise agile, long-horizon behavior by coordinating base and arm motion, yet whole-body trajectory optimization in cluttered, confined spaces remains difficult…
Soft Responsive Materials Enhance Humanoid Safety
Chunzheng Wang, Yiyuan Zhang, Annan Tang +12
Humanoid robots are envisioned as general-purpose platforms in human-centered environments, yet their deployment is limited by vulnerability to falls and the risks posed by rigid m…
Fast and Reliable Gradients for Deformables Across Frictional Contact Regimes
Ziqiu Zeng, Gang Yang, Zhenhao Huang +5
Differentiable simulation establishes the mathematical foundation for solving challenging inverse problems in computer graphics and robotics, such as physical system identification…
An Angular-Temporal Interaction Network for Light Field Object Tracking in Low-Light Scenes
Mianzhao Wang, Fan Shi, Xu Cheng +2
High-quality 4D light field representation with efficient angular feature modeling is crucial for scene perception, as it can provide discriminative spatial-angular cues to identif…
CoLa: Chinese Character Decomposition with Compositional Latent Components
Fan Shi, Haiyang Yu, Bin Li +1
Humans can decompose Chinese characters into compositional components and recombine them to recognize unseen characters. This reflects two cognitive principles: Compositionality, t…
RARM: Confidence-Gated Progress Reward Modeling for RL in Manipulation
Pengzhi Yang, Xinyu Wang, Pengyu Jing +7
Reinforcement learning for robot manipulation is often bottlenecked by reward design, especially in long-horizon tasks: sparse success rewards provide weak supervision, while hand-…
LFM: Leveraging Foundation Models for Source-Free Universal Domain Adaptation
Jing Li, Pan Liu, Meng Zhao +7
Source-free universal domain adaptation (SF-UniDA) adapts a pre-trained source model to an unlabeled target domain under both covariate and label shifts, without access to source d…
Learning to Walk in Costume: Adversarial Motion Priors for Aesthetically Constrained Humanoids
Arturo Flores Alvarez, Fatemeh Zargarbashi, Havel Liu +7
We present a Reinforcement Learning (RL)-based locomotion system for Cosmo, a custom-built humanoid robot designed for entertainment applications. Unlike traditional humanoids, ent…
An End-to-End Model for Time Series Classification In the Presence of Missing Values
Pengshuai Yao, Mengna Liu, Xu Cheng +4
Time series classification with missing data is a prevalent issue in time series analysis, as temporal data often contain missing values in practical applications. The traditional…
Revolutionizing Personalized Voice Synthesis: The Journey towards Emotional and Individual Authenticity with DIVSE (Dynamic Individual Voice Synthesis Engine)
Fan Shi
This comprehensive paper delves into the forefront of personalized voice synthesis within artificial intelligence (AI), spotlighting the Dynamic Individual Voice Synthesis Engine (…
Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning
Haiyang Yu, Yuchuan Wu, Fan Shi +18
Chinese ancient documents, invaluable carriers of millennia of Chinese history and culture, hold rich knowledge across diverse fields but face challenges in digitization and unders…
FLASH: Fast Learning via GPU-Accelerated Simulation for High-Fidelity Deformable Manipulation in Minutes
Siyuan Luo, Bingyang Zhou, Chong Zhang +9
Simulation frameworks such as Isaac Sim have enabled scalable robot learning for locomotion and rigid-body manipulation; however, contact-rich simulation remains a major bottleneck…
Prototype-based Heterogeneous Federated Learning for Blade Icing Detection in Wind Turbines with Class Imbalanced Data
Lele Qi, Mengna Liu, Xu Cheng +3
Wind farms, typically in high-latitude regions, face a high risk of blade icing. Traditional centralized training methods raise serious privacy concerns. To enhance data privacy in…
Fast But Accurate: A Real-Time Hyperelastic Simulator with Robust Frictional Contact
Ziqiu Zeng, Siyuan Luo, Fan Shi +1
We present a GPU-friendly framework for real-time implicit simulation of elastic material in the presence of frictional contacts. The integration of hyperelasticity, non-interpenet…
Towards Generative Abstract Reasoning: Completing Raven's Progressive Matrix via Rule Abstraction and Selection
Fan Shi, Bin Li, Xiangyang Xue
Endowing machines with abstract reasoning ability has been a long-term research topic in artificial intelligence. Raven's Progressive Matrix (RPM) is widely used to probe abstract…
Staircase Cascaded Fusion of Lightweight Local Pattern Recognition and Long-Range Dependencies for Structural Crack Segmentation
Hui Liu, Chen Jia, Fan Shi +4
Accurately segmenting structural cracks at the pixel level remains a major hurdle, as existing methods fail to integrate local textures with pixel dependencies, often leading to fr…
Beyond Task-Specific Reasoning: A Unified Conditional Generative Framework for Abstract Visual Reasoning
Fan Shi, Bin Li, Xiangyang Xue
Abstract visual reasoning (AVR) enables humans to quickly discover and generalize abstract rules to new scenarios. Designing intelligent systems with human-like AVR abilities has b…
GGTalker: Talking Head Systhesis with Generalizable Gaussian Priors and Identity-Specific Adaptation
Wentao Hu, Shunkai Li, Ziqiao Peng +6
Creating high-quality, generalizable speech-driven 3D talking heads remains a persistent challenge. Previous methods achieve satisfactory results for fixed viewpoints and small-sca…
Enabling Robust Cloth Manipulation via Inference-Time Simulator-in-the-Loop Refinement
Xin Liu, Yulin Li, Ziming Li +7
Simulator-in-the-loop optimization offers a promising inference-time mechanism for robot manipulation. It uses a physical simulator as a backend rollout engine to evaluate candidat…
Compositional Law Parsing with Latent Random Functions
Fan Shi, Bin Li, Xiangyang Xue
Human cognition has compositionality. We understand a scene by decomposing the scene into different concepts (e.g., shape and position of an object) and learning the respective law…
GO-Flock: Goal-Oriented Flocking in 3D Unknown Environments with Depth Maps
Yan Rui Tan, Wenqi Liu, Wai Lun Leong +4
Artificial Potential Field (APF) methods are widely used for reactive flocking control, but they often suffer from challenges such as deadlocks and local minima, especially in the…
LIDAR: Lightweight Adaptive Cue-Aware Fusion Vision Mamba for Multimodal Segmentation of Structural Cracks
Hui Liu, Chen Jia, Fan Shi +4
Achieving pixel-level segmentation with low computational cost using multimodal data remains a key challenge in crack segmentation tasks. Existing methods lack the capability for a…
Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing
Jiaren Peng, Zeqin Li, Chang You +17
The rapid advancement of Large Language Models (LLMs) has created new opportunities for Automated Penetration Testing (AutoPT), spawning numerous frameworks aimed at achieving end-…
Beyond Jailbreak: Unveiling Risks in LLM Applications Arising from Blurred Capability Boundaries
Yunyi Zhang, Shibo Cui, Baojun Liu +4
LLM applications (i.e., LLM apps) leverage the powerful capabilities of LLMs to provide users with customized services, revolutionizing traditional application development. While t…
Using large language models for embodied planning introduces systematic safety risks
Tao Zhang, Kaixian Qu, Zhibin Li +4
Large language models are increasingly used as planners for robotic systems, yet how safely they plan remains an open question. To evaluate safe planning systematically, we introdu…
Compass: Degradation-Simulated Reciprocal Learning with Lightweight Needle RWKV for Multimodal Crack Segmentation under Missing Modalities
Hui Liu, Chen Jia, Fan Shi +3
In multimodal crack segmentation for industrial facilities, the key challenge is preventing missing modalities from degrading pixel-level performance while maintaining low computat…
REBot: Reflexive Evasion Robot for Instantaneous Dynamic Obstacle Avoidance
Zihao Xu, Ce Hao, Chunzheng Wang +3
Dynamic obstacle avoidance (DOA) is critical for quadrupedal robots operating in environments with moving obstacles or humans. Existing approaches typically rely on navigation-base…
Abstracting Concept-Changing Rules for Solving Raven's Progressive Matrix Problems
Fan Shi, Bin Li, Xiangyang Xue
The abstract visual reasoning ability in human intelligence benefits discovering underlying rules in the novel environment. Raven's Progressive Matrix (RPM) is a classic test to re…
ViDA-MAN: Visual Dialog with Digital Humans
Tong Shen, Jiawei Zuo, Fan Shi +7
We demonstrate ViDA-MAN, a digital-human agent for multi-modal interaction, which offers realtime audio-visual responses to instant speech inquiries. Compared to traditional text o…
Few-Shot Neural Differentiable Simulator: Real-to-Sim Rigid-Contact Modeling
Zhenhao Huang, Siyuan Luo, Bingyang Zhou +3
Accurate physics simulation is essential for robotic learning and control, yet analytical simulators often fail to capture complex contact dynamics, while learning-based simulators…
Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents
Guanxiong Chen, Qianjun Xia, Jiawei Peng +21
Real-to-sim conversion for robotic interaction with objects remains labor-intensive because it requires more than visual reconstruction: a streamlined real2sim process must recover…
SAC-Loco: Safe and Adjustable Compliant Quadrupedal Locomotion
Aoqian Zhang, Zixuan Zhuang, Chunzheng Wang +3
Quadruped robots are designed to achieve agile and robust locomotion by drawing inspiration from legged animals. However, most existing control methods for quadruped robots lack a…
Learning Quiet Walking for a Small Home Robot
Ryo Watanabe, Takahiro Miki, Fan Shi +5
As home robotics gains traction, robots are increasingly integrated into households, offering companionship and assistance. Quadruped robots, particularly those resembling dogs, ha…
SCRWKV: Ultra-Compact Structure-Calibrated Vision-RWKV for Topological Crack Segmentation
Hanxu Zhang, Chen Jia, Hui Liu +3
Achieving pixel-level accurate segmentation of structural cracks across diverse scenarios remains a formidable challenge. Existing methods face significant bottlenecks in balancing…
Kling-MotionControl Technical Report
Kling Team, Jialu Chen, Yikang Ding +21
Character animation aims to generate lifelike videos by transferring motion dynamics from a driving video to a reference image. Recent strides in generative models have paved the w…
Sensing and Navigation of Aerial Robot for Measuring Tree Location and Size in Forest Environment
Tomoki Anzai, Moju Zhao, Fan Shi +2
This paper shows the achievement of a sensing and navigation system of aerial robot for measuring location and size of trees in a forest environment autonomously. Although forestry…
DiffPhD: A Unified Differentiable Solver for Projective Heterogeneous Materials in Elastodynamics with Contact-Rich GPU-Acceleration
Shih-Yu Lai, Sung-Han Tien, Jui-I Huang +9
Differentiable simulation of soft bodies is a foundation for system identification, trajectory optimization, and Real2Sim transfer. Yet, existing methods such as the differentiable…
Generation of femtosecond optical vortex beams in all-fiber mode-locked fiber laser using mode selective coupler
Teng Wang, Feng Wang, Fan Shi +4
We experimentally demonstrated a high-order optical vortex pulsed laser based on a mode selective all-fiber fused coupler composed of a single-mode fiber (SMF) and a few-mode fiber…
Real-IKEA: Physical Fidelity is the Prerequisite for Robust Manipulation
Kunqi Xu, Zhenhao Huang, Siyuan Luo +2
Robotic manipulation robustness often founders on the physics gap between simplified simulations and the resistance-laden real world. In this work, we emphasize that physical reali…