papers

Publications (25)

cs.CV2023

Improving Neural Indoor Surface Reconstruction with Mask-Guided Adaptive Consistency Constraints

Xinyi Yu, Liqin Lu, Jintao Rong +2

3D scene reconstruction from 2D images has been a long-standing task. Instead of estimating per-frame depth maps and fusing them in 3D, recent research leverages the neural implici…

cs.CV2026

When Distillation Breaks Motion Control: Restoring Generative Trajectories for Fast Video Generators

Jintao Rong, Xin Xie, Xinyi Yu +4

Training-free motion customization imposes motion patterns from reference videos onto video generators through test-time computation. Most existing methods target full diffusion mo…

cs.CV2024

Retrieval-Enhanced Visual Prompt Learning for Few-shot Classification

Jintao Rong, Hao Chen, Linlin Ou +3

The Contrastive Language-Image Pretraining (CLIP) model has been widely used in various downstream vision tasks. The few-shot learning paradigm has been widely adopted to augment i…

cs.CV2023

ShiftNAS: Improving One-shot NAS via Probability Shift

Mingyang Zhang, Xinyi Yu, Haodong Zhao +1

One-shot Neural architecture search (One-shot NAS) has been proposed as a time-efficient approach to obtain optimal subnet architectures and weights under different complexity case…

eess.SY2025

Explicit Solution of Tunable Input-to-State Safe-Based Controller Under High-Relative-Degree Constraints

Yan Wei, Yu Feng, Linlin Ou +2

This paper investigates the safety analysis and verification of nonlinear systems subject to high-relative-degree constraints and unknown disturbance. The closed-form solution of t…

cs.AI2022

Efficient Re-parameterization Operations Search for Easy-to-Deploy Network Based on Directional Evolutionary Strategy

Xinyi Yu, Xiaowei Wang, Jintao Rong +2

Structural re-parameterization (Rep) methods has achieved significant performance improvement on traditional convolutional network. Most current Rep methods rely on prior knowledge…

cs.CV2022

MKIoU Loss: Towards Accurate Oriented Object Detection in Aerial Images

Xinyi Yu, Jiangping Lu, Mi Lin +1

Oriented bounding box regression is crucial for oriented object detection. However, regression-based methods often suffer from boundary problems and the inconsistency between loss…

cs.CV2021

Pedestrian Attribute Recognition in Video Surveillance Scenarios Based on View-attribute Attention Localization

Weichen Chen, Xinyi Yu, Linlin Ou

Pedestrian attribute recognition in surveillance scenarios is still a challenging task due to the inaccurate localization of specific attributes. In this paper, we propose a novel…

cs.CV2022

Real-time Rail Recognition Based on 3D Point Clouds

Xinyi Yu, Weiqi He, Xuecheng Qian +2

Accurate rail location is a crucial part in the railway support driving system for safety monitoring. LiDAR can obtain point clouds that carry 3D information for the railway enviro…

cs.RO2021

A Self-adaptive SAC-PID Control Approach based on Reinforcement Learning for Mobile Robots

Xinyi Yu, Yuehai Fan, Siyu Xu +1

Proportional-integral-derivative (PID) control is the most widely used in industrial control, robot control and other fields. However, traditional PID control is not competent when…

cs.RO2022

Multi-subgoal Robot Navigation in Crowds with History Information and Interactions

Xinyi Yu, Jianan Hu, Yuehai Fan +2

Robot navigation in dynamic environments shared with humans is an important but challenging task, which suffers from performance deterioration as the crowd grows. In this paper, mu…

cs.LG2021

Across-Task Neural Architecture Search via Meta Learning

Jingtao Rong, Xinyi Yu, Mingyang Zhang +1

Adequate labeled data and expensive compute resources are the prerequisites for the success of neural architecture search(NAS). It is challenging to apply NAS in meta-learning scen…

cs.RO2026

3DGSNav: Enhancing Vision-Language Model Reasoning for Object Navigation via Active 3D Gaussian Splatting

Wancai Zheng, Hao Chen, Xianlong Lu +2

Object navigation is a core capability of embodied intelligence, enabling an agent to locate target objects in unknown environments. Recent advances in vision-language models (VLMs…

cs.CV2023

CrossFusion: Interleaving Cross-modal Complementation for Noise-resistant 3D Object Detection

Yang Yang, Weijie Ma, Hao Chen +2

The combination of LiDAR and camera modalities is proven to be necessary and typical for 3D object detection according to recent studies. Existing fusion strategies tend to overly…

cs.CV2022

Conditional Generative Data-free Knowledge Distillation

Xinyi Yu, Ling Yan, Yang Yang +2

Knowledge distillation has made remarkable achievements in model compression. However, most existing methods require the original training data, which is usually unavailable due to…

cs.LG2022

RepNAS: Searching for Efficient Re-parameterizing Blocks

Mingyang Zhang, Xinyi Yu, Jingtao Rong +1

In the past years, significant improvements in the field of neural architecture search(NAS) have been made. However, it is still challenging to search for efficient networks due to…

cs.CV2021

Effective Model Compression via Stage-wise Pruning

Mingyang Zhang, Xinyi Yu, Jingtao Rong +1

Automated Machine Learning(Auto-ML) pruning methods aim at searching a pruning strategy automatically to reduce the computational complexity of deep Convolutional Neural Networks(d…

cs.RO2025

GSORB-SLAM: Gaussian Splatting SLAM benefits from ORB features and Transmittance information

Wancai Zheng, Xinyi Yu, Jintao Rong +3

The emergence of 3D Gaussian Splatting (3DGS) has recently ignited a renewed wave of research in dense visual SLAM. However, existing approaches encounter challenges, including sen…

cs.CV2021

Graph Pruning for Model Compression

Mingyang Zhang, Xinyi Yu, Jingtao Rong +1

Previous AutoML pruning works utilized individual layer features to automatically prune filters. We analyze the correlation for two layers from the different blocks which have a sh…

cs.RO2025

UP-SLAM: Adaptively Structured Gaussian SLAM with Uncertainty Prediction in Dynamic Environments

Wancai Zheng, Linlin Ou, Jiajie He +3

Recent 3D Gaussian Splatting (3DGS) techniques for Visual Simultaneous Localization and Mapping (SLAM) have significantly progressed in tracking and high-fidelity mapping. However,…

cs.LG2024

LoRAPrune: Structured Pruning Meets Low-Rank Parameter-Efficient Fine-Tuning

Mingyang Zhang, Hao Chen, Chunhua Shen +4

Large Language Models (LLMs), such as LLaMA and T5, have shown exceptional performance across various tasks through fine-tuning. Although low-rank adaption (LoRA) has emerged to ch…

cs.CL2024

Channel Merging: Preserving Specialization for Merged Experts

Mingyang Zhang, Jing Liu, Ganggui Ding +3

Lately, the practice of utilizing task-specific fine-tuning has been implemented to improve the performance of large language models (LLM) in subsequent tasks. Through the integrat…

cs.CV2024

Boosting Box-supervised Instance Segmentation with Pseudo Depth

Xinyi Yu, Ling Yan, Pengtao Jiang +4

The realm of Weakly Supervised Instance Segmentation (WSIS) under box supervision has garnered substantial attention, showcasing remarkable advancements in recent years. However, t…

cs.CV2021

Oriented Object Detection in Aerial Images Based on Area Ratio of Parallelogram

Xinyi Yu, Mi Lin, Jiangping Lu +1

Oriented object detection is a challenging task in aerial images since the objects in aerial images are displayed in arbitrary directions and are frequently densely packed. The mai…

cs.RO2021

A Self-adaptive LSAC-PID Approach based on Lyapunov Reward Shaping for Mobile Robots

Xinyi Yu, Siyu Xu, Yuehai Fan +1

To solve the coupling problem of control loops and the adaptive parameter tuning problem in the multi-input multi-output (MIMO) PID control system, a self-adaptive LSAC-PID algorit…