Publications (36)
Real-time Stereo-based 3D Object Detection for Streaming Perception
Changcai Li, Zonghua Gu, Gang Chen +3
The ability to promptly respond to environmental changes is crucial for the perception system of autonomous driving. Recently, a new task called streaming perception was proposed.…
Multi-Teacher Knowledge Distillation with Reinforcement Learning for Visual Recognition
Chuanguang Yang, Xinqiang Yu, Han Yang +4
Multi-teacher Knowledge Distillation (KD) transfers diverse knowledge from a teacher pool to a student network. The core problem of multi-teacher KD is how to balance distillation…
Echo-Forcing: A Scene Memory Framework for Interactive Long Video Generation
Mingqiang Wu, Weilun Feng, Zhefeng Zhang +8
Autoregressive video diffusion models enable open-ended generation through local attention and KV caching. However, existing training-free long-video optimization methods mainly fo…
QuantSparse: Comprehensively Compressing Video Diffusion Transformer with Model Quantization and Attention Sparsification
Weilun Feng, Chuanguang Yang, Haotong Qin +8
Diffusion transformers exhibit remarkable video generation capability, yet their prohibitive computational and memory costs hinder practical deployment. Model quantization and atte…
Ships, Splashes, and Waves on a Vast Ocean
Libo Huang, Ziyin Qu, Xun Tan +3
The simulation of large open water surface is challenging using a uniform volumetric discretization of the Navier-Stokes equations. Simulating water splashes near moving objects, w…
Wavelet-based Mamba with Fourier Adjustment for Low-light Image Enhancement
Junhao Tan, Songwen Pei, Wei Qin +3
Frequency information (e.g., Discrete Wavelet Transform and Fast Fourier Transform) has been widely applied to solve the issue of Low-Light Image Enhancement (LLIE). However, exist…
Lifelong Generative Learning via Knowledge Reconstruction
Libo Huang, Zhulin An, Xiang Zhi +1
Generative models often incur the catastrophic forgetting problem when they are used to sequentially learning multiple tasks, i.e., lifelong generative learning. Although there are…
Relational Diffusion Distillation for Efficient Image Generation
Weilun Feng, Chuanguang Yang, Zhulin An +4
Although the diffusion model has achieved remarkable performance in the field of image generation, its high inference delay hinders its wide application in edge devices with scarce…
Continual Learning in the Frequency Domain
Ruiqi Liu, Boyu Diao, Libo Huang +3
Continual learning (CL) is designed to learn new tasks while preserving existing knowledge. Replaying samples from earlier tasks has proven to be an effective method to mitigate th…
Exemplar-Free Class Incremental Learning via Incremental Representation
Libo Huang, Zhulin An, Yan Zeng +3
Exemplar-Free Class Incremental Learning (efCIL) aims to continuously incorporate the knowledge from new classes while retaining previously learned information, without storing any…
MPQ-DMv2: Flexible Residual Mixed Precision Quantization for Low-Bit Diffusion Models with Temporal Distillation
Weilun Feng, Chuanguang Yang, Haotong Qin +10
Diffusion models have demonstrated remarkable performance on vision generation tasks. However, the high computational complexity hinders its wide application on edge devices. Quant…
WorldCache: Accelerating World Models for Free via Heterogeneous Token Caching
Weilun Feng, Guoxin Fan, Haotong Qin +10
Diffusion-based world models have shown strong potential for unified world simulation, but the iterative denoising remains too costly for interactive use and long-horizon rollouts.…
TJ4DRadSet: A 4D Radar Dataset for Autonomous Driving
Lianqing Zheng, Zhixiong Ma, Xichan Zhu +9
The next-generation high-resolution automotive radar (4D radar) can provide additional elevation measurement and denser point clouds, which has great potential for 3D sensing in au…
Fast Convolution based on Winograd Minimum Filtering: Introduction and Development
Gan Tong, Libo Huang
Convolutional Neural Network (CNN) has been widely used in various fields and played an important role. Convolution operators are the fundamental component of convolutional neural…
DFM: Difference Feature Modeling with Text-Guided Gated Contrastive Loss for Remote Sensing Image Change Captioning
Yelin Wang, Zijia Song, Chuanguang Yang +4
The primary goal of Remote Sensing Image Change Captioning (RSICC) is to automatically generate descriptions of changes between remote sensing images captured at different time poi…
eTag: Class-Incremental Learning with Embedding Distillation and Task-Oriented Generation
Libo Huang, Yan Zeng, Chuanguang Yang +3
Class-Incremental Learning (CIL) aims to solve the neural networks' catastrophic forgetting problem, which refers to the fact that once the network updates on a new task, its perfo…
Online Policy Distillation with Decision-Attention
Xinqiang Yu, Chuanguang Yang, Chengqing Yu +3
Policy Distillation (PD) has become an effective method to improve deep reinforcement learning tasks. The core idea of PD is to distill policy knowledge from a teacher agent to a s…
Quantized Visual Geometry Grounded Transformer
Weilun Feng, Haotong Qin, Mingqiang Wu +8
Learning-based 3D reconstruction models, represented by Visual Geometry Grounded Transformers (VGGTs), have made remarkable progress with the use of large-scale transformers. Their…
Fast-SAM3D: 3Dfy Anything in Images but Faster
Weilun Feng, Mingqiang Wu, Zhiliang Chen +10
SAM3D enables scalable, open-world 3D reconstruction from complex scenes, yet its deployment is hindered by prohibitive inference latency. In this work, we conduct the \textbf{firs…
Multi-party Collaborative Attention Control for Image Customization
Han Yang, Chuanguang Yang, Qiuli Wang +4
The rapid advancement of diffusion models has increased the need for customized image generation. However, current customization methods face several limitations: 1) typically acce…
Multi-Scale Cost Volumes Cascade Network for Stereo Matching
Xiaogang Jia, Wei Chen, Zhengfa Liang +3
Stereo matching is essential for robot navigation. However, the accuracy of current widely used traditional methods is low, while methods based on CNN need expensive computational…
SQ-VDiT: Accurate Quantized Video Diffusion Transformer with Salient Data and Sparse Token Distillation
Weilun Feng, Haotong Qin, Chuanguang Yang +7
Diffusion transformers have emerged as the mainstream paradigm for video generation models. However, the use of up to billions of parameters incurs significant computational costs.…
PrePrompt: Predictive prompting for class incremental learning
Libo Huang, Zhulin An, Chuanguang Yang +5
Class Incremental Learning (CIL) based on pre-trained models offers a promising direction for open-world continual learning. Existing methods typically rely on correlation-based st…
A Survey on Causal Reinforcement Learning
Yan Zeng, Ruichu Cai, Fuchun Sun +2
While Reinforcement Learning (RL) achieves tremendous success in sequential decision-making problems of many domains, it still faces key challenges of data inefficiency and the lac…
A Unified Model of Feature Extraction and Clustering for Spike Sorting
Libo Huang, Lu Gan, Bingo Wing-Kuen Ling
Spike sorting plays an irreplaceable role in understanding brain codes. Traditional spike sorting technologies perform feature extraction and clustering separately after spikes are…
EstRTL: Functional Estimation Guided RTL Code Generation
Qi Xiong, Renzhi Chen, Bowei Wang +3
Optimizing register transfer level (RTL) code is of vital importance in hardware design. Large language models (LLMs) provide new methods for the automatic generation and optimizat…
DA-Nav: Direction-Aware City-Scale Vision-Language Navigation
Ye Yuan, Kehan Chen, Xinqiang Yu +7
The paper presents DA-Nav, a direction-aware vision‑language navigation system that uses commercial map directions and reformulates navigation as discrete spatial grounding on an e…
Parameterized Prompt for Incremental Object Detection
Zijia An, Boyu Diao, Ruiqi Liu +5
Recent studies have demonstrated that incorporating trainable prompts into pretrained models enables effective incremental learning. However, the application of prompts in incremen…
Teacher-Guided Student Self-Knowledge Distillation Using Diffusion Model
Yu Wang, Chuanguang Yang, Zhulin An +6
Existing Knowledge Distillation (KD) methods often align feature information between teacher and student by exploring meaningful feature processing and loss functions. However, due…
Q-VDiT: Towards Accurate Quantization and Distillation of Video-Generation Diffusion Transformers
Weilun Feng, Chuanguang Yang, Haotong Qin +8
Diffusion transformers (DiT) have demonstrated exceptional performance in video generation. However, their large number of parameters and high computational complexity limit their…
CLIP-KD: An Empirical Study of CLIP Model Distillation
Chuanguang Yang, Zhulin An, Libo Huang +5
Contrastive Language-Image Pre-training (CLIP) has become a promising language-supervised visual pre-training framework. This paper aims to distill small CLIP models supervised by…
IOR: Inversed Objects Replay for Incremental Object Detection
Zijia An, Boyu Diao, Libo Huang +3
Existing Incremental Object Detection (IOD) methods partially alleviate catastrophic forgetting when incrementally detecting new objects in real-world scenarios. However, many of t…
Confounded Causal Imitation Learning with Instrumental Variables
Yan Zeng, Shenglan Nie, Feng Xie +3
Imitation learning from demonstrations usually suffers from the confounding effects of unmeasured variables (i.e., unmeasured confounders) on the states and actions. If ignoring th…
Low-redundancy Distillation for Continual Learning
RuiQi Liu, Boyu Diao, Libo Huang +4
Continual learning (CL) aims to learn new tasks without erasing previous knowledge. However, current CL methods primarily emphasize improving accuracy while often neglecting traini…
MPQ-DM: Mixed Precision Quantization for Extremely Low Bit Diffusion Models
Weilun Feng, Haotong Qin, Chuanguang Yang +7
Diffusion models have received wide attention in generation tasks. However, the expensive computation cost prevents the application of diffusion models in resource-constrained scen…
Efficient Continual Learning through Frequency Decomposition and Integration
Ruiqi Liu, Boyu Diao, Libo Huang +4
Continual learning (CL) aims to learn new tasks while retaining past knowledge, addressing the challenge of forgetting during task adaptation. Rehearsal-based methods, which replay…