activity
20242026
collaborators

8 papers

eess.IV2026

TaQ-DiT: Time-aware Quantization for Diffusion Transformers

Xinyan Liu, Huihong Shi, Yang Xu +1

Transformer-based diffusion models, dubbed Diffusion Transformers (DiTs), have achieved state-of-the-art performance in image and video generation tasks. However, their large model…

cs.CV2025

StripDet: Strip Attention-Based Lightweight 3D Object Detection from Point Cloud

Weichao Wang, Wendong Mao, Zhongfeng Wang

The deployment of high-accuracy 3D object detection models from point cloud remains a significant challenge due to their substantial computational and memory requirements. To addre…

cs.CV2025

A Memory-Efficient Framework for Deformable Transformer with Neural Architecture Search

Wendong Mao, Mingfan Zhao, Jianfeng Guan +2

Deformable Attention Transformers (DAT) have shown remarkable performance in computer vision tasks by adaptively focusing on informative image regions. However, their data-dependen…

cs.GR2025

CDM-QTA: Quantized Training Acceleration for Efficient LoRA Fine-Tuning of Diffusion Model

Jinming Lu, Minghao She, Wendong Mao +1

Fine-tuning large diffusion models for custom applications demands substantial power and time, which poses significant challenges for efficient implementation on mobile devices. In…

cs.AR2025

AccLLM: Accelerating Long-Context LLM Inference Via Algorithm-Hardware Co-Design

Yanbiao Liang, Huihong Shi, Haikuo Shao +1

Recently, large language models (LLMs) have achieved huge success in the natural language processing (NLP) field, driving a growing demand to extend their deployment from the cloud…

cs.AR2025

An Efficient Sparse Hardware Accelerator for Spike-Driven Transformer

Zhengke Li, Wendong Mao, Siyu Zhang +2

Recently, large models, such as Vision Transformer and BERT, have garnered significant attention due to their exceptional performance. However, their extensive computational requir…