8 papers
TaQ-DiT: Time-aware Quantization for Diffusion Transformers
Xinyan Liu, Huihong Shi, Yang Xu +1
Transformer-based diffusion models, dubbed Diffusion Transformers (DiTs), have achieved state-of-the-art performance in image and video generation tasks. However, their large model…
StripDet: Strip Attention-Based Lightweight 3D Object Detection from Point Cloud
Weichao Wang, Wendong Mao, Zhongfeng Wang
The deployment of high-accuracy 3D object detection models from point cloud remains a significant challenge due to their substantial computational and memory requirements. To addre…
A Memory-Efficient Framework for Deformable Transformer with Neural Architecture Search
Wendong Mao, Mingfan Zhao, Jianfeng Guan +2
Deformable Attention Transformers (DAT) have shown remarkable performance in computer vision tasks by adaptively focusing on informative image regions. However, their data-dependen…
CDM-QTA: Quantized Training Acceleration for Efficient LoRA Fine-Tuning of Diffusion Model
Jinming Lu, Minghao She, Wendong Mao +1
Fine-tuning large diffusion models for custom applications demands substantial power and time, which poses significant challenges for efficient implementation on mobile devices. In…
AccLLM: Accelerating Long-Context LLM Inference Via Algorithm-Hardware Co-Design
Yanbiao Liang, Huihong Shi, Haikuo Shao +1
Recently, large language models (LLMs) have achieved huge success in the natural language processing (NLP) field, driving a growing demand to extend their deployment from the cloud…
An Efficient Sparse Hardware Accelerator for Spike-Driven Transformer
Zhengke Li, Wendong Mao, Siyu Zhang +2
Recently, large models, such as Vision Transformer and BERT, have garnered significant attention due to their exceptional performance. However, their extensive computational requir…