8 papers
AdaFuse: Accelerating Dynamic Adapter Inference via Token-Level Pre-Gating and Fused Kernel Optimization
Qiyang Li, Rui Kong, Yuchen Li +5
The integration of dynamic, sparse structures like Mixture-of-Experts (MoE) with parameter-efficient adapters (e.g., LoRA) is a powerful technique for enhancing Large Language Mode…
DVD-Quant: Data-free Video Diffusion Transformers Quantization
Zhiteng Li, Hanxuan Li, Junyi Wu +6
Diffusion Transformers (DiTs) have emerged as the state-of-the-art architecture for video generation, yet their computational and memory demands hinder practical deployment. While…
FlexSpec: Frozen Drafts Meet Evolving Targets in Edge-Cloud Collaborative LLM Speculative Decoding
Yuchen Li, Rui Kong, Zhonghao Lyu +11
Deploying large language models (LLMs) in mobile and edge computing environments is constrained by limited on-device resources, scarce wireless bandwidth, and frequent model evolut…
B2LoRa: Boosting LoRa Transmission for Satellite-IoT Systems with Blind Coherent Combining
Yimin Zhao, Weibo Wang, Xiong Wang +6
With the rapid growth of Low Earth Orbit (LEO) satellite networks, satellite-IoT systems using the LoRa technique have been increasingly deployed to provide widespread Internet ser…
QuantFace: Efficient Quantization for Face Restoration
Jiatong Li, Libo Zhu, Haotong Qin +5
Diffusion models have been achieving remarkable performance in face restoration. However, the heavy computations hamper the widespread adoption of these models. In this work, we pr…
Low-bit Model Quantization for Deep Neural Networks: A Survey
Kai Liu, Qian Zheng, Kaiwen Tao +9
With unprecedented rapid development, deep neural networks (DNNs) have deeply influenced almost all fields. However, their heavy computation costs and model sizes are usually unacc…