2 papers
cs.LG2026
DRiffusion: Draft-and-Refine Process Parallelizes Diffusion Models with Ease
Runsheng Bai, Chengyu Zhang, Yangdong Deng
Diffusion models have achieved remarkable success in generating high-fidelity content but suffer from slow, iterative sampling, resulting in high latency that limits their use in i…
cs.LG2024
SKIM: Any-bit Quantization Pushing The Limits of Post-Training Quantization
Runsheng Bai, Bo Liu, Qiang Liu
Large Language Models (LLMs) exhibit impressive performance across various tasks, but deploying them for inference poses challenges. Their high resource demands often necessitate c…