3 papers
cs.LG2026
DRiffusion: Draft-and-Refine Process Parallelizes Diffusion Models with Ease
Runsheng Bai, Chengyu Zhang, Yangdong Deng
Diffusion models have achieved remarkable success in generating high-fidelity content but suffer from slow, iterative sampling, resulting in high latency that limits their use in i…
cs.CV2026
IDPruner: Harmonizing Importance and Diversity in Visual Token Pruning for MLLMs
Yifan Tan, Yifu Sun, Shirui Huang +4
Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities, yet they encounter significant computational bottlenecks due to the massive volume of visual tok…
cs.LG2024
AlignedKV: Reducing Memory Access of KV-Cache with Precision-Aligned Quantization
Yifan Tan, Haoze Wang, Chao Yan +1
Model quantization has become a crucial technique to address the issues of large memory consumption and long inference times associated with LLMs. Mixed-precision quantization, whi…