5 papers
Attribution-Guided and Coverage-Maximized Pruning for Structural MoE Compression
Yifu Ding, Jiacheng Wang, Ge Yang +4
Mixture-of-Experts (MoE) models scale compute efficiently, yet remain expensive to deploy due to their substantial memory footprint and inference overhead. Prior compression method…
SPA-Cache: Singular Proxies for Adaptive Caching in Diffusion Language Models
Wenhao Sun, Rong-Cheng Tu, Yifu Ding +4
While Diffusion Language Models (DLMs) offer a flexible, arbitrary-order alternative to the autoregressive paradigm, their non-causal nature precludes standard KV caching, forcing…
BWTA: Accurate and Efficient Binarized Transformer by Algorithm-Hardware Co-design
Yifu Ding, Xianglong Liu, Shenghao Jin +2
Ultra low-bit quantization brings substantial efficiency for Transformer-based models, but the accuracy degradation and limited GPU support hinder its wide usage. In this paper, we…
Diagonal-Tiled Mixed-Precision Attention for Efficient Low-Bit MXFP Inference
Yifu Ding, Xinhao Zhang, Jinyang Guo
Transformer-based large language models (LLMs) have demonstrated remarkable performance across a wide range of real-world tasks, but their inference cost remains prohibitively high…
VORTA: Efficient Video Diffusion via Routing Sparse Attention
Wenhao Sun, Rong-Cheng Tu, Yifu Ding +4
Video diffusion transformers have achieved remarkable progress in high-quality video generation, but remain computationally expensive due to the quadratic complexity of attention o…