3 papers
cs.AR2026
APT: Accelerating Diffusion Transformers via Attention Probability-Guided Pruning and Quantization
Sungyeob Yoo, Seeyeon Kim, Joonyong Park +2
Recent advances in generative AI have significantly increased the demand for high-resolution image and video generation, positioning diffusion models as a core technology. Among th…
cs.AR2026
Beyond Capacity: Scalable MoE LLM Inference via High-Bandwidth Flash with Direct GPU and HBM Paths
Seeyeon Kim, Juhyeong Jin, Joo-Young Kim
Modern mixture-of-experts (MoE) language models increasingly strain the capacity and cost efficiency of high-bandwidth memory (HBM), as rapidly growing expert weights must be provi…
cs.AR2026
MASQ: Accelerating Masked Diffusion via Stage-Wise Multi-Precision Quantization
Seeyeon Kim, Jaehun Lee, Sungyeob Yoo +1
Masked diffusion enables region-specific image synthesis but suffers from computational redundancy, since the entire image is processed each timestep even though only the masked re…