4 papers
ASSERT: Adaptive Stochastic Sampling for Robust Diffusion Models on Analog Compute-in-Memory Hardware
Yuannuo Feng, Yizhe Chen, Wenshuai Yao +4
Diffusion models achieve strong image generation quality but incur high iterative denoising costs. Analog compute-in-memory (CIM) can accelerate matrix-vector multiplications, yet…
Approximate Speculative Decoding
Yuannuo Feng, Zegang Peng, Yuxin Xie +5
Speculative decoding accelerates autoregressive generation by verifying a draft block with a target model in parallel. Under standard greedy verification, decoding stops at the fir…
When Guidance Goes Off-Scale: Recalibrating Diffusion Transformers under Analog Compute-in-Memory Nonidealities
Wenshuai Yao, Wenyong Zhou
Diffusion Transformers (DiTs) incur high memory traffic and energy costs because sampling repeatedly evaluates large denoisers dominated by linear operations. Analog compute-in-mem…
Selective KV Cache Protection for Noise-Resilient LLM Inference on Analog Compute-In-Memory Systems
Yuannuo Feng, Wenyong Zhou, Yuang Ma +5
Analog compute-in-memory (CIM) arrays have emerged as a promising substrate for energy-efficient LLM inference, particularly for weight-stationary computations in linear layers. Ho…