activation sparsity 1diffusion models 1dynamic batching 1inference acceleration 1multimodal large language models 1sequence truncation 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.AI2026
Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation
Qicheng Zhao, Qi Sun, Zheyu Yan
The paper presents Seer, a training‑free approach that detects the true end of generated sequences in diffusion multimodal large language models by monitoring MLP activation sparsi…
cs.AI2026
ResilPhase: Plug-and-Play Phase Mapping and Noise-Resilient Macro-Trajectory Extrapolation for Diffusion Acceleration
Qicheng Zhao, Yu Li, Qi Sun +1
The adoption of powerful diffusion models is hindered by their significant inference latency. Recent ``cache-then-forecast'' schemes alleviate this issue by accelerating DiTs using…