1 paper
Wei-Hsing Huang, Kiseok Lee, Ming-Yen Lee +7
Single-stream (batch-one) edge inference cannot amortize weight traffic across requests. Full-attention diffusion LLMs recompute the entire sequence at every step; native block dif…