1 paper
Omin Kwon, Yeonjae Kim, Doyeon Kim +3
Block diffusion LLMs are an emerging paradigm for parallel language generation, but their KV caching makes memory access the dominant bottleneck in long-context inference. Sparse a…