2 papers
cs.CL2025
FlashDLM: Accelerating Diffusion Language Model Inference via Efficient KV Caching and Guided Diffusion
Zhanqiu Hu, Jian Meng, Yash Akhauri +4
Diffusion language models offer parallel token generation and inherent bidirectionality, promising more efficient and powerful sequence modeling compared to autoregressive approach…
cs.DC2025
EcoServe: Designing Carbon-Aware AI Inference Systems
Yueying Li, Zhanqiu Hu, Esha Choukse +3
The rapid increase in LLM ubiquity and scale levies unprecedented demands on computing infrastructure. These demands not only incur large compute and memory resources but also sign…