Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
BlockServe: Block-Grained Continuous Batching for High-Throughput Diffusion LLM Serving
Yuanjie Zhu, Liangwei Yang, Ke Xu +4
Efficient serving of diffusion large language models (dLLMs) is hindered by convergence heterogeneity: when batching multiple requests, different sequences converge at different ra…
cs.LG2024
Do We Really Need Graph Convolution During Training? Light Post-Training Graph-ODE for Efficient Recommendation
Weizhi Zhang, Liangwei Yang, Zihe Song +4
The efficiency and scalability of graph convolution networks (GCNs) in training recommender systems (RecSys) have been persistent concerns, hindering their deployment in real-world…