activity
20242026
collaborators

5 papers

cs.DC2026

RolloutPipe: Overlapping Pipelined Rollout and Training in Disaggregated On-Policy LLM Reinforcement Learning

Rongjian Chen, Jianmin Hu, Kejiang Ye +1

Large language model (LLM) post-training for reasoning increasingly relies on reinforcement learning with verifiable rewards (RLVR), where models learn from ground-truth feedback o…

cs.DC2026

SwiftCache: Efficient LLM Serving for Multi-turn Conversations with Heterogeneous KV Cache Sharing

Jianmin Hu, Minxian Xu, Sa Wang +5

Multi-turn conversation is a fundamental scenario in LLM applications, widely used in chatbots and AI agents. As the conversation evolves, historical tokens accumulate continuously…

cs.DC2025

BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure

Yiyuan He, Minxian Xu, Jingfeng Wu +7

Large language models (LLMs) are increasingly deployed in AI infrastructure, driving the need for high throughput, resource efficient serving systems. Disaggregated LLM serving, wh…

cs.DC2025

BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs

Jianmin Hu, Minxian Xu, Kejiang Ye +1

In recent years, the Mixture-of-Experts (MoE) architecture has been widely applied to large language models (LLMs), providing a promising solution that activates only a subset of t…

cs.DC2024

Edge AI: A Taxonomy, Systematic Review and Future Directions

Sukhpal Singh Gill, Muhammed Golec, Jianmin Hu +12

Edge Artificial Intelligence (AI) incorporates a network of interconnected systems and devices that receive, cache, process, and analyze data in close communication with the locati…