57 citations · 75 across the 25 of their papers we have counts for
21 papers · 1 filter
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents
Ran Yan, Wei Fu, Jiale Li +21
LLM agents are rapidly being deployed in production, including coding assistants, customer-support chatbots, and scientific research assistants, yet they remain fundamentally stati…
TurboServe: Serving Streaming Video Generation Efficiently and Economically
Youhe Jiang, Haoxu Wang, Haotong Bao +5
Streaming video generation is emerging as a new serving workload in which users interact with long-lived sessions that generate video progressively, chunk by chunk. Unlike offline…
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
You Peng, Youhe Jiang, Wenshuang Li +5
Agentic LLM applications increasingly execute user requests as multi-step workflows involving planning, tool use, branching, refinement, and synthesis. In such settings, users expe…
HexiSeq: Accommodating Long Context Training of LLMs over Heterogeneous Hardware
Yan Liang, Youhe Jiang, Ran Yan +3
Long-context training of large language models (LLMs) is commonly distributed with Context Parallelism (CP) and Head Parallelism (HP), but existing training systems largely assume…
Autopoiesis: A Self-Evolving System Paradigm for LLM Serving Under Runtime Dynamics
Youhe Jiang, Ran Yan, You Peng +4
Modern Large Language Model (LLM) serving operates in highly volatile environments characterized by severe runtime dynamics, such as workload fluctuations and elastic cluster autos…
BOute: Cost-Efficient LLM Serving with Heterogeneous LLMs and GPUs via Multi-Objective Bayesian Optimization
Youhe Jiang, Fangcheng Fu, Eiko Yoneki
The rapid growth of large language model (LLM) deployments has made cost-efficient serving systems essential. Recent efforts to enhance system cost-efficiency adopt two main perspe…