180 citations · 241 across the 12 of their papers we have counts for
12 papers
Cloud Native System for LLM Inference Serving
Minxian Xu, Junhan Liao, Jingfeng Wu +3
Large Language Models (LLMs) are revolutionizing numerous industries, but their substantial computational demands create challenges for efficient deployment, particularly in cloud…
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling
Jingfeng Wu, Yiyuan He, Minxian Xu +3
The rise of large language models (LLMs) has created new opportunities across various fields but has also introduced significant challenges in resource management. Current LLM serv…
SealOS+: A Sealos-based Approach for Adaptive Resource Optimization Under Dynamic Workloads for Securities Trading System
Haojie Jia, Zhenhao Li, Gen Li +2
As securities trading systems transition to a microservices architecture, optimizing system performance presents challenges such as inefficient resource scheduling and high service…
LSRAM: A Lightweight Autoscaling and SLO Resource Allocation Framework for Microservices Based on Gradient Descent
Kan Hu, Minxian Xu, Kejiang Ye +1
Microservices architecture has become the dominant architecture in cloud computing paradigm with its advantages of facilitating development, deployment, modularity and scalability.…
DRPC: Distributed Reinforcement Learning Approach for Scalable Resource Provisioning in Container-based Clusters
Haoyu Bai, Minxian Xu, Kejiang Ye +2
Microservices have transformed monolithic applications into lightweight, self-contained, and isolated application components, establishing themselves as a dominant paradigm for app…
An Interference-aware Approach for Co-located Container Orchestration with Novel Metric
Xiang Li, Linfeng Wen, Minxian Xu +1
Container orchestration technologies are widely employed in cloud computing, facilitating the co-location of online and offline services on the same infrastructure. Online services…