activity
20242026
collaborators
Showing 2025Show all

7 papers · 1 filter

cs.DC2025

ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments

Youhe Jiang, Fangcheng Fu, Xiaozhe Yao +4

Recent developments in large language models (LLMs) have demonstrated their remarkable proficiency in a range of tasks. Compared to in-house homogeneous GPU clusters, deploying LLM…

cs.DC2025

Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUs

Guoliang He, Youhe Jiang, Wencong Xiao +8

The scaling law for large language models (LLMs) depicts that the path towards machine intelligence necessitates training at large scale. Thus, companies continuously build large-s…

cs.DB2025

AutoIndexer: A Reinforcement Learning-Enhanced Index Advisor Towards Scaling Workloads

Taiyi Wang, Eiko Yoneki

Efficiently selecting indexes is fundamental to database performance optimization, particularly for systems handling large-scale analytical workloads. While deep reinforcement lear…

cs.DC2025

Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs

Youhe Jiang, Fangcheng Fu, Xiaozhe Yao +6

Recent advancements in Large Language Models (LLMs) have led to increasingly diverse requests, accompanied with varying resource (compute and memory) demands to serve them. However…

cs.LG2025

Navigating in High-Dimensional Search Space: A Hierarchical Bayesian Optimization Approach

Wenxuan Li, Taiyi Wang, Eiko Yoneki

Optimizing black-box functions in high-dimensional search spaces has been known to be challenging for traditional Bayesian Optimization (BO). In this paper, we introduce HiBO, a no…

cs.DB2025

A New Paradigm in Tuning Learned Indexes: A Reinforcement Learning Enhanced Approach

Taiyi Wang, Liang Liang, Guang Yang +2

Learned Index Structures (LIS) have significantly advanced data management by leveraging machine learning models to optimize data indexing. However, designing these structures ofte…